Srijith Ravikumar is a Principal Engineer at Amazon building AI-powered recommendation systems at scale. Published researcher at AAAI

Your best-sellers will show up when a customer asks an AI assistant for a recommendation. The product you launched last month will not, and the reason is not the one most brands have been sold.

The story going around is that the model has never seen it. Every large language model is trained up to a fixed date, so last month’s launch is not in its parameters. Tidy, and for shopping, mostly wrong. The assistants your customers use do not answer product questions from memory alone. They retrieve. OpenAI takes catalog feeds directly from merchants and asks for a full snapshot at least daily. Where a merchant is feeding an assistant, the new product is already in the index. Absence from the answer is not absence from the catalog.

Your new product is missing something harder to manufacture. Ask what an assistant needs before it will name a product, and the answer is evidence: ratings, review volume, sales history, independent coverage it can cite as a reason. Studying Amazon’s shopping agent from the outside, Profitero, Mars United Commerce and Publicis Commerce reported that content optimization moved a product’s share of recommendations only after it cleared a baseline of credibility, which in their data sat near a top-15,000 best-seller rank, a 4.6-star rating and substantial review volume.

That is not an Amazon quirk. A preregistered audit across 12 open-weight and proprietary models, in hotels rather than retail, found a top guest rating raised an assistant’s odds of recommending a property by 31.6 percentage points and review volume by another 8.3, concluding that this “advantages established properties over new entrants with thin review histories.” The same audit found management response to reviews, a cue the optimization industry actively sells, moved nothing at all.

Anyone who has built a recommender system knows this problem by name. The name is cold start, and it was never about whether a model had read about the product. It is about interaction history, the behavioral record a ranker trains on. An item that nobody has bought, rated or reviewed gives a collaborative ranker nothing to work with, so it leans on what it can measure, which is precisely what a new product has not accumulated yet. Generative shopping did not invent the problem. It inherited it, and it made the penalty visible in a single sentence instead of buried on page three.

That sentence is the whole shift. A search engine returns 10 blue links and lets the shopper choose. A shopping assistant returns two or three named products, with reasons. The shelf collapses from a page to a sentence, and being on page one no longer matters if you are not in the sentence. The channel is still small next to paid search or email, which is exactly why the growth curve is the part to watch: In July 2026, AI-referred traffic to U.S. retail sites was up 62% year over year, and those visits converted at a rate 60% higher than non-AI traffic did. That was the eleventh consecutive month in which AI traffic converted better.

Now look at what this does to a launch. The products you most want to push are the ones with the thinnest evidence behind them: the new line, the seasonal release, the SKU you just built a campaign around. Your back catalog carries years of ratings and rank. The products carrying next quarter’s revenue have none.

The penalty lands hardest exactly where the spending is.

None of this is fixed by a better model, and none of it is fixed by one department. Three things decide whether a launch is visible, and not one of them sits where you would expect.

Put a clock on freshness.

Time the gap between a product going live and its listing being complete and correct in every feed an assistant can reach. In my experience building these systems, that interval is rarely instrumented at all, and teams that measure it tend to find an approval or content-enrichment step upstream of the feed holding the record for days. That one belongs to operations.

Seed the evidence, not just the listing.

Ratings, reviews and independent coverage are ranking inputs, on the same footing as price and availability. Sampling and sanctioned early-review programs are what carry a new product from invisible to recommended. Note the word sanctioned. Incentivized reviews are prohibited on most major marketplaces and the compliant channels are capped, which makes this a launch-quarter program. That one is merchandising’s.

Measure inclusion for launch cohorts.

Track how often your products appear in assistant answers, and track launch-window products apart from the catalog average, because the aggregate hides the cohort at risk. Assistant answers vary between runs, so the instrument is a sampled panel of buying-intent prompts producing a distribution, not a rank you look up. Tools exist for the tracking. Few cut it by launch cohort, which is where the signal is. That one is marketing’s.

Every product you launch ships with a disadvantage that has nothing to do with how good the product is. It does not lift on its own. Best-seller rank is built from sales, not from ratings, as Amazon states plainly, so a product that is not being recommended is not making the sales that earn the rank that earns the recommendation.

Recommender researchers have a name for that loop, too, “popularity bias,” and the Profitero research found products that went unrecommended at the start typically stayed unrecommended even after their content was optimized. Cold start is a problem you work, not one you wait out. The question is whether you spend the launch window driving that disadvantage down deliberately, or spend the launch budget driving demand toward a shelf that will not say your name.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Share.
Leave A Reply

Exit mobile version