How ChatGPT picks products: we decoded a real shopping answer

Ask ChatGPT for a product and sometimes you get plain text, sometimes a rich carousel with prices and ratings. We captured and decoded a real streamed shopping answer to show the machinery underneath: how it classifies the question, where the products actually come from, what the model 'sees' about each one, and what it takes to get your product on that shelf.

Ask ChatGPT "recommend me a sunscreen" and one of two things happens. Sometimes you get a paragraph of prose. Sometimes you get a rich carousel - product cards with photos, prices, star ratings and review counts you can scroll through like a shelf. Same kind of question, two completely different answers. What decides which one you get, and how does the engine choose the products it puts in front of millions of shoppers?

Most writing on this topic guesses. We didn't want to guess, so we did the boring, useful thing: captured a real streamed shopping answer in the browser's network tab and took it apart, field by field. No insider access, no leaked docs - just the raw data the page itself receives and renders. What follows is the whole machine, laid out in the order it actually runs: how ChatGPT decides your question is a shopping question, how it rewrites it, where the products really come from, what the model "sees" about each one, how it ranks them, and the second, separate shelf of cited sources running alongside. Then what every piece of that means if you want your brand on the shelf.

One worked example runs through this piece: a single "best sunscreen" answer captured in Turkey in June 2026, on the free tier. Treat it as a snapshot, not a study - the specific brands and prices change by the day, the country and the user. The mechanics, on the other hand, are stable, and they're what you can actually build a strategy on.

Updated October 2026: in August ChatGPT changed how it picks brands for prose answers - it now names candidates before it searches, then checks their own websites. The carousel teardown below still holds; what changed, and what it means for you, is in the last section.

The whole pipeline, in one breath

Before we zoom in, here's the entire journey from your question to that scrollable shelf. Six stages, each of which left a fingerprint in the captured data:

  1. Classify - the turn is tagged with an intent. Ours came back shopping.
  2. Rewrite - your casual question is rephrased into a precise, attribute-rich search query.
  3. Invoke a tool - a browse/shopping tool is called to go fetch live products.
  4. Retrieve - products come back from merchant feeds and aggregators, each as a structured card.
  5. Rank & render - the products are ordered and injected into the answer as a carousel.
  6. Cite - in parallel, a normal web search attaches a list of source links under the answer.

Two things about that list are worth holding onto. First, the product carousel and the source citations are two different systems with two different supply chains - you can win one and lose the other. Second, almost everything that determines whether you appear is decided in stages 1–4, before the model writes a single word of prose. Let's walk through each.

How we captured it (and how you can too)

You don't need a proxy, a certificate or any paid tool. A shopping answer is delivered to your browser as a long stream of text events, and your browser's built-in developer tools can read it. The steps:

  1. Open ChatGPT in a desktop browser (Chrome or Safari), then open developer tools - F12, or right-click → Inspect.
  2. Go to the Network tab and leave it recording. This logs every request the page makes.
  3. Ask a shopping question - something with clear buying intent, like "recommend me a sunscreen." Watch for the carousel to appear.
  4. Find the conversation request in the network log - it's the long-lived one that streams the answer back.
  5. Open its Response and you'll see the raw stream: hundreds of small data: events arriving one after another.
  6. Read the fields. The interesting ones are the product entries, the rewritten queries, the source list and the metadata block. That's the whole teardown.

That stream is delivered as server-sent events - the same technology a live sports ticker uses. The answer isn't sent in one piece; it's built up from a flood of tiny deltas: "append this word," "add this product to the metadata," "patch this field." The text you watch typing out on screen is your browser replaying those deltas in real time. Crucially, the product data and source data arrive as their own structured deltas alongside the prose - which is exactly why we can lift them out cleanly.

The hidden tokens that place the cards

One detail surprises everyone the first time they see it. The carousel doesn't live in a separate part of the response - it's stitched into the prose using invisible marker characters. In the raw text, where a card should appear, the model emits something like entity["turn0product0", "turn0product3", "turn0product7"], wrapped in non-printing Unicode characters your browser intercepts. When it sees that marker, it doesn't print the text - it swaps in the rendered product cards for those IDs. A second marker, cite turn0product0, drops a little citation badge inline.

So the model's real job isn't to "write about products." It's to retrieve a set of product objects, give each an ID like turn0product0, and then decide where in its sentence to drop the markers that pull those cards in. The brand never controls the prose; it controls whether its product is one of the objects available to be referenced in the first place. (These tokens leak into copied text often enough that people google them - we've written a short explainer on what turn0product7 actually means.)

Stage 1: ChatGPT decides your question is "shopping"

The first meaningful thing in the response isn't text - it's a classification. Every turn is tagged with what kind of request it is, and our captured answer carried one unambiguous label in its metadata: turn_use_case: "shopping". That tag is the fork in the road. Once the turn is classified as commercial intent, ChatGPT stops behaving like a chatbot and starts behaving like a storefront - it reaches for a tool and returns structured products instead of a paragraph.

This is the single most important thing to internalise, because it sits upstream of everything else: whether you get a carousel at all is decided by intent classification, not by the exact words you typed. "Best running shoes for flat feet" trips the shopping classifier; "how do I lace running shoes" usually doesn't. Same product category, opposite outcome. If your buyers phrase their questions in ways that don't read as buying intent - "is X worth it," "how do I choose a Y" - you may never reach the product shelf at all, and you'll be competing in the prose answer instead, where the rules are different.

The practical implication is that the question itself is a battleground. Part of getting recommended is understanding which phrasings your category triggers the carousel on, and making sure you're strong for those. That's not something you control directly - but it's something you can measure, and measuring which prompts produce a shelf in your category is step one of any serious effort.

Stage 2: It rewrites your question before it shops

The shopper didn't get searched for what they asked. Our captured query was a casual "recommend me a sunscreen," but the engine rewrote it into something far more specific before going to look. The captured rewrite, stored in the metadata as the model's search query, was roughly "sunscreen SPF50 facial sunscreen." The carousel was built for that rewritten query - not the words the human typed.

Read that again, because it's a quiet bombshell for anyone doing product marketing. You are not competing for the literal question. You're competing for the engine's machine-tightened version of it, which silently injects the attributes it has decided matter - here, the SPF level and the "facial" use case. The engine took a vague request and made three product decisions on the shopper's behalf before any brand was even considered.

If your product data speaks in those attributes - if your titles, specs and structured data actually say "SPF50" and "face" - you're legible to the rewrite and eligible to be retrieved. If your data is vaguer than the rewrite, you can be the objectively best answer to the human's question and still miss the query the engine actually ran. The lesson: write your product data in the language of attributes and use cases, not just brand and benefit.

Stage 3: The browse tool goes and fetches

With an intent and a rewritten query in hand, the model called a tool. Our capture named it plainly in the metadata: a browse/shopping tool was invoked (tool_invoked: true), and the model itself was identified by its internal slug. This is the moment the system stops relying on what it already "knows" and goes to look something up - which is the entire reason a brand can influence the outcome at all.

It's worth dwelling on why this is good news. A pure language model can only recommend what it absorbed in training, which is frozen, months stale, and impossible to influence after the fact. A tool-using model recommends what it can fetch right now. That means accurate, fresh, well-structured data about your products is not a nice-to-have - it's the only lever that exists. The engine outsourced the actual product knowledge to live sources, and live sources are something you can feed.

Stage 4: The products don't come from the model's memory

Here's the part most people get wrong, and it's the heart of the teardown. The model is not "remembering" brands from its training and listing its favourites. Each product card in the capture arrived as structured data, retrieved at answer time, and stamped with provider and source markers showing it came from merchant feeds and product aggregators - not from the model's weights. In the raw fields, each product carried anonymised source tags (provider and metadata-source codes) that identify the commerce pipelines it was pulled from. ChatGPT was reading a live catalogue and selecting from it.

Here are the three products it fully retrieved for that one query, exactly as the fields described them:

Pos.ProductPriceRatingReviewsMerchants
1La Roche-Posay Anthelios UVMune 400 SPF50+TRY 862.434.814,000DermoEczanem.com + others
2Avène Ultra Fluid Invisible SPF50TRY 720.004.6937Sachane + others
3SVR Sun Secure Aqua Fluid SPF50+TRY 608.304.144DermoEczanem.com + others

(A fourth, an L'Oréal Paris SPF50+, was referenced in the prose without a full card.) Notice what these have in common: they're carried by Turkish dermo-pharmacy retailers, priced in lira, and described in local terms. This wasn't a global brand-fame contest; it was a retrieval from feeds that happened to be well-populated for this market. A smaller brand with clean feed data sat on the same shelf as giants.

That reframes the whole game. You don't earn the shelf by being famous enough for the model to recall you. You earn it by being present and clean in the feeds the engine reads - accurate prices, real availability, complete attributes, carried by retailers the aggregators ingest. It's closer to merchandising and feed management than to brand awareness. The brand-building still matters for the reviews and reputation that feed the ranking - but the entry ticket is data.

Stage 5: What the model actually "sees" about your product

Buried in the stream is the most revealing field of all - the compact text snapshot the model reads for each product before deciding anything. Not your beautifully designed product page. Not your campaign photography or your brand story. This, and only this:

FieldWhat the model saw for the top-ranked product
TitleLa Roche-Posay Anthelios UVMune 400 SPF50+
DescriptionNone
PriceTRY 862.43
Number of Reviews14,000
Rating4.8
Featured TagNone
MerchantsDermoEczanem.com + others

That's the entire dossier. A title, a price, a review count, a rating, the merchants who carry it - and two blank fields. Every None is a missing signal the brand failed to supply. The striking part: this product had no description and no featured tag and still won position one, because its review count and rating were overwhelming. The lesson isn't "descriptions don't matter" - it's that the engine ranks on the structured fields it can actually read, and the brands that fill all of them hand it the most reasons to choose them. Imagine the same card with a sharp description and a "dermatologist-recommended" featured tag: same product, more surface area to win on.

Stage 6: Position is a ranking, and reviews dominate it

The carousel isn't a random gallery - it's an ordered ranking, and position one is the recommendation the overwhelming majority of shoppers act on. So what decided the order? In our snapshot, the ratings were clustered tightly together, while the review counts were not even close:

Star ratings were nearly level across the three products - they barely separate first from third.
Review volume, by contrast, was a landslide - and it tracked the ranking almost perfectly. One captured “best sunscreen” carousel (TR, June 2026); a single snapshot, not a study.

Put the two charts side by side and the story writes itself: when the quality scores are close, the weight of social proof behind them is what separates the winner from the also-rans. Fourteen thousand reviews versus forty-four is not a rounding difference; it's a different universe of credibility. One reading isn't a law - LLM shopping results fluctuate, and a different day, country or phrasing would reshuffle the deck - but the shape rhymes with everything the captured fields imply. If you're trying to climb the carousel, "accumulate more genuine reviews" is almost never the wrong answer.

The other shelf: organic citations

Running in parallel with the product retrieval was an ordinary web search, and its results were attached to the answer as a separate list of sources. These are not the product cards - they're the cited links, and in our capture they came from a different pipeline entirely (the fields tagged them with their own search source). The eight domains it surfaced:

Cited source domains (organic search track)
watsons.com.tr
gratis.com
dermoeczanem.com
eveline.com.tr
dermo.com.tr
kbrfarma.com
dermoak.com
cerave.com.tr

Two details matter here. First, every one of those links was stamped with a utm_source=chatgpt.com tracking parameter - which means any clicks back to those sites show up in their analytics as ChatGPT referral traffic. If you've ever wondered how to prove AI is sending you visitors, that parameter is the receipt - and here's how to find that traffic in GA4. Second, and more strategically: this is a completely separate shelf from the carousel. A brand can be cited as a source here without appearing as a product card, or appear as a card without being cited. Winning AI shopping means competing on both - and they reward different things. The carousel rewards clean feed data and reviews; the citation list rewards the same content and authority signals that earn citations across every AI answer, shopping or not.

The single biggest mental-model upgrade from this teardown: "AI shopping visibility" is not one number. It's at least two - are you a product card, and are you a cited source - and they have different owners, different fixes and different supply chains. Measuring them as one blurs exactly the distinction you need to act on.

Why the same question gives different answers

If you run our exact query yourself, you'll almost certainly get a different result - maybe even plain prose with no carousel at all. That's not a bug, and it's the thing brands most need to make peace with. Several knobs move between any two runs:

  • Non-determinism. Language models sample their output; the same prompt on the same model can rank brands differently minute to minute.
  • Plan and model. Our capture ran on the free tier with a specific model build. A paid tier or a newer model can classify intent differently and retrieve from different surfaces.
  • Country and language. Our answer was steeped in the Turkish market - lira prices, local pharmacies. The same question in another country pulls from that market's feeds.
  • Phrasing. As we saw, a small wording change can flip whether the shopping classifier even fires.

The honest consequence: never trust a single screenshot. One answer is an anecdote. The only reliable picture comes from running the same prompts repeatedly over time and watching the averages and trends - which is precisely why the tools in this space exist, and why "I asked once and we showed up" is not a strategy.

Translating the teardown into metrics you can track

Every field we pulled apart maps onto a number a brand can actually monitor. Lined up, the raw stream becomes a scorecard:

What we found in the streamThe metric it becomes
Is your product a card at all?Visibility / presence rate
Were you in position one?Win rate (first recommendation)
Where in the carousel?Average position
The price shown on your cardMentioned price (and price-drift over time)
You vs. the other cardsShare of voice vs. competitors
Were you a cited source?Citation rate (the second shelf)

This is the bridge from "interesting teardown" to "thing you manage." A single capture gives you one row of that scorecard for one query on one day; tracking turns it into a trend you can move. If you want the broader version of which numbers are worth the effort, we go deep in the KPIs actually worth tracking.

What this means if you sell something

Strip away the mechanics and the to-do list is unglamorous, concrete, and very doable. Every item below maps directly to a stage of the teardown:

  • Get into the feeds, cleanly. The carousel is sourced from merchant feeds and aggregators (Stage 4). If your products aren't in them - or are in them with stale prices and thin attributes - you are simply not selectable. This is upstream of everything else; fix it first.
  • Fill every field. Title, price, availability, description, and especially structured product data (Stage 5). The model ranks on what it can read; every blank field is a point you chose not to score.
  • Stack real reviews. Review count and rating are visible ranking inputs and, in our snapshot, the single strongest separator of position (Stage 6). They're also the evidence the model trusts. This is the slowest lever and the highest-leverage one.
  • Speak in attributes. The engine rewrites vague questions into attribute-rich queries (Stage 2). Data that explicitly names its SPF, its size, its use case and its format survives the rewrite; data that hides behind brand language doesn't.
  • Compete on price and availability honestly. Both are shown on the card and both are plausible ranking inputs. A card with a stale price or "out of stock" is a card the engine has a reason to skip.
  • Win the citation shelf too. The cited-sources list is a separate prize with its own rules - earn it with the authority and content work that wins AI citations generally.
  • Win the question, not just the topic. If your buyers' natural phrasing doesn't trip the shopping classifier (Stage 1), work the surrounding shopping playbook so you still show up in the prose answer.

The myths this teardown kills

  • "The AI recommends brands it remembers being famous." No - it retrieves from live feeds at answer time. Fame helps only insofar as it produces reviews and citations.
  • "You can pay for placement." Not in the organic carousel - it's unsponsored and ranked on merit signals. ChatGPT now tests labeled ads, but OpenAI says they sit apart from the answer and don't influence it. You buy your way onto the shelf with data quality, not ad spend.
  • "It's just SEO with a new coat of paint." Partly - the citation shelf rewards SEO-like authority. But the product carousel rewards feed completeness and reviews, which most SEO programs never touch.
  • "We showed up once, so we're winning." One run is noise. Visibility here is a distribution, not a fact, and only trends over many runs are trustworthy.

Beyond ChatGPT

Everything here pays off well beyond one engine. The same clean feeds, structured data and reviews feed Google's shopping surfaces, Perplexity's product answers, and the buying agents coming next - which raise the stakes further, because they remove the human who used to forgive a thin product page for a good product. Treat ChatGPT's carousel as the first, most-developed instance of a pattern - AI-mediated buying - rather than a one-off integration to chase. The brands that get their data right while these surfaces are still early and ranked purely on merit will establish relevance that compounds as the shelves get crowded.

None of this is a hack. It's the same discipline that wins AI product recommendations generally: be present in the data, be complete, be well-reviewed, be accurate. The teardown just shows you, concretely and from the wire, why those four things are the whole game - because that compact structured snapshot is all the model ever sees, and you are the only one who can make it better.

What changed since this capture (October 2026)

The capture above is from June 2026. Three things have been reported since that affect how ChatGPT picks brands - mostly for the prose answer around the shelf. None is official OpenAI documentation, so we label each source.

  • It names candidates before it searches. One independent analysis read the raw traffic of 27 conversations on one ChatGPT Plus account (July 2026): in 21 of them, ChatGPT's first search already contained brands the user never typed. A brand ChatGPT named in its own query reached the final answer 68.9% of the time; a brand it only found during the search reached it 2.1%. One account, so the author calls every percentage directional - but the shape matters: the shortlist is often decided from memory, and search is used to confirm it.
  • It verifies on the brand's own site. On August 8, 2026 the share of ChatGPT's search queries using the site: operator jumped from 0.37% to 16.8% in a day, per Promptwatch (vendor, live-interface data). A separate Peec AI analysis (vendor) found 84% of the domains those queries target are branded - your own pages are where ChatGPT now checks what it is about to say.
  • It has its own index. Peec AI reports (vendor, October 7, 2026) that ChatGPT runs a family of its own retrieval indexes - web, news, YouTube, local, shopping and more - rather than relying only on Bing or Google.

For a merchant, that adds two jobs to the list above. Getting into the candidate set is a memory problem: reviews, third-party roundups and mentions across the web are what put your brand in ChatGPT's first query. Winning the verification step is a first-party problem: product, pricing and help pages that state the facts plainly in HTML, so the site: check confirms you instead of finding nothing. The carousel mechanics - feeds, filled fields, reviews - are unchanged by any of this, and OpenAI's new ad formats sit outside the organic shelf.


Curious whether ChatGPT puts your products on the shelf - or your competitors'? A free Zene audit shows what the engines recommend in your category today, which competitors hold position one, and where your gap is.

Frequently asked questions

How does ChatGPT decide which products to show?

Before it answers, ChatGPT classifies the intent behind your question. If it reads as a buying question, it tags the turn as a shopping use case, calls a browse tool, and returns a structured product carousel instead of plain prose. So whether you get a carousel at all is decided by that classification - not by the exact words you typed. Within the carousel, products are ordered on merit signals like reviews, rating, price and availability.

Where do ChatGPT's product recommendations come from?

Not from the model's memory. The products in a shopping carousel are pulled live from merchant feeds and product aggregators at answer time, each carrying a title, price, rating, review count, merchants and images. The model isn't recalling your brand from training - it's reading a current catalogue and selecting from it, which is why accurate, complete product data is what gets you included.

Why does ChatGPT show products for some questions but not others?

Because the shopping carousel only appears when the question is classified as commercial intent. 'Best running shoes for flat feet' triggers it; 'how do I lace running shoes' usually doesn't. Same topic, different intent, completely different answer format. If your buyers ask in ways that don't trip the shopping classifier, you won't appear on the product shelf no matter how good your data is.

Can you influence which products ChatGPT recommends?

You can't pay for placement - the carousel is organic and unsponsored, and OpenAI says the ads it now tests in ChatGPT are labeled, separate and don't influence answers. What you can influence is whether your data makes you selectable: complete and accurate product feeds, correct prices and availability, genuine reviews and ratings, and structured product markup the engine can read unambiguously. The teardown shows the model sees a compact, structured snapshot of each product, so the brands with the cleanest snapshot win the shelf.

Muhammet İLBAŞ
Written by
Muhammet İLBAŞ
Founder & engineer, building Zene in public
Share

Put it into practice

Free AI visibility checker →Pricing page checker →Schema markup generator →All free toolsCompare Zene vs alternatives

Keep reading

All articles →
Teardowns

Tools for Google AI Mode (2026): what tracks it, what wins it

"Tools for AI Mode" hides two different questions: how do I use Google's AI Mode well, and how do I know whether it mentions my brand? A quick answer to the first, and an honest map of the second - what Search Console shows (and hides), what SEO suites actually capture, what AI visibility trackers measure instead, and which of it is free. (We make a tool in the last category, and we say so.)

Teardowns

The best AI SEO tools in 2026 - sorted by the job you're hiring them for

"AI SEO tools" is one label hiding three different jobs: AI that helps you write and optimize content, AI features inside classic SEO suites, and tools that track your visibility inside AI answers. An honest map of all three - what each category actually does, who it's for, and how to avoid buying the wrong one. (We make a tool in the third category, and we say so.)

Teardowns

The ChatGPT your buyers use isn't the one the API shows you

Ask the same question through the ChatGPT API and through the app your customers actually open, and you can get two different answers - one from the model's frozen memory, one from a live web search with real citations. Most GEO tools only ever measure the first. Here's exactly why the two disagree, why the gap decides where AI sends customers, and what it really costs to capture the answer people see.

Find out if AI recommends you.

Put this guide into practice - get your free visibility score in minutes.

Zene dashboard: a brand's AI visibility score, the next automatic scan, and per-engine cards for ChatGPT, Claude, Gemini and Perplexity