Skip to content

Resources

How to check what ChatGPT, Gemini and Copilot say about your store (manually, and with a simulated check)

A single screenshot of one assistant answer tells you almost nothing. Here is the fresh-session method merchants use to check what ChatGPT, Gemini, Copilot and Perplexity say about their products, the buyer-intent questions worth asking, why the answers differ between sessions and apps, and what a repeatable simulated check adds on top.

· 9 min read · Arve Solland

Short answer: open ChatGPT with no history, ask the questions a buyer would ask about one specific product, and write down the variant, price and stock state it describes and the date. Repeat in Gemini, Copilot and Perplexity, and again next week. Then run the same questions as a simulated check against your store’s own tools, so you can see which fact the assistant lacked and re-test after you fix it.

Why one screenshot is not a check

On 9 June 2026 a merchant opened a thread in the Shopify Community titled “Is your store actually showing up in ChatGPT shopping results? Here’s how I checked mine”. The method was simple: ask “the kind of question a real buyer would” and see which stores come back. Over the following weeks the replies added the caveats that matter. One person logged a single product across four assistants and found it in three of them. Another pointed out that the same query asked twice can return two different answers, and that traffic from these answers often lands in analytics as direct, so the order that came from an assistant looks like any other.

Two lessons follow. First, an assistant answer is a sample, not a fact about your store; you need several samples before you can say anything. Second, there are two different things you can check. One is whether your store appears at all when a shopper asks a generic question (“best rain jacket for cycling”). The other is whether the assistant describes your product correctly once it has found it: the right variant, the right price, in stock or not, will it fit. This article is mostly about the second, because that is the part you control from your product data, and because it is where the errors sit. In a simulated check we ran on 4 September 2026 across 34 specialist storefronts (340 questions, one shopper model, OpenAI’s GPT-5.4 mini via the API), the assistant almost always found the product; 119 of the 186 answers that were not fully right had missed one fact about it. The full breakdown by question type is in What AI shoppers get wrong about specialist products.

How do you check by hand?

The fresh-session method takes twenty minutes the first time and ten each week after.

  1. Start clean. Use a chat that does not carry your history: a temporary chat, a signed-out window, or a private browser window, whichever the app offers. Your own account has months of context about your store, and the assistant will use it. A shopper’s does not.
  2. Pick three products. One bestseller, one with several variants at different prices, and one whose deciding detail sits in a spec table or a size chart. The last two are where answers go wrong.
  3. Ask as a buyer, in two forms. Once without naming your store (“which 35 mm black-and-white film under $10 is in stock?”) and once naming it (“is the 5-roll bundle in stock at [store]?”). The first tells you about discovery, the second about accuracy.
  4. Write down what it described. Not “it got it right” but the variant it named, the price it quoted, the stock state it gave, and any spec or compatibility claim. Add the app, the date and whether the answer linked to your product page.
  5. Repeat in the other apps. ChatGPT, Gemini, Microsoft Copilot and Perplexity retrieve differently (more below), so a product that reads fine in one can be described wrongly in another. Four apps, three products, two forms each: 24 rows in a spreadsheet.
  6. Ask again in a week. The answers will move. What you are looking for is the fact that is wrong every time, across apps. That one is yours to fix.

Keep the log. When you change a product’s data, the log is the only way to tell whether the change reached the assistants.

Which questions should you ask?

Generic “best X” questions are easy to ask and hard to learn from; the answer depends on the whole market and on the assistant’s mood. Questions about your own products are the ones that show a fixable gap. Use these templates, with the brackets filled from a real product.

Type Template What to record
Fit “Does the [product] come in [colour or option], and is that one available now?” The option it names; available or not
Size “My [measurement] is [value]. Which size of [product] should I order?” The size it recommends; whether it cites your chart
Compatibility “I have [thing they own]. Will [product] work with it?” Yes, no or unsure; what it based that on
Comparison “What is actually different between [product A] and [product B]?” The differences it lists; the ones it misses
Stock and price filter “Show me in-stock [category] under [price] with [property].” Which products it returns; any that are sold out or over budget
Policy “If [product] does not fit, can I return it, and who pays shipping?” The window and the cost it states

The types are not equally hard. In our September batch, fit questions were fully right 81% of the time; comparisons 24%, filtered searches 28%. If you only have time for two questions per product, ask a comparison and a filter.

Why do answers differ between sessions and apps?

Four reasons, and all four are ordinary.

The model samples. Ask the same question twice and you get different wording, sometimes a different product, because the answer is generated fresh each time. This is why single results mean little.

Each app retrieves differently. An assistant can know your product from a syndicated catalog entry, from a crawl it made days ago, or from fetching the page while you wait. Shopify’s Agentic Storefronts, on by default for eligible stores and managed under Sales channels > Agentic, syndicates each product’s title, description, options, images, price, availability “and other key attributes” to ChatGPT, Google AI Mode and Gemini, Microsoft Copilot and Meta (help.shopify.com, read 12 September 2026). Perplexity is not in that list and reads your pages the way a search engine does. Two apps looking at two copies of your product, taken at two different times, will not agree.

The copy may be stale. A syndicated entry or a cached crawl lags your admin. If you changed a price or a variant sold out this morning, an assistant can carry the old value for days. Shopify’s own note that changes to Catalog access take up to seven days to apply gives a sense of the scale.

Context leaks in. Your history, your location and the session’s earlier turns all shape the answer. The ChatGPT channel, for instance, applies to stores that sell to customers in the United States; a shopper elsewhere may be served from a crawl instead of the catalog. Hence the fresh session.

None of this is a fault in the assistants or in the catalog. It is the reason a check has to be repeated and logged rather than done once.

What does “discovery-only” mean for the ChatGPT channel?

Shopify’s help page for the channel describes it as “a discovery-focused referrer platform”, and says ChatGPT users “complete their purchase on your online store checkout” (help.shopify.com, read 12 September 2026). Your store must sell to customers in the United States, though it can be based anywhere, and there are no fees beyond your usual payment processing.

For a check, this matters in two ways. First, the answer a shopper sees in ChatGPT is a referral: it describes your product, then sends the shopper to your product page and your checkout. If the description said $69.99 and the variant they wanted is $74.99, they find out on your page, not in the chat, and the mismatch is yours to absorb. So the accuracy of the description is worth checking even though the sale happens on your storefront. Second, the same page says that if you turn the channel off, your products “might still be displayed or referenced” through web crawling and indexing. Opting out removes the catalog copy, not the crawl, so a manual check is still worth running on a store that has opted out.

What does a simulated check add?

A manual check tells you what an app said today. It cannot tell you why, and it cannot be re-run exactly after a fix. A simulated check is built for those two things.

  • Same questions, generated from your data. The questions are the ones in the table above, drawn from your own products, each with a reference answer built from the store’s public product data before the assistant is asked.
  • Tools only. The simulated shopper uses the storefront tools every Liquid store has served since 21 August 2026 (search, browse, get product, show variant, cart, policies) and your public pages. Nothing else. So an answer reflects what your store makes readable, not what a model remembered from a crawl.
  • Graded against the store. A separate judge model marks each answer fully right, partly right or wrong against the reference, and records which fact was missing. That is how you get from “it was wrong” to “it quoted the first variant’s price instead of the one the shopper asked for”, which is the most common shape we see and the subject of An AI assistant says your product is out of stock. It isn’t.
  • Repeatable. Every tool call and every answer is stored. Change the product, run again, compare. On a 144-product dev store on 5 September 2026, the same 14 questions went from 7 to 13 fully right with Google’s Gemini 3.8 Flash via the API and from 4 to 8 with OpenAI’s GPT-5.4 mini via the API, one run each, after the deciding specs were made readable to the tools. Runs vary; treat that as a direction, not a score.

The manual log and the simulated check answer different questions. Keep both.

What a simulated check cannot tell you

It runs models through their developer APIs, not the consumer apps. So it does not tell you what ChatGPT, Gemini, Copilot or Perplexity will say to a particular shopper this afternoon, whether your store appears for a generic “best X” question, or how each app’s own retrieval and caching will treat your product. One model, one run, ten or fourteen questions: the percentages describe a shape, not a benchmark. And a right answer in a simulated check is a necessary condition, not a sufficient one; the fresh-session log is still how you confirm that the fix reached the apps.

A routine that fits in a week

Monday, ten minutes: the fresh-session log, four apps, the same three products. Once a month, a simulated check on the whole catalog to find the products whose one deciding fact is unreadable. After any change to variants, prices, size charts or compatibility data, both. StoreKnows runs the simulated check on your own store for free, shows every answer beside the product data behind it, and lets you try a fix with your own question before anything is published.

Results in this article come from simulated checks run by StoreKnows on 4 September 2026 against 34 public storefronts and on 5 September 2026 against a development store. Questions were answered by OpenAI’s GPT-5.4 mini and Google’s Gemini 3.8 Flash, called through their developer APIs, not the consumer apps, and graded by a separate judge model. StoreKnows is independently developed and not affiliated with, endorsed by or sponsored by OpenAI, Google, Anthropic or Shopify.

Questions this article answers

How do I check what ChatGPT says about my store?
Open ChatGPT in a session with no history (a temporary chat, or signed out), ask the questions a buyer would ask about a specific product, and write down the variant, price and stock state it describes, plus the date. Repeat the same questions in Gemini, Copilot and Perplexity, and again a week later. One answer is an anecdote; a log is a check.
Why does ChatGPT give a different answer to the same question in a new session?
Assistants sample their wording, retrieve differently each time (a syndicated catalog entry, a cached crawl, a live page fetch), and may use your chat history and location. The same question can return a different variant or price twice in a row. That is why you record several runs and look for the pattern, not the single result.
What does a simulated check add that asking the apps by hand does not?
It asks the same buyer questions through the models' developer APIs using only your store's own storefront tools, grades each answer against your product data, and records every tool call, so you can see which fact was missing and re-run it after a fix. It does not tell you what the consumer apps will say to a particular shopper.