Skip to content

Resources

Shopify's readiness scanner, free AI checkers and StoreKnows: what each one actually tests

Five tools will tell you whether your store is ready for AI shoppers, and they measure five different things. Shopify's scanner and Craftshift's checker test files, schema and endpoints. AI Catalog Score grades catalog fields. Verity Score logs what four models say. StoreKnows grades answers to fit, size, spec and compatibility questions against your own product data. Here is what each one checks, what each costs, and when the free one is enough.

· 12 min read · Arve Solland

Short answer: Shopify’s readiness scanner and Craftshift’s checker test whether files, schema and endpoints are present. AI Catalog Score grades how complete your catalog fields are. Verity Score asks buyer questions across four models and logs the replies. None of them grade an assistant’s answer to “does this fit”, “which size”, “what is the spec” or “will it work with mine” against your own product record, show the data behind each answer, let you try a fix, and re-run the same questions after publishing. That is what StoreKnows does, and only that. All five have a place; this page says which.

Why five tools give five different answers

“Is my store ready for AI shoppers” sounds like one question. The tools that answer it measure different layers of the same store, and a store can pass one layer and fail the next.

  • Presence. Does /llms.txt exist, is the Product JSON-LD complete, is guest checkout on, does robots.txt admit GPTBot? Yes/no facts about files and markup.
  • Completeness. Per product, is the title clear, the category set, the metafields filled? Grades on the data you maintain.
  • Visibility. When a shopper asks a generic question, which brands get named? What the assistants say, aggregated across many stores.
  • Accuracy. Asked about one of your products, does an assistant give the right variant, price, stock state, size or spec? A grade on one answer against your own data.

Presence and completeness are necessary. They are not accuracy. In a simulated check we ran on 4 September 2026 across 34 specialist storefronts (340 questions, one shopper model, OpenAI’s GPT-5.4 mini via the API, native storefront tools only), the assistant almost always found the product; 119 of the 186 answers that were not fully right had missed a single fact about it. Every one of those stores served the standard files. The breakdown by question type is in What AI shoppers get wrong about specialist products.

So the useful comparison is not “which tool scores my store highest” but “which layer does each tool test, and which layer is my problem”.

Shopify’s agentic readiness scanner

Shopify’s own page at shopify.com/agentic-readiness (read 12 September 2026) is short. It asks for a product URL and says: “We check your product page for the structured data that AI agents read to answer shoppers’ questions.” It lists no individual checks. The detail comes from two secondary sources that opened the tool and wrote down what it does.

The first is a community thread of 26 April 2026 by Rahul of FoundGPT, who reported that the scanner “runs 31 checks on any public storefront across five categories”: AI discoverability, product schema, transaction readiness, trust signals and operational maturity. It is free, needs no login, and takes about thirty seconds per URL. His caveat: it diagnoses and “does not fix anything”.

The second is Craftshift’s guide of 12 May 2026, which lists the checks per category. Agent discovery: robots.txt, llms.txt, sitemaps, the UCP discovery endpoint. Product schema: whether the JSON-LD carries name, image, description, offers, review, aggregateRating, brand, sku, shippingDetails and itemCondition. Transaction readiness: guest checkout, current pricing and inventory, whether sign-in is required. Trust: policy pages, About, contact, FAQ depth. Operational: shipping clarity, return windows, currency, locale sitemaps. The guide also records what the scanner does not test: catalog breadth (it samples pages), client-side-rendered schema, variant image and count problems, translations, and the competitive signals that decide which product an agent recommends. Its summary line is fair: a perfect 100 “measures technical readiness, not competitive readiness”.

Two things follow. This is a thorough scanner for the presence layer, and since it is Shopify’s own, its checklist is the closest thing to a statement of what Shopify’s channels want to find. And none of the 31 checks asks a question about a product and grades the reply. It confirms the Product schema has an offers block; it does not check that the price in that block is the one for the variant a shopper asked about.

Craftshift’s AI readiness checker

Craftshift, a Shopify Partner, publishes a free checker at craftshift.com/ai-readiness-checker (read 12 September 2026). It takes a store URL, scans the homepage HTML and robots.txt, and returns a score out of 100 with pass/fail per check and suggested fixes. The nine checks: JSON-LD schema, Speakable schema, meta descriptions, and whether robots.txt admits GPTBot, PerplexityBot, ClaudeBot, OAI-SearchBot and ChatGPT-User, plus the presence of llms.txt and llms-full.txt.

Narrower than Shopify’s scanner and quicker to read, it is the tool for finding out whether you have accidentally blocked a crawler, a common problem that no answer-grading tool will diagnose. It does not open product pages and does not test answers. Craftshift’s dated guide to the native llms.txt and agents.md files is worth reading on its own.

AI Catalog Score

AI Catalog Score (aicatalogscore.com, read 12 September 2026) does two distinct things, and it helps to keep them apart.

The catalog audit grades each product on eight dimensions: title clarity, description richness, alt text completeness, JSON-LD structured data, category assignment, metafields completeness, review signals and pricing structure, rolled up to a 0–100 score with a letter grade; the public audit samples roughly 250 products per store. This is the completeness layer at scale, and its emphasis on filled metafields over prose matches what we see in our own data.

The visibility tracking is what earns the company its dataset. Its “State of AI Commerce Q2 2026” report (20 May 2026) draws on 1,047,024 captures across 17,863 brands from six agents (ChatGPT, Claude, Perplexity, Gemini, Mistral and DeepSeek) over 22 days: Gemini produced 37% of captures, Amazon was the most-mentioned merchant at 64,426 mentions, the average audit score among the top 200 brands was 59/100, and position 50 in a category typically gets under a tenth of the mentions position 1 gets. That is original, large and worth citing. It measures which brands get named, not whether what is said about them is right; the report says it runs no controlled experiments.

Pricing, from its homepage on 12 September 2026 (there is no separate pricing page; /pricing returned 404): free up to 50 SKUs; Growth USD 49 a month for 500 SKUs; Pro USD 149 a month for 3,000; Performance USD 399 a month, unlimited, with an alternative of 5% of AI-attributed revenue uplift capped at USD 5,000 a month. Yearly billing is discounted 30%. It advertises a “+10 points within 30 days or full refund” promise on paid plans.

Verity Score

Verity Score (verityscore.io, read 12 September 2026) is the closest of the four to grading answers, so it is worth being precise about where it stops.

Its audit has three layers. The source layer checks what crawlers can read: rendering, schema.org, price visibility, availability, catalog metadata. The simulation layer runs “buyer decision questions” on ChatGPT, Perplexity, Gemini and Claude and keeps each reply “verbatim with its model, timestamp and sources”. The repair layer drafts fixes from existing product content and applies them after approval, reversible for 30 days.

Its published AI Buyer Score, explained in a post of 5 March 2026 (updated 12 April 2026), is a set of criteria a shopping agent needs to see on a product page: clear structured price, credible reviews, documented shipping, returns possible, confirmed availability, identifiable brand, complete specifications, coherent claims and proofs, crawlability. The post lists nine; the homepage now says eight. Either way, these are signal-presence criteria scored per page. The logged replies sit beside them as evidence of visibility, and Verity publishes original benchmarks from them: 237 audits across 102 Shopify stores, a GEO Barometer of 475 French stores, and fifteen vertical benchmarks.

What the published material does not describe is grading a reply to a fit, size, spec or compatibility question against the store’s own record and naming the fact that was wrong. The score is about what the page exposes; the simulations show what models say.

Pricing, from verityscore.io/en/pricing on 12 September 2026: Free at USD 0 (25 products, one ChatGPT purchase simulation, 60-day history); Lite USD 39 a month (100 products, four simulations on ChatGPT and Perplexity); Essential USD 149 a month (350 products, twelve simulations on three engines); Growth USD 449 a month (5,000 products, thirty simulations on all engines, weekly scans); Scale USD 899 a month (10,000 products); Enterprise on quote. Annual billing gives two months free.

Shopify’s Knowledge Base app

The Knowledge Base app (apps.shopify.com/shopify-knowledge-base, read 12 September 2026; by Shopify, free, launched 16 May 2025, 3.6 stars from 29 reviews) is not a scanner, but merchants reach for it when a scanner flags “thin policy content”. It lets you “view and customize the FAQs that AI shopping agents use to answer questions” about your store: shipping, returns, payment and policies generated from your settings, plus FAQs you write, stored as metaobjects, with a count of how often agents request the store’s information.

It is store-level. It holds no per-product spec, size chart or compatibility list, and it tests nothing. For policy questions, fully right only 36% of the time in our September batch, it is the right first step and it is free.

StoreKnows

StoreKnows tests the accuracy layer, on your own store, inside the Shopify admin. The check reads your catalog (products, variants, metafields, metaobjects, size charts), builds up to fourteen shopper questions with a recorded reference answer each (fit, size, spec filter, comparison, compatibility, policy), and has simulated shoppers on OpenAI, Google and Anthropic models ask them through the storefront tools every Liquid store already serves. A separate judge model grades each reply against the reference as fully right, partly right or wrong, and you see the reply, the model, the product data it should have found, and the gap.

Then three things the other tools do not do. You can type your own question into a proposed read-only tool and see the answer and the product fields it used, before anything is published. If you enable, the tools are added beside Shopify’s own through a theme app embed, and StoreKnows visits the storefront daily to confirm they are still available. And the same questions run again after publishing, so the before and after is per question, per model. On a 144-product development store on 5 September 2026 that took Google’s Gemini 3.8 Flash via the API from 7 to 13 of 14 fully right and OpenAI’s GPT-5.4 mini via the API from 4 to 8, one run each; runs vary.

Its limits: it does not check files, schema or robots.txt (run Shopify’s scanner for that); it does not track which brands get named for generic questions (Verity’s and AI Catalog Score’s work); a simulated check through developer APIs does not predict what a consumer assistant says to a real shopper; the published tools use a saved catalog copy, not live inventory, and work with compatible agents on theme-based stores.

Pricing: the check is free, one per store. Enabling is a one-time charge per store through Shopify billing, with no subscription, no usage charges and no per-order fee; the current figure is on our Shopify App Store listing, so this table gives the other tools’ published prices and not ours.

What each tool tests

Shopify scanner Craftshift checker AI Catalog Score Verity Score Knowledge Base StoreKnows
Files and robots.txt (llms.txt, crawler access) Yes Yes No Partial (crawlability) No No
Product JSON-LD completeness Yes Partial (homepage only) Yes Yes No No
Endpoint and checkout readiness (UCP, guest checkout) Yes No No No No No
Per-product field completeness (metafields, category, alt text) No No Yes Partial No No
Brand visibility across models, generic questions No No Yes Yes No No
Model replies logged verbatim No No No Yes No Yes
Answers graded against the store’s own product data No No No No No Yes
Fit, size, spec, compatibility questions on your products No No No Partial (specs as a presence check) No Yes
Product data behind each answer shown No No No No No Yes
Try a fix with your own question before publishing No No No Partial (fix previews) No Yes
Same questions re-run after publishing No No No No No Yes
Store-level policy FAQs editable No No No No Yes No
Runs on any public store without installing Yes Yes Yes No No No
Price Free Free Free to 50 SKUs, then USD 49–399/mo Free tier, then USD 39–899/mo Free Check free, then one-time (see listing)

“Partial” means the tool touches the row without doing what the row says in full; the sections above give the specifics. Prices as read on 12 September 2026.

When the free scanner is enough

Run Shopify’s scanner first, whatever else you do. It is free, it takes thirty seconds, and it is the checklist from the company that runs the channels. It is enough on its own when:

  • Your products have one or two options and one price each, so the first variant in the schema is the only variant.
  • The facts a shopper needs are in the description, not in a spec table, a size chart or a metafield.
  • Your question is “have I blocked a crawler, forgotten llms.txt, or left a schema field empty”, which is exactly what it and Craftshift’s checker find.
  • Nothing has changed and you want a monthly confirmation that the files are still there.

On a text-rich 150-product coffee store we checked on 2 September 2026, the native tools alone answered 10 of 10 questions for Anthropic’s Claude Opus 5 and 9 of 10 for OpenAI’s GPT-5.6, in a simulated check. A store like that learns little from an answer grader that a file check has not already told it.

When to choose the others

Verity Score when you want to see what four consumer-facing engines say about your products over time, replies kept verbatim and dated, and you sell in a market it benchmarks. Its free plan gives one ChatGPT simulation on 25 products, enough to see whether the format suits you.

AI Catalog Score when you have hundreds or thousands of SKUs and want a consistent completeness grade per product so a team can work through the worst ones, plus a picture of which brands in your category the agents name. The free 50-SKU tier is a real audit.

Knowledge Base when the wrong answers are about shipping, returns or payment rather than products. Write the FAQ there before you buy anything.

Craftshift’s checker when you want a one-minute robots.txt and llms.txt check without a login, or you have just changed a theme and want to confirm nothing blocked a bot.

StoreKnows when the scanner passes and shoppers still get the wrong variant, price, size or compatibility. That is the attribute problem, four of five classifiable gaps in our September batch, and it lives below the layer the file and field tools measure. If you are not sure which layer your problem is on, the fresh-session method in How to check what ChatGPT, Gemini and Copilot say about your store takes twenty minutes and will tell you.

Use them in order

Presence, then completeness, then accuracy. Shopify’s scanner and Craftshift’s checker in the first ten minutes; Knowledge Base for the policy FAQs the scanner flags; AI Catalog Score or Verity Score if your catalog is large or you want the visibility picture; a simulated check that grades answers when the facts that decide a sale sit in variants, charts and specs. StoreKnows runs that check on your own store for free, shows every answer beside the product data behind it, and lets you try a fix with your own question before anything is published.

Results in this article come from simulated checks run by StoreKnows on 2 and 4 September 2026 against public storefronts and on 5 September 2026 against a development store. Questions were answered by OpenAI’s GPT-5.4 mini and GPT-5.6, Google’s Gemini 3.8 Flash and Anthropic’s Claude Opus 5, called through their developer APIs, not the consumer apps, and graded by a separate judge model. Third-party check lists, datasets and prices are as published on the pages named, read on 12 September 2026. StoreKnows is independently developed and not affiliated with, endorsed by or sponsored by OpenAI, Google, Anthropic, Shopify, Craftshift, AI Catalog Score or Verity Score.

Questions this article answers

Does Shopify's free agentic readiness scanner test what AI assistants answer about my products?
No. Per the community write-up of 26 April 2026 and Craftshift's guide of 12 May 2026, it runs 31 checks in five categories: agent discovery files, product JSON-LD fields, checkout readiness, trust pages and operational signals. It tells you whether the structured data is present, not whether an assistant reads the right variant, price, size or spec from it.
Which AI readiness tool is free?
Shopify's scanner, Craftshift's checker and the Knowledge Base app are free with no tier. AI Catalog Score is free up to 50 SKUs, then USD 49, 149 or 399 a month. Verity Score has a free plan (25 products, one ChatGPT simulation), then USD 39 to 899 a month. StoreKnows' check is free and enabling is a one-time charge per store, shown on our Shopify App Store listing. Other tools' prices as read on 12 September 2026.
When is the free Shopify scanner enough?
When your products have few options, the deciding facts are in the description, and you mainly want to confirm that llms.txt, product schema and guest checkout are in place. Run it first. If your catalog has variants at different prices, size charts, spec tables or compatibility data, a file check cannot tell you whether an assistant reads those, and you need a tool that grades answers.