Blog

How AI assistants choose which Shopify stores to recommend

When a shopper asks ChatGPT, Perplexity, Claude or Gemini what to buy, the answer cites a few stores. Here's what decides whether yours is one of them.

· 5 min read

A shopper types best trail runner for wide feet? into an AI assistant. A few seconds later the answer is on screen: three shoes, three stores, three source links. There's no page two to scroll to. If you sell a wide-fit trail runner and your store isn't one of the three, that shopper never hears about you.

Getting into those answers is what people now call generative engine optimization, or GEO. The name is new, but it rests on three questions an assistant has to answer "yes" to before it can recommend a product:

  1. Can it reach the page? Its crawler has to be allowed in.
  2. Can it read the product? The facts a shopper asks about have to be on the page as text.
  3. Can it tell what the store offers? The catalog, policies and product identifiers have to be clear enough to match to a question.

Most stores that never come up fail one of these, and usually in a way that can be fixed.

1. Let the right crawlers in

Assistants that answer with live sources send crawlers to fetch pages. The ones that matter for shopping answers include:

A second group collects text to train future models: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl) and Google-Extended, which is a control token rather than a crawler of its own. Blocking these doesn't take you out of today's answers. Blocking the first group does.

That difference is where most problems start. A "block AI bots" snippet pasted into robots.txt often lists search crawlers next to training crawlers, so a store that only meant to opt out of training quietly drops out of ChatGPT search and Perplexity too. A leftover Disallow: / under User-agent: * does the same to every crawler at once.

Shopify's default robots.txt doesn't block the search crawlers above, so if yours does, someone or something edited it. Open yourstore.com/robots.txt and look for Disallow: / under User-agent: * or under any of those names. Whether to let the training crawlers in is a separate decision, and it's yours to make.

2. Put the answer on the page, in words

Once a crawler has fetched a page, the model works from what it can read. Anything it can't read, it can't match to a question.

Go back to the wide-feet question. A product page answers it only if the text says the shoe comes in a wide fit. If "2E" lives in a size-chart image, or the description is a one-line tagline, the page doesn't answer the question as far as the assistant can tell, however well the shoe sells.

The usual gaps:

None of this is new SEO advice. What has changed is that an assistant reads a page far more literally than a shopper skimming it, so the gaps cost more.

3. Give assistants a map: llms.txt and agents.md

llms.txt is a proposed convention: a short Markdown file at the root of a site that says what the site is and links to its most useful pages. Shopify stores already serve one by default, with agents.md alongside it, and both can be customized from your theme.

A good one names what you sell, points to your main collections, and links your shipping and returns policies, which shoppers ask assistants about all the time. Keep personal details such as private emails and phone numbers out of it: the file is public and widely cached.

Be realistic about what it does. How much assistants rely on these files is still unclear, so treat them as a cheap extra rather than a strategy. They're worth ten minutes, not more attention than a blocked crawler or a thin description.

4. Measure what actually arrives

Assistant traffic shows up in your analytics if you know where to look. Visits arrive with referrers such as chatgpt.com, perplexity.ai, gemini.google.com and claude.ai, and ChatGPT adds utm_source=chatgpt.com to many of the links it shows. Watching those numbers tells you whether the fixes above are moving anything.

Measure over weeks, not days. The same question can bring up different sources from one day to the next, so a single answer is a sample, not a verdict.

A 15-minute check you can run today

  1. Open yourstore.com/robots.txt. Make sure nothing says Disallow: / for User-agent: * or for OAI-SearchBot, Claude-SearchBot or PerplexityBot.
  2. Open your five best sellers and read only the text. Does each one say who it's for, what it's made of, and the specs a shopper would ask about?
  3. Check that your variants have barcodes (GTINs) filled in.
  4. Open yourstore.com/llms.txt. Does it describe your store and link your main collections and policies?
  5. Ask an assistant the question your customer would ask, and note which stores it cites. Ask again next week.

Doing this across a whole catalog

That check works for five products. For three hundred, it's a weekend. That's the job we built Discoverly for.

Discoverly runs inside your Shopify admin. A read-only scan takes under ninety seconds and scores your store out of 100 across five areas, from whether AI crawlers can reach it to whether your products carry details assistants can recognize. Every deduction shows the signal behind it, and every fix arrives as a draft: an unblocked crawler, an llms.txt or agents.md file, a rewritten description. You read the full diff, you press publish, and the original is saved so any change rolls back in one click.

The score and the audit are free, and so are the robots.txt, llms.txt and agents.md fixes. AI-written product content uses credits on paid plans, and every install starts with a 14-day trial and 20 credits. The score measures readiness, not traffic: it shows what stands between your catalog and the answer, and it can't promise that an assistant will pick you.

Install Discoverly on the Shopify App Store

Bring your store into the answer.

Install on Shopify 14-day trial, 20 AI credits. The audit stays free.