Baike.dev
Log in
> 返回资讯列表
news_article.exe
📰
#OpenAI#GPT#Google#Claude#Anthropic

133 of 10,099 Shopify stores block an AI crawler. Six block the one ChatGPT shops with.

2026年9月6日2 次浏览来源:Dev.to 阅读原文

Merchants are told two opposite things about AI crawlers: block them, because they take your content and give nothing back; and admit them, because that is how a store gets into an AI shopping answer. Both assume a decision is being made. We wanted to know how many stores have made one, and which way. Method Every store in a corpus of 10,099 known Shopify storefronts has its read as part of a scan. The file is parsed the way the major crawlers document parsing it: most specific group wins, is the fallback, longest matching path rule wins, beats on a tie. Each of twelve crawler names is asked one question: may it fetch ? A store counts as blocking a crawler when the answer is no. A second reading looks at the product page for a tag carrying or . Readings were taken between 29 August and 2...

Merchants are told two opposite things about AI crawlers: block them, because they take your content and give nothing back; and admit them, because that is how a store gets into an AI shopping answer. Both assume a decision is being made. We wanted to know how many stores have made one, and which way. Method Every store in a corpus of 10,099 known Shopify storefronts has its read as part of a scan. The file is parsed the way the major crawlers document parsing it: most specific group wins, is the fallback, longest matching path rule wins, beats on a tie. Each of twelve crawler names is asked one question: may it fetch ? A store counts as blocking a crawler when the answer is no. A second reading looks at the product page for a tag carrying or . Readings were taken between 29 August and 2 September 2026. Every store that blocked at least one crawler, or carried the tag, is one row in the CSV at the end. The rest of the corpus blocked nothing and is the denominator. What this cannot see. A robots.txt is a request. A store can also block a crawler at the edge, with a bot-management rule or a firewall, and that block is invisible here because the scanner is not the crawler being blocked. Every count below is a floor. Result 133 of 10,099 stores block at least one crawler: 1.32%. Crawler What it feeds Fetches at answer time Stores blocking CCBot Common Crawl no 81 GPTBot OpenAI training & retrieval no 77 Bytespider TikTok / Doubao no 72 Amazonbot Alexa+ / Rufus no 60 Google-Extended AI Overviews & AI Mode grounding no 58 ClaudeBot Claude retrieval & citations no 53 Applebot-Extended Apple Intelligence no 48 meta-externalagent Meta AI no 47 ChatGPT-User Live fetches during a ChatGPT chat yes 11 PerplexityBot Perplexity search & shopping yes 10 OAI-SearchBot ChatGPT search & shopping results yes 6 Perplexity-User Live fetches when a Perplexity user asks yes 3 "Fetches at answer time" marks the four crawlers that read a page, or index for a search result, at the moment a person is asking. The other eight crawl ahead of time, to train or ground a model. The split follows each operator's published description of the name. The twelve are the names merchants' block lists actually carry, not every answer-time agent that exists; Anthropic's and Google's live-fetch agents are not among them. How they block 130 of 133 block with . The whole site, not the product pages. 3 block alone. 54 block exactly one crawler. 28 block the same eight, verbatim: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, GPTBot, Google-Extended, meta-externalagent. 4 more block those eight and others. Identical lists do not arise independently; this one reads like a copied snippet, and it contains none of the four answer-time crawlers. 1 store blocks all twelve. Training or answering Of the 133 stores that block anything, 120 block only training and grounding crawlers. 13 block at least one crawler that fetches at answer time; 1 of those blocks only answer-time crawlers, 12 block both kinds. The clearest pair is OpenAI's. trains; is what ChatGPT's search and shopping results are built from. 77 stores block the first. 6 block the second, and every one of those 6 also blocks the first. The reverse, keep the shopping crawler out and let the training crawler in, happens on 0 stores. No store in the corpus has decided to stay out of AI shopping answers. The stores that are out of them by robots.txt are there because a copied training opt-out happened to include the name. The noai tag 8 stores carry a or meta robots tag on their product page. 6 of the 8 block nothing at all in robots.txt. By their names, six of the eight are musicians' merchandise stores, which again looks like one template rather than eight decisions. The tag is a separate mechanism and this study makes no claim about which crawlers honour it. What it does and does not mean Blocking AI crawlers is rare on Shopify. 98.7% of stores block none of the twelve by robots.txt. Where it happens it is mos

> 分享:
Baike.dev

baike.dev helps you discover great languages, frameworks, databases, DevOps and cloud-native tools.

Quick links

About

Contribute

Found a great developer tool? Share it with the community.

Submit a tool
© 2026 baike.dev Developer EncyclopediaUpdated daily · Discover great developer tools