robots.txt for AI crawlers: which bots to allow on an online store
Which AI user agents to allow or block in robots.txt (ChatGPT-User, OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot and more), the difference between shopping agents, AI search and training crawlers, and a ready-to-use robots.txt.
Short answer: allow AI assistant and AI search user agents, because they fetch your store when a shopper asks about it, and decide separately about AI training crawlers. For most online stores that means allowing ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot, Perplexity-User and PerplexityBot, and blocking training crawlers such as GPTBot, ClaudeBot, Google-Extended or CCBot only if you want to opt out of training.
You can check what your current robots.txt blocks with the free agent-readiness audit.
Three kinds of AI crawler
Not all AI bots do the same job, so blocking them has very different costs.
- Assistant agents fetch a page live because a person asked the assistant something, for example “is this jacket in stock in medium?”. Block these and the assistant can’t answer about your store.
- AI search crawlers build the index that AI answers and shopping results are drawn from. Block these and your products are less likely to be recommended at all.
- Training crawlers collect text to train future models. Blocking them is a legitimate choice and has little effect on whether shoppers find you today.
The user agents to know
| User agent | Operator | Kind |
|---|---|---|
ChatGPT-User |
OpenAI | Assistant agent |
OAI-SearchBot |
OpenAI | AI search |
GPTBot |
OpenAI | Training |
Claude-User |
Anthropic | Assistant agent |
Claude-SearchBot |
Anthropic | AI search |
ClaudeBot |
Anthropic | Training |
Perplexity-User |
Perplexity | Assistant agent |
PerplexityBot |
Perplexity | AI search |
Google-Extended |
Training (Gemini) | |
Applebot-Extended |
Apple | Training |
Amazonbot |
Amazon | AI search |
meta-externalfetcher |
Meta | Assistant agent |
meta-externalagent |
Meta | Training |
MistralAI-User |
Mistral | Assistant agent |
DuckAssistBot |
DuckDuckGo | AI search |
CCBot |
Common Crawl | Training |
Bytespider |
ByteDance | Training |
Google-Extended and Applebot-Extended are control tokens rather than separate crawlers: blocking them opts your content out of AI training without affecting normal Google or Apple search.
A robots.txt for a store that wants AI shoppers
This allows assistants and AI search, opts out of training, and keeps cart, checkout and account pages out of every crawler.
# Assistants and AI search: let them read the store
User-agent: ChatGPT-User
User-agent: OAI-SearchBot
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: Perplexity-User
User-agent: PerplexityBot
Disallow: /cart
Disallow: /checkout
Disallow: /my-account
# Training crawlers: opt out (optional, delete this group to allow training)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /
# Everyone else
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /my-account
Sitemap: https://yourstore.example/sitemap.xml
Change the paths to match your platform. On Magento these are usually /checkout/cart, /checkout and /customer/account.
Mistakes that block AI agents by accident
- A blanket
User-agent: *withDisallow: /left over from a staging site. It blocks every agent that doesn’t have its own group. - Copying a “block all AI” list that includes
ChatGPT-UserorPerplexityBotwhen you only meant to stop training. - A firewall or bot-protection rule that returns
403,429or503to anything that isn’t a browser. robots.txt can say “allow” while the firewall says “no”. Check your CDN or security plugin and allow verified AI agents. - robots.txt returning a 5xx error. Many crawlers treat a server error on robots.txt as “keep out”. Serve it with a
200, or a404if you have none. - No
Sitemap:line. It’s the quickest way for any crawler to find your product sitemap.
Frequently asked questions
Does blocking GPTBot remove my store from ChatGPT?
No. GPTBot is OpenAI’s training crawler. ChatGPT’s search results use OAI-SearchBot, and live browsing on a user’s request uses ChatGPT-User. Keep those allowed if you want ChatGPT to recommend your products.
Do AI crawlers obey robots.txt?
The major operators listed above say their crawlers do. Assistant agents acting on a user’s direct request may be treated differently by some operators, which is another reason not to rely on robots.txt alone to control access to sensitive pages.
How do I test my robots.txt against AI crawlers?
Run the free agent-readiness audit. It checks your robots.txt against every user agent in the table above, flags which kind each blocked one is, and also detects firewalls that turn bots away.