Most robots.txt files written for AI crawlers are copied from a list somewhere, and many of them do something the owner did not intend. A SaaS company blocks every bot with “AI” in its description and disappears from ChatGPT search. A publisher blocks Google-Extended and expects to leave AI Overviews, which that token does not control. A retailer adds an allow rule for one search bot and accidentally opens its checkout and account pages to it. The fix is to start from the business goal, use the tokens each provider documents, and check the result in your server logs.
Why training and search crawlers are separate, and where each one feeds, is covered in the AI search guide. Everything below assumes that split and gets into the file itself.
The tokens, grouped by what blocking them does
These are the tokens I would consider in October 2026, checked against each operator's documentation. Group them by effect, because that is how you will write the file.
Naming these tokens has become more common, though it is still a minority practice. HTTP Archive's 2025 Web Almanac found GPTBot named in 4.5% of the desktop robots.txt files it analysed, up from 2.9% in 2024, and ClaudeBot in 3.6%, up from 1.9%.1 Those figures count mentions, so they include files that allow a bot as well as files that block it.
| Token | Operator | What it controls | Follows robots.txt? |
|---|---|---|---|
GPTBot | OpenAI | Crawling for model training. Disallowing it signals that content should not be used in training.2 | Yes |
OAI-SearchBot | OpenAI | Surfacing sites in ChatGPT search. Changes take about 24 hours.2 | Yes |
ChatGPT-User | OpenAI | User-initiated visits from ChatGPT and Custom GPTs.2 | OpenAI says robots.txt rules may not apply |
ClaudeBot, Claude-SearchBot, Claude-User | Anthropic | Training, search indexing and user-requested fetches respectively. Anthropic also supports Crawl-delay.3 | Yes, set per subdomain |
PerplexityBot | Perplexity | Surfacing and linking sites in Perplexity results. Not used for model training.4 | Yes |
Perplexity-User | Perplexity | User-requested visits to answer a question.4 | Generally ignores it |
Google-Extended | Training future Gemini models, and grounding in Gemini Apps and Grounding with Google Search on Vertex AI. No effect on Google Search inclusion or ranking.5 | Read from robots.txt; no crawler of its own | |
Applebot-Extended | Apple | Whether content Applebot crawls can train Apple's foundation models. Not used in Search ranking.6 | Read from robots.txt; does not crawl |
Meta-ExternalAgent | Meta | Crawling for uses such as training foundation AI models or improving products.7 | Yes. Meta-ExternalFetcher, for user requests, may bypass it. |
CCBot | Common Crawl | Crawling for Common Crawl's open repository of web data, which anyone can access.8 | Yes |
CCBot needs a judgement call. Common Crawl describes itself as an open repository for research, and its archive is open to anyone,8 so you cannot know every use made of it. If your goal is to stay out of training data generally, I would block it along with the named training crawlers.
Four parsing rules that catch people out
Most broken AI configurations come from how robots.txt is parsed. The standard is RFC 9309, and Google publishes its own reading of it, which helps where the RFC leaves room.
Configuration one: be cited, but not trained on
This is the setup I would recommend for most SaaS, ecommerce, B2B and local businesses. You want assistants to find, quote and link to your pages, and you would prefer future models not to train on them. The simplest version adds one group for the training tokens and leaves the search crawlers to follow your normal rules.
Cloudflare's traffic data shows why the split matters. Across a fixed set of its customers, it classed 80% of AI crawling in the 12 months to July 2025 as training, 18% as search and 2% as user actions, assigning purpose from operator disclosures. In July 2025 Anthropic's crawlers visited 38,000 pages for every page visit Anthropic referred, against 1,091.4 for OpenAI and 194.8 for Perplexity.12 Training is most of the crawling and is the part with no route back to your site, so blocking it while leaving the search crawlers alone keeps the traffic that can send visitors.
| Group | Lines | Why |
|---|---|---|
| Training crawlers | User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Meta-ExternalAgent User-agent: CCBot then Disallow: / | One group, several user agents, one rule. Each operator documents these as the training opt-out. |
| Everyone else | User-agent: * then your usual rules, for example Disallow: /checkout/ and Disallow: /account/ | OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot are not named, so they follow this group. |
Leaving the search crawlers unnamed is deliberate. Naming them in an Allow: / group looks tidy but removes your default rules for them. If you do want them named, copy every shared Disallow line into their group as well.
Know the cost of one line. Blocking Google-Extended also opts you out of grounding in Gemini Apps and in Grounding with Google Search on Vertex AI.5 A business that wants to be cited in the Gemini app may decide to leave that token out, accepting Gemini training as the price.
Configuration two: block AI use as far as robots.txt allows
A publisher with licensing deals, or a membership site, may want to keep content out of AI products altogether. Add the search and user-triggered tokens to the training group: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. Expect three gaps.
- User-triggered fetchers may ignore the file. OpenAI says robots.txt rules may not apply to ChatGPT-User,2 Perplexity says Perplexity-User generally ignores them,4 and Meta says Meta-ExternalFetcher may bypass them.7 Enforcing a block on those needs firewall rules matched to their published IP ranges.
- Google's AI Overviews and AI Mode are not controlled by any AI token. Blocking Googlebot would remove you from Google Search. The documented route is the Search generative AI control in Search Console, which excludes your links and content from AI Overviews, AI Mode and generative AI features in Discover, and does not affect other parts of Search.13
- robots.txt is not access control. RFC 9309 says its rules are not a form of access authorisation, and listing paths in it makes them discoverable.10 Anything truly private belongs behind a login.
Configuration three: allow everything
Some businesses want every model and assistant to know them: a developer tools company whose documentation should be in every coding assistant, or a destination marketing body that wants visitors from anywhere. Here the file needs no AI-specific lines at all. Keep your normal User-agent: * group and do not name the AI crawlers.
The work in this configuration happens at the edge. OpenAI says that to be eligible for ChatGPT search, the host or CDN must allow traffic from its published search bot IP addresses.14 Bot management that challenges unfamiliar user agents can block crawlers your robots.txt welcomes, so check the firewall rules as well as the file.
Checking it works in your server logs
A robots.txt that parses correctly is only half the job. These checks show whether crawlers are behaving as you intended.
I would put a short version of this check on the list for every CDN change, platform move and security review. A new bot management rule can undo a careful robots.txt overnight, and nobody notices until AI referrals fall.
Common questions
01Will blocking GPTBot remove my site from ChatGPT search?
02Does blocking Google-Extended keep my site out of AI Overviews?
03Why does a bot I allowed now crawl pages I blocked for everyone?
User-agent: * rules, so repeat any shared disallows in its group.04Can I stop ChatGPT-User or Perplexity-User with robots.txt?
05Should I block AI crawlers by IP address?
Sources
- Platform docs11
- Regulator1
- Industry study2
- 01SEO: 2025 Web AlmanacHTTP ArchiveIndustry study
- 02Overview of OpenAI CrawlersOpenAIPlatform docs
- 03Does Anthropic crawl data from the web, and how can site owners block the crawler?Claude Help CenterPlatform docs
- 04Perplexity CrawlersPerplexityPlatform docs
- 05Google's common crawlersGoogle Crawling InfrastructurePlatform docs
- 06About ApplebotApple SupportPlatform docs
- 07Meta Web CrawlersMeta for DevelopersPlatform docs
- 08CCBotCommon CrawlPlatform docs
- 09How Google interprets the robots.txt specificationGoogle Crawling InfrastructurePlatform docs
- 10RFC 9309: Robots Exclusion ProtocolIETF (September 2022)Regulator
- 11Publishers and Developers - FAQOpenAI Help CenterPlatform docs
- 12The crawl-to-click gap: Cloudflare data on AI bots, training, and referralsCloudflare Blog (August 2025)Industry study
- 13Search generative AI controlSearch Console HelpPlatform docs
- 14Searching the web with ChatGPTOpenAI Help CenterPlatform docs
Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the AI search guide.