A site can have a carefully written robots.txt and still be invisible to AI search, because the crawler's request is stopped before it reaches the server. Bot management at a CDN or web application firewall sits in front of the origin and decides, request by request, whether a crawler gets the page, a challenge or a 403. Since 2025 that layer has increasingly been set up to stop AI crawlers by default, and a rule meant for training crawlers can also catch search crawlers, user-triggered fetchers and, now and then, Googlebot. robots.txt records what you would like. The edge decides what happens.
Choosing the tokens for the file itself is covered in how to set up robots.txt for AI crawlers, and the wider mechanics of crawling and retrieval in the AI search guide. What follows assumes the file is right and checks whether the edge agrees with it.
How a crawler request gets past the edge, or does not
Two layers decide what a crawler receives, and they work differently. robots.txt is read voluntarily. Cloudflare's own documentation says compliance is voluntary and that the file does not prevent crawlers from accessing content at a technical level.1 The edge enforces whatever rules it holds, whether or not the crawler read robots.txt first, and whether or not it would have obeyed it.
One crawler request, from robots.txt to your server
robots.txt
A compliant crawler fetches /robots.txt and decides which paths to request. That fetch passes through the edge too, so the file itself can be blocked or challenged.
The last step explains why so many teams miss the problem. They check their server logs, see no errors for OAI-SearchBot, and conclude that all is well, when the requests were being stopped one hop earlier and never logged at all.
What Cloudflare blocks by default, and the switches that do it
Cloudflare is the clearest case because it documents its defaults in detail. In July 2025 it said it would block AI crawlers that access content without permission or compensation by default, and that every new domain would be asked whether to allow AI crawlers.2 In July 2026 it gave every plan, including Free, three AI behaviours to control separately: Search, for crawlers that index content to answer questions later; Agent, for automated activity on a person's behalf such as chat fetch bots and browser-use agents; and Training, for crawlers that take content to train or fine-tune models. Each can be blocked on all pages, blocked only on pages that display ads, or allowed.3
The same announcement set new defaults from 15 September 2026. New domains will have Training and Agent bots blocked on pages that display ads, with Search allowed, and multi-purpose crawlers that combine Search and Training will be affected by the block on Training.3 By Cloudflare's definition, user-triggered fetchers belong under Agent. So a publisher site that onboards after that date can stop assistants fetching an article when a reader asks about it, without anyone choosing to.
The defaults follow what Cloudflare measures on its own network. Across a fixed set of customers, it classed 80% of AI crawling in the 12 months to July 2025 as training, 18% as search and 2% as user actions, and in July 2025 Anthropic's crawlers visited 38,000 pages for every page visit Anthropic referred.4 They come from a company that sells AI crawler controls, so read them as one well-placed vendor's view. They do explain why training crawlers are the first thing an edge blocks, and why the search and user-triggered traffic that does send visitors can get caught alongside it.
| Feature | What it does | What to watch for |
|---|---|---|
| AI bot policies (Search, Agent, Training) | Blocks each behaviour on all pages, on pages with ads, or not at all.3 | Agent covers the fetches an assistant makes while answering a person, so blocking it has a visible cost. |
| Managed robots.txt | Prepends Cloudflare's rules to your existing file. The managed rules disallow crawlers including GPTBot, ClaudeBot, Google-Extended and CCBot, and add a content signal of search=yes, ai-train=no.1 | The file crawlers receive differs from the one in your repository. Fetch the live /robots.txt before you debug anything else. |
| AI Crawl Control, per-crawler actions | Allows or blocks each named crawler. A block creates or updates a WAF custom rule. On the Free plan crawlers are identified by user agent string, and paid plans can choose a 403 or 402 response.5 | The block lives as a WAF rule, so it can surprise whoever manages the firewall later. |
| Bot Fight Mode | Challenges traffic matching known bot patterns, and cannot be adjusted with WAF custom rules.6 | Its documentation does not say how it treats verified crawlers, so check its events before assuming they pass. |
Other providers make similar choices with different defaults. In AWS WAF's Bot Control rule group, the category rules block only unverified bots, with one exception: the CategoryAI rule blocks AI bots whether they are verified or not.7 AWS also notes that Bot Control verifies bots using the IP address of the request origin, so verified bots arriving through a proxy or load balancer need a rule placed before the group.7 Perplexity's crawler documentation gives example rules for both Cloudflare and AWS WAF that match the user agent and the published IP ranges together.8 Whatever your provider, find the equivalent of each row in the table above and write down how it is set.
Search, training and user-triggered crawlers need different rules
The operators split their crawlers by purpose, and the cost of a block at the edge follows the purpose. The robots.txt post lists every token. The table below covers what an edge block does to each kind.
| Kind | Examples | Cost of an edge block | Is robots.txt enough? |
|---|---|---|---|
| Training | GPTBot, ClaudeBot | No documented effect on search visibility. OpenAI describes GPTBot as crawling content that may be used to train its models.9 | For compliant crawlers, yes. Anthropic says blocking its IP addresses may not work reliably because it can stop it reading robots.txt.10 |
| Search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | You leave that assistant's search index. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from its published IP ranges.9 | Yes, and an edge block overrides an allow in the file. |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User | The assistant cannot read your page when a person asks it to, even if it already knows the URL. | Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User,9 and Perplexity says Perplexity-User generally ignores them.8 |
| Control token | Google-Extended | None possible. Google-Extended has no user agent of its own, and crawling uses existing Google user agents.11 | Yes. robots.txt is the only place it works. |
The last row has a practical consequence. An edge rule cannot target Google-Extended, so any rule written to stop "Google AI" by matching Google user agents or IP ranges is blocking Googlebot. That costs more than AI features, which themselves need crawling to be allowed in robots.txt and by any CDN or hosting infrastructure.12 The status code the edge returns matters too. A 403 means the server understood the request and refuses to authorise it,13 and Google does not index URLs that return a 4xx status code and removes indexed URLs that start returning one.14 A 429 signals too many requests in a given amount of time,15 and Googlebot treats it like a server error that slows crawling. Google's documentation also advises against using 401 or 403 to limit crawl rate.14
When Googlebot is caught at the edge, the cause is rarely the AI controls. I would look first at rate limits tuned for human traffic, country blocks, custom rules that match "bot" in the user agent, and verification that fails because the CDN sees a proxy's address instead of the crawler's.
How to check that a crawler is who it says it is
Anyone can send a request with GPTBot or Googlebot in the user agent. Cloudflare treats a bot as verified only when it identifies itself in a way that can be checked: a Web Bot Auth signature, a published IP list with a stable user agent, or reverse DNS.16 You can run the same checks on any IP address in your logs.
A request that claims to be OAI-SearchBot from an address outside OpenAI's list is not OAI-SearchBot. Blocking it is correct, and it should not count as evidence of what OpenAI's crawler receives.
Testing what each crawler receives with curl
The quickest test is to request a page while pretending to be each crawler, and compare the results with a normal browser request. In bash:
for ua in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User PerplexityBot Perplexity-User Googlebot; do printf '%-18s' $ua; curl -s -o /dev/null -w '%{http_code}\n' -A "Mozilla/5.0 (compatible; $ua)" https://www.example.com/; done- The same in PowerShell:
foreach ($ua in 'GPTBot','OAI-SearchBot','ClaudeBot','PerplexityBot','Googlebot') { try { $r = Invoke-WebRequest 'https://www.example.com/' -UserAgent "Mozilla/5.0 (compatible; $ua)" -UseBasicParsing; "$ua $($r.StatusCode)" } catch { "$ua $($_.Exception.Response.StatusCode.value__)" } } - Fetch the live file as the crawlers see it:
curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot)" https://www.example.com/robots.txt. If a managed block appears above your own rules, the CDN is adding it. - Run each test against a product or article page as well as the home page, because rules for ads or particular paths only show up there.
Read the results with care, because this test only exercises rules that look at the user agent. On Cloudflare's Free plan, AI Crawl Control identifies crawlers by user agent string,5 so a spoofed request will usually meet the same block as the real crawler. Rules that verify identity behave the other way: your laptop is not on OpenAI's list, so a well-configured edge should treat your fake OAI-SearchBot as unverified.
| Result | What it tells you | Next step |
|---|---|---|
| 403 or 402 for one crawler, 200 for a browser | A rule matches that user agent. The real crawler will almost certainly get the same. | Find the rule in the CDN's security events and decide whether you meant it. |
| A challenge page, whatever the status code | A bot mode is challenging the request. Crawlers do not solve challenges. | Check whether the real crawler is challenged too, in the events log. |
| 429 | A rate limit. Google treats 429 like a server error and slows crawling.14 | Exempt verified crawlers from limits tuned for people. |
| 200 for everything | No user agent rule blocks these crawlers. It proves nothing about rules based on verified identity or IP. | Confirm with logs that the real crawlers arrive and get 200s. |
Reading the evidence at the edge and at the origin
Your server's access log shows what got through. To count status codes by crawler in a log in the common combined format, where the status is the ninth field: for b in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot PerplexityBot Googlebot; do echo "$b"; grep "$b" access.log | awk '{print $9}' | sort | uniq -c; done. A crawler you allow that is missing entirely, or that only ever reaches /robots.txt, is a sign of a block upstream. Verify a sample of its IP addresses before you trust the counts.
Blocked requests are only in the CDN's records. On Cloudflare, requests challenged by Bot Fight Mode appear under Security, Analytics, in the Events tab, labelled Bot Fight Mode in the Service field.6 AI Crawl Control lists each AI crawler with its requests and robots.txt violations, and notes that unsuccessful requests can come from any rule or error, not only its own blocks.5 Read both views together: the origin log for what arrived, the edge log for what was turned away and by which rule.
The same check belongs on the list for every CDN change, security review and replatform. The website migrations guide makes the same point for Googlebot when a site changes host. The Cite skill in my open-source pack includes AI crawler access in its checks, and how to measure visibility in AI search covers spotting the referral drop that is often the first symptom.
Common questions
01Why is OAI-SearchBot blocked when my robots.txt allows it?
02Does Cloudflare block Googlebot by default?
03Can I block Google-Extended at my CDN?
04Is a curl request with a crawler's user agent a reliable test?
05Should I block ChatGPT-User and other user-triggered fetchers at the edge?
Sources
- Platform docs14
- Regulator2
- Industry study1
- 01robots.txt settingCloudflare Docs (updated August 2026)Platform docs
- 02Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for A New Business ModelCloudflare press release (July 2025)Platform docs
- 03New options to manage AI trafficCloudflare Changelog (July 2026)Platform docs
- 04The crawl-to-click gap: Cloudflare data on AI bots, training, and referralsCloudflare Blog (August 2025)Industry study
- 05Manage AI crawlersCloudflare AI Crawl Control docs (updated July 2026)Platform docs
- 06Get started with Bot Fight ModeCloudflare Docs (updated August 2026)Platform docs
- 07AWS WAF Bot Control rule groupAWS WAF Developer GuidePlatform docs
- 08Perplexity CrawlersPerplexityPlatform docs
- 09Overview of OpenAI CrawlersOpenAIPlatform docs
- 10Does Anthropic crawl data from the web, and how can site owners block the crawler?Claude Help CenterPlatform docs
- 11Google's common crawlersGoogle Crawling InfrastructurePlatform docs
- 12AI features and your websiteGoogle Search CentralPlatform docs
- 13RFC 9110: HTTP Semantics, section 15.5.4 (403 Forbidden)IETF (June 2022)Regulator
- 14How HTTP Status Codes Affect Google's CrawlersGoogle Crawling InfrastructurePlatform docs
- 15RFC 6585: Additional HTTP Status Codes, section 4 (429 Too Many Requests)IETF (April 2012)Regulator
- 16Verified botsCloudflare Docs (updated July 2026)Platform docs
- 17Verify requests from Google crawlers and fetchersGoogle Crawling InfrastructurePlatform docs
Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the AI search guide.