Skip to content
← All writing
AI search10 September 2026Updated 7 October 202611 min read17 sources

When your CDN blocks AI crawlers, whatever robots.txt says

How bot management at the edge decides what AI crawlers and Googlebot receive, how to check a crawler is real, and how to test what each one gets.

By Mani Bharij, SEO & AI Search Consultant

In this postHover or tap a point
010203040506

Chapter 011 min

How a crawler request gets past the edge, or does not

A site can have a carefully written robots.txt and still be invisible to AI search, because the crawler's request is stopped before it reaches the server. Bot management at a CDN or web application firewall sits in front of the origin and decides, request by request, whether a crawler gets the page, a challenge or a 403. Since 2025 that layer has increasingly been set up to stop AI crawlers by default, and a rule meant for training crawlers can also catch search crawlers, user-triggered fetchers and, now and then, Googlebot. robots.txt records what you would like. The edge decides what happens.

Choosing the tokens for the file itself is covered in how to set up robots.txt for AI crawlers, and the wider mechanics of crawling and retrieval in the AI search guide. What follows assumes the file is right and checks whether the edge agrees with it.

Chapter 01 / 061 min

How a crawler request gets past the edge, or does not

Two layers decide what a crawler receives, and they work differently. robots.txt is read voluntarily. Cloudflare's own documentation says compliance is voluntary and that the file does not prevent crawlers from accessing content at a technical level.1 The edge enforces whatever rules it holds, whether or not the crawler read robots.txt first, and whether or not it would have obeyed it.

Diagram

One crawler request, from robots.txt to your server

01

robots.txt

A compliant crawler fetches /robots.txt and decides which paths to request. That fetch passes through the edge too, so the file itself can be blocked or challenged.

Step 1 of 5
Simplified. Each CDN evaluates its features in its own order. The point is that identity checks, bot policies and firewall rules all run before the origin sees the request.

The last step explains why so many teams miss the problem. They check their server logs, see no errors for OAI-SearchBot, and conclude that all is well, when the requests were being stopped one hop earlier and never logged at all.

Chapter 02 / 063 min

What Cloudflare blocks by default, and the switches that do it

Cloudflare is the clearest case because it documents its defaults in detail. In July 2025 it said it would block AI crawlers that access content without permission or compensation by default, and that every new domain would be asked whether to allow AI crawlers.2 In July 2026 it gave every plan, including Free, three AI behaviours to control separately: Search, for crawlers that index content to answer questions later; Agent, for automated activity on a person's behalf such as chat fetch bots and browser-use agents; and Training, for crawlers that take content to train or fine-tune models. Each can be blocked on all pages, blocked only on pages that display ads, or allowed.3

The same announcement set new defaults from 15 September 2026. New domains will have Training and Agent bots blocked on pages that display ads, with Search allowed, and multi-purpose crawlers that combine Search and Training will be affected by the block on Training.3 By Cloudflare's definition, user-triggered fetchers belong under Agent. So a publisher site that onboards after that date can stop assistants fetching an article when a reader asks about it, without anyone choosing to.

The defaults follow what Cloudflare measures on its own network. Across a fixed set of customers, it classed 80% of AI crawling in the 12 months to July 2025 as training, 18% as search and 2% as user actions, and in July 2025 Anthropic's crawlers visited 38,000 pages for every page visit Anthropic referred.4 They come from a company that sells AI crawler controls, so read them as one well-placed vendor's view. They do explain why training crawlers are the first thing an edge blocks, and why the search and user-triggered traffic that does send visitors can get caught alongside it.

FeatureWhat it doesWhat to watch for
AI bot policies (Search, Agent, Training)Blocks each behaviour on all pages, on pages with ads, or not at all.3Agent covers the fetches an assistant makes while answering a person, so blocking it has a visible cost.
Managed robots.txtPrepends Cloudflare's rules to your existing file. The managed rules disallow crawlers including GPTBot, ClaudeBot, Google-Extended and CCBot, and add a content signal of search=yes, ai-train=no.1The file crawlers receive differs from the one in your repository. Fetch the live /robots.txt before you debug anything else.
AI Crawl Control, per-crawler actionsAllows or blocks each named crawler. A block creates or updates a WAF custom rule. On the Free plan crawlers are identified by user agent string, and paid plans can choose a 403 or 402 response.5The block lives as a WAF rule, so it can surprise whoever manages the firewall later.
Bot Fight ModeChallenges traffic matching known bot patterns, and cannot be adjusted with WAF custom rules.6Its documentation does not say how it treats verified crawlers, so check its events before assuming they pass.
Cloudflare features that change what AI crawlers receive, as documented in summer 2026

Other providers make similar choices with different defaults. In AWS WAF's Bot Control rule group, the category rules block only unverified bots, with one exception: the CategoryAI rule blocks AI bots whether they are verified or not.7 AWS also notes that Bot Control verifies bots using the IP address of the request origin, so verified bots arriving through a proxy or load balancer need a rule placed before the group.7 Perplexity's crawler documentation gives example rules for both Cloudflare and AWS WAF that match the user agent and the published IP ranges together.8 Whatever your provider, find the equivalent of each row in the table above and write down how it is set.

Chapter 03 / 062 min

Search, training and user-triggered crawlers need different rules

The operators split their crawlers by purpose, and the cost of a block at the edge follows the purpose. The robots.txt post lists every token. The table below covers what an edge block does to each kind.

KindExamplesCost of an edge blockIs robots.txt enough?
TrainingGPTBot, ClaudeBotNo documented effect on search visibility. OpenAI describes GPTBot as crawling content that may be used to train its models.9For compliant crawlers, yes. Anthropic says blocking its IP addresses may not work reliably because it can stop it reading robots.txt.10
SearchOAI-SearchBot, Claude-SearchBot, PerplexityBotYou leave that assistant's search index. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from its published IP ranges.9Yes, and an edge block overrides an allow in the file.
User-triggeredChatGPT-User, Claude-User, Perplexity-UserThe assistant cannot read your page when a person asks it to, even if it already knows the URL.Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User,9 and Perplexity says Perplexity-User generally ignores them.8
Control tokenGoogle-ExtendedNone possible. Google-Extended has no user agent of its own, and crawling uses existing Google user agents.11Yes. robots.txt is the only place it works.
What blocking each kind of crawler at the edge costs you

The last row has a practical consequence. An edge rule cannot target Google-Extended, so any rule written to stop "Google AI" by matching Google user agents or IP ranges is blocking Googlebot. That costs more than AI features, which themselves need crawling to be allowed in robots.txt and by any CDN or hosting infrastructure.12 The status code the edge returns matters too. A 403 means the server understood the request and refuses to authorise it,13 and Google does not index URLs that return a 4xx status code and removes indexed URLs that start returning one.14 A 429 signals too many requests in a given amount of time,15 and Googlebot treats it like a server error that slows crawling. Google's documentation also advises against using 401 or 403 to limit crawl rate.14

When Googlebot is caught at the edge, the cause is rarely the AI controls. I would look first at rate limits tuned for human traffic, country blocks, custom rules that match "bot" in the user agent, and verification that fails because the CDN sees a proxy's address instead of the crawler's.

Chapter 04 / 061 min

How to check that a crawler is who it says it is

Anyone can send a request with GPTBot or Googlebot in the user agent. Cloudflare treats a bot as verified only when it identifies itself in a way that can be checked: a Web Bot Auth signature, a published IP list with a stable user agent, or reverse DNS.16 You can run the same checks on any IP address in your logs.

Checklist0 of 5 done

A request that claims to be OAI-SearchBot from an address outside OpenAI's list is not OAI-SearchBot. Blocking it is correct, and it should not count as evidence of what OpenAI's crawler receives.

Chapter 05 / 061 min

Testing what each crawler receives with curl

The quickest test is to request a page while pretending to be each crawler, and compare the results with a normal browser request. In bash:

  • for ua in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User PerplexityBot Perplexity-User Googlebot; do printf '%-18s' $ua; curl -s -o /dev/null -w '%{http_code}\n' -A "Mozilla/5.0 (compatible; $ua)" https://www.example.com/; done
  • The same in PowerShell: foreach ($ua in 'GPTBot','OAI-SearchBot','ClaudeBot','PerplexityBot','Googlebot') { try { $r = Invoke-WebRequest 'https://www.example.com/' -UserAgent "Mozilla/5.0 (compatible; $ua)" -UseBasicParsing; "$ua $($r.StatusCode)" } catch { "$ua $($_.Exception.Response.StatusCode.value__)" } }
  • Fetch the live file as the crawlers see it: curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot)" https://www.example.com/robots.txt. If a managed block appears above your own rules, the CDN is adding it.
  • Run each test against a product or article page as well as the home page, because rules for ads or particular paths only show up there.

Read the results with care, because this test only exercises rules that look at the user agent. On Cloudflare's Free plan, AI Crawl Control identifies crawlers by user agent string,5 so a spoofed request will usually meet the same block as the real crawler. Rules that verify identity behave the other way: your laptop is not on OpenAI's list, so a well-configured edge should treat your fake OAI-SearchBot as unverified.

ResultWhat it tells youNext step
403 or 402 for one crawler, 200 for a browserA rule matches that user agent. The real crawler will almost certainly get the same.Find the rule in the CDN's security events and decide whether you meant it.
A challenge page, whatever the status codeA bot mode is challenging the request. Crawlers do not solve challenges.Check whether the real crawler is challenged too, in the events log.
429A rate limit. Google treats 429 like a server error and slows crawling.14Exempt verified crawlers from limits tuned for people.
200 for everythingNo user agent rule blocks these crawlers. It proves nothing about rules based on verified identity or IP.Confirm with logs that the real crawlers arrive and get 200s.
How to read a spoofed curl test
Chapter 06 / 062 min

Reading the evidence at the edge and at the origin

Your server's access log shows what got through. To count status codes by crawler in a log in the common combined format, where the status is the ninth field: for b in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot PerplexityBot Googlebot; do echo "$b"; grep "$b" access.log | awk '{print $9}' | sort | uniq -c; done. A crawler you allow that is missing entirely, or that only ever reaches /robots.txt, is a sign of a block upstream. Verify a sample of its IP addresses before you trust the counts.

Blocked requests are only in the CDN's records. On Cloudflare, requests challenged by Bot Fight Mode appear under Security, Analytics, in the Events tab, labelled Bot Fight Mode in the Service field.6 AI Crawl Control lists each AI crawler with its requests and robots.txt violations, and notes that unsuccessful requests can come from any rule or error, not only its own blocks.5 Read both views together: the origin log for what arrived, the edge log for what was turned away and by which rule.

The same check belongs on the list for every CDN change, security review and replatform. The website migrations guide makes the same point for Googlebot when a site changes host. The Cite skill in my open-source pack includes AI crawler access in its checks, and how to measure visibility in AI search covers spotting the referral drop that is often the first symptom.

Questions5 answered

Common questions

01Why is OAI-SearchBot blocked when my robots.txt allows it?
Because a CDN or firewall rule is stopping it before robots.txt matters. OpenAI recommends allowing OAI-SearchBot in robots.txt and also allowing requests from its published IP ranges.9 Check the CDN's security events for the crawler and the rule that matched.
02Does Cloudflare block Googlebot by default?
Its 2026 AI defaults for new domains block Training and Agent bots on pages with ads and leave Search allowed.3 Googlebot is more often caught by custom firewall rules, rate limits or country blocks, so confirm in your logs that verified Googlebot requests get 200s.
03Can I block Google-Extended at my CDN?
No. Google-Extended has no separate user agent, and crawling uses existing Google user agents.11 It only works as a robots.txt token. A CDN rule aimed at it would block Googlebot.
04Is a curl request with a crawler's user agent a reliable test?
It reliably shows rules that match the user agent, which is how Cloudflare's Free plan identifies AI crawlers.5 It cannot show rules based on verified identity, because your IP is not on any crawler's list. Confirm with server logs and the CDN's events.
05Should I block ChatGPT-User and other user-triggered fetchers at the edge?
Only if you want assistants unable to read your pages when people ask about them. OpenAI says robots.txt rules may not apply to ChatGPT-User,9 so the edge is the only control that is enforced. For most businesses that want to be recommended, I would allow them.
References17 sources

Sources

  • Platform docs14
  • Regulator2
  • Industry study1
  1. 01robots.txt settingCloudflare Docs (updated August 2026)Platform docs
  2. 02Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for A New Business ModelCloudflare press release (July 2025)Platform docs
  3. 03New options to manage AI trafficCloudflare Changelog (July 2026)Platform docs
  4. 04The crawl-to-click gap: Cloudflare data on AI bots, training, and referralsCloudflare Blog (August 2025)Industry study
  5. 05Manage AI crawlersCloudflare AI Crawl Control docs (updated July 2026)Platform docs
  6. 06Get started with Bot Fight ModeCloudflare Docs (updated August 2026)Platform docs
  7. 07AWS WAF Bot Control rule groupAWS WAF Developer GuidePlatform docs
  8. 08Perplexity CrawlersPerplexityPlatform docs
  9. 09Overview of OpenAI CrawlersOpenAIPlatform docs
  10. 10Does Anthropic crawl data from the web, and how can site owners block the crawler?Claude Help CenterPlatform docs
  11. 11Google's common crawlersGoogle Crawling InfrastructurePlatform docs
  12. 12AI features and your websiteGoogle Search CentralPlatform docs
  13. 13RFC 9110: HTTP Semantics, section 15.5.4 (403 Forbidden)IETF (June 2022)Regulator
  14. 14How HTTP Status Codes Affect Google's CrawlersGoogle Crawling InfrastructurePlatform docs
  15. 15RFC 6585: Additional HTTP Status Codes, section 4 (429 Too Many Requests)IETF (April 2012)Regulator
  16. 16Verified botsCloudflare Docs (updated July 2026)Platform docs
  17. 17Verify requests from Google crawlers and fetchersGoogle Crawling InfrastructurePlatform docs

Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the AI search guide.

Keep reading

Questions about any of this?

LinkedIn is the easiest place to reach me. Send a message or connect, and I will reply there.