Technical SEO is the part of search work that decides whether a page can be found at all. Before relevance, links or content quality come into it, a crawler has to discover the URL, be allowed to fetch it, get a usable response, see the content, and decide that this URL (and not one of its duplicates) is the one to keep. Each of those steps can fail quietly, and a failure at any one of them makes everything after it irrelevant.
The audience for that machinery has grown. Googlebot and Bingbot are still the crawlers that matter most for traffic, but OpenAI, Anthropic and Perplexity now run their own crawlers for search features, training and user-requested fetches, and they do not all behave like Googlebot. This guide covers both. It teaches the mechanism first, then what to do about it, then how to check that it worked, chapter by chapter in roughly the order a crawler meets each problem. The last chapter turns it into an audit you can run.
Several chapters link to deeper pages. Site moves have their own guide in website migrations, large catalogues in ecommerce SEO, and visibility in assistants in the AI search guide. If you work with an AI coding agent, the free skill pack on this site has skills for several of these checks, linked where they fit.
The gates a URL passes through before it can rank or be cited
Discover
The crawler learns the URL exists, usually from a link on a page it already knows, a sitemap, or a push notification such as IndexNow.
How search engines and AI crawlers find, fetch, render and index a page
Google describes its own process in three stages: crawling, indexing and serving.1 Crawling covers discovery and download. Googlebot finds new URLs mainly by following links from pages it already knows, for example a category page linking to a new article, and from sitemaps that site owners submit.1 During crawling it renders the page with a recent version of Chrome, because many sites rely on JavaScript to bring content onto the page.1 Indexing is where the content is analysed, similar pages are grouped, and the most representative one is selected as the canonical.1
None of this is guaranteed. Google does not promise to crawl, index or serve a page even when it follows every guideline, and not every page it processes is indexed.1 That matters for how you read reports. A URL listed as "Crawled - currently not indexed" in Search Console has passed the technical gates and been judged not worth keeping for now.2 No amount of technical work will force it in. Technical SEO removes the reasons a page cannot be indexed; it cannot create a reason it should be.
Not every crawler does every stage
Googlebot runs the full pipeline. AI crawlers are more varied, and the useful way to think about them is by purpose. OpenAI runs OAI-SearchBot to surface sites in ChatGPT’s search features, GPTBot to collect content that may be used to train its models, and ChatGPT-User for actions a person takes in ChatGPT, where robots.txt rules may not apply because a user started the request.3 Anthropic runs ClaudeBot for training data, Claude-SearchBot to improve search results and Claude-User for fetches a person asks for, and states that all three respect robots.txt.4 Perplexity runs PerplexityBot for its search results, and Perplexity-User for user requests, which generally ignores robots.txt.5
Training traffic is most of it. Cloudflare’s analysis of AI crawling across a fixed set of its customers found that in July 2025, 79% of AI crawling was for training, 17% for search and 3.2% for user actions, with the training share up from 72% a year earlier.6 The split matters because the crawler you allow or block decides which outcome you get. Blocking a training crawler does not remove you from that company’s search product, and blocking a search crawler can.
| Crawler | Operator and purpose | Follows robots.txt | Runs JavaScript |
|---|---|---|---|
| Googlebot (Smartphone and Desktop) | Google Search, Images, Video, News and Discover7 | Yes, when crawling automatically7 | Yes, in a separate rendering phase8 |
| Google-Extended | A robots.txt token controlling use of content for Gemini training; does not affect inclusion or ranking in Google Search7 | It is a robots.txt token | Not applicable |
| Bingbot | Microsoft Bing search | Yes | Generally yes, though not every framework a modern browser supports9 |
| OAI-SearchBot | OpenAI: ChatGPT search results3 | Yes3 | Not observed in the Vercel and MERJ analysis10 |
| GPTBot | OpenAI: model training3 | Yes3 | Not observed10 |
| ChatGPT-User | OpenAI: user-initiated actions3 | Rules may not apply3 | Not observed10 |
| ClaudeBot, Claude-SearchBot, Claude-User | Anthropic: training, search quality, user fetches4 | Yes, including Crawl-delay4 | ClaudeBot not observed rendering10 |
| PerplexityBot, Perplexity-User | Perplexity: search results, user requests5 | PerplexityBot yes; Perplexity-User generally not5 | Not observed10 |
The rest of this guide follows a URL through those gates. For Google, each one has a report or a tool that shows you what happened. For AI crawlers there is usually no report, so you test the response yourself and read your server logs.
HTTP status codes and what each one tells a crawler
The status code is the first thing a crawler reads in a response, and it decides whether the content is used at all. The meanings come from the HTTP standard, RFC 9110. Search engines layer their own handling on top, and for SEO you need both: what the code means, and what a crawler does with it.
| Code | Meaning in HTTP | What Google’s crawlers do | Use it for |
|---|---|---|---|
| 200 OK | The request succeeded | Passes the content on for processing; indexing is still not guaranteed11 | Every page you want indexed |
| 301 Moved Permanently | The resource has a new permanent URI that future requests should use12 | Follows it; a strong signal that the target should be canonical11 | Pages that have moved for good |
| 308 Permanent Redirect | Like 301, but the request method must not change12 | Treated the same as 30111 | Moved pages, especially where forms or APIs post to them |
| 302 Found and 307 Temporary Redirect | The resource is temporarily elsewhere; 307 keeps the method12 | Follows it; a weak signal that the target should be canonical11 | Genuinely temporary moves |
| 304 Not Modified | For a conditional request: nothing has changed, use the cached copy12 | Signals the content is unchanged11 | Saving crawl effort on unchanged pages |
| 404 Not Found | No current representation found at this URI12 | Content not used; crawl frequency of the URL falls over time11 | Pages that never existed or have no replacement |
| 410 Gone | No longer available, and likely permanent12 | Handled like 40411 | Deliberately removed pages |
| 429 Too Many Requests | The client sent too many requests | Treated as a sign of server overload11 | Short-term crawl throttling only |
| 500 and 503 | Server error; 503 means temporarily unavailable, often with a Retry-After header12 | Crawling slows down; indexed URLs are kept for a while, then dropped11 | Real outages and planned maintenance (503) |
Soft 404s: the error that returns 200
A soft 404 is a page that tells people it is missing, or is empty, while the server returns 200.11 Typical causes are an out-of-stock template that renders "product not available", a search results page with no results, or a single-page application that shows its own "not found" view on a URL the server happily answers. Search Console lists these under "Soft 404" in the Page indexing report.2 The fix is to return a real 404 or 410 for pages that do not exist, and to give pages that do exist enough content to stop looking empty.
Server errors and planned downtime
5xx responses make Google’s crawlers slow down, in proportion to how many URLs are affected, and URLs that keep returning errors are eventually dropped from the index.11 For a planned maintenance window, a 503 with a Retry-After header is the correct response in HTTP terms.12 Keep it short. Returning 500, 503 or 429 is the documented way to cut Googlebot’s crawl rate in an emergency, but only for a couple of hours or a day or two, and URLs that return those codes for several days may be dropped.13 A maintenance page served with a 200 is worse than either: it tells crawlers the error page is the content.
Two more rules catch people out. A 401 or 403 is not a way to slow crawling down.11 And a 204 No Content response means Google has nothing to process.11
Why AI crawlers make clean status codes more urgent
AI crawlers appear to be worse than Googlebot at avoiding dead URLs. In Vercel and MERJ’s December 2024 analysis of crawler traffic on Vercel’s network, 34.82% of ChatGPT’s crawler fetches and 34.16% of Claude’s hit 404 pages, against 8.22% for Googlebot. ChatGPT also spent 14.36% of its fetches on redirects, against 1.49% for Googlebot.10 That was one network over one period, and crawlers change. Still, it suggests that old, broken and redirected URLs waste a larger share of AI crawler visits than of Googlebot’s, which is one more reason to keep redirects tidy and internal links pointing at final URLs.
How to check the status code a crawler gets
Browsers hide status codes and follow redirects silently, so check from the command line. The first command shows the headers of a single response. The second follows every redirect and prints each hop. Running them with a crawler’s user agent string shows whether the server or CDN treats that user agent differently, with one limit covered in the chapter on CDNs: protection that verifies IP addresses will treat your test as an impostor.
# Headers only, no redirects followed
curl -sI https://www.example.com/page
# Follow redirects and print every hop
curl -sIL https://www.example.com/old-page | grep -iE "^(HTTP|location)"
# The same request with a crawler user agent string (copy the current one from the operator’s documentation)
curl -sI -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" https://www.example.com/pageRedirects: which kind, how many hops, and how long to keep them
A redirect tells a crawler that a URL’s content lives somewhere else. For SEO there are three decisions: permanent or temporary, server-side or client-side, and one hop or several.
Use a permanent redirect (301 or 308) whenever the move is permanent. Google’s crawlers treat both as a strong signal that the target should become the canonical URL, and treat 302 and 307 as a weak one.11 Some content management systems and plugins default to 302, so check the code the server actually returns. The admin screen is not proof.
Do redirects on the server. A server-side redirect is an HTTP response, so every crawler sees it, including the ones that never run JavaScript. A JavaScript redirect only works for a crawler that renders the page, and chapter 7 shows that most AI crawlers do not. Meta refresh sits in between: it is in the HTML, but it is still a weaker and slower signal than a status code.
Chains and loops
Google follows up to 10 redirect hops by default and uses the content of the final URL.11 Other crawlers may give up sooner. For robots.txt files specifically, the standard only asks crawlers to follow at least five.14 Long redirect chains are also one of the documented ways sites waste crawl budget.15 The practical rule is one hop: every old URL redirects straight to its final destination, and every internal link points at the final URL, never at a redirect. Sites that have been migrated more than once almost always carry chains, because each migration adds a layer of redirects on top of the last.
Never redirect large numbers of unrelated URLs to the homepage. If a removed page has no close equivalent, a 404 or 410 is the honest answer. How long to keep redirects, and how to build a full map for a site move, are covered in the migrations guide and in how to build a redirect map. If you are mid-move, the migration tutorial takes you through it stage by stage.
robots.txt, meta robots and X-Robots-Tag: crawling versus indexing
There are two separate controls, and most indexing accidents come from mixing them up. robots.txt controls whether a crawler may fetch a URL. Meta robots and the X-Robots-Tag header control what a search engine does with a URL it has fetched: whether to index it, follow its links or show a snippet. Use the first to keep crawlers away from URLs that waste their time. Use the second to keep pages out of results.
How robots.txt is parsed
robots.txt was a de facto convention for decades and became a standard, RFC 9309, in 2022. The rules that matter most in practice:
- A crawler finds the group whose user-agent line matches its product token, case-insensitively. If several groups match, their rules are combined into one.14 A crawler that has its own group ignores the
*group, so rules you want applied to everyone have to be repeated in each named group. - Within a group, the most specific matching rule wins, measured by the length of the path. When an allow and a disallow rule are equally specific, allow should be used.14
- Crawlers must parse at least 500 kibibytes of the file.14 Google ignores anything past 500 KiB.16
- Crawlers should not use a cached copy for more than 24 hours, unless the file is unreachable.14 So a change can take up to a day to apply, and OpenAI gives about 24 hours for OAI-SearchBot to pick up a robots.txt update.3
- Paths are case-sensitive. Google supports
*for any sequence of characters and$for the end of the URL, and supports only the user-agent, allow, disallow and sitemap fields. It does not support crawl-delay.16 Anthropic’s crawlers do honour Crawl-delay.4 - The rules are not a form of access authorisation.14 Well-behaved crawlers obey them; nothing forces anyone else to.
What happens when robots.txt itself fails
The status code of the robots.txt file has its own rules, and they surprise people. If robots.txt returns a 4xx, the crawler may access anything on the host, as if no file existed.14 If it returns a 5xx, the crawler must assume everything is disallowed.14 Google adds detail: a 429 is handled like a server error, crawling stops for the first 12 hours of errors, the last good copy is used for up to 30 days, and after that Google either assumes no restrictions or stops crawling, depending on whether the rest of the site is generally available.16 Robots.txt redirects are followed for at least five hops, after which Google treats the file as a 404.16
How Google handles a robots.txt that keeps returning server errors
Crawling stops
For the first 12 hours of errors, Google stops crawling the site while it keeps trying to fetch robots.txt.
- Crawling stops, day 1 to 1: For the first 12 hours of errors, Google stops crawling the site while it keeps trying to fetch robots.txt.
- Last good copy used, day 2 to 30: For up to 30 days, Google uses the last version of robots.txt it fetched successfully, and keeps trying to fetch a new one.
- Fallback, day 31 to 32: After 30 days, Google assumes there are no crawl restrictions if the site is generally available, or stops crawling if the site has wider availability problems.
The practical lesson is that robots.txt needs to be the most reliable file on the site. A misconfigured CDN or firewall that returns 403 for robots.txt quietly opens the whole site to crawling. One that returns 500 or 503 can stop it.
# Everyone: keep crawlers out of internal search and filtered listings
User-agent: *
Disallow: /search
Disallow: /*?*sort=
Disallow: /*?*colour=
# OpenAI: allow ChatGPT search, opt out of model training
# A named group replaces the * group, so repeat the shared rules
User-agent: OAI-SearchBot
Disallow: /search
Disallow: /*?*sort=
Disallow: /*?*colour=
User-agent: GPTBot
Disallow: /
Sitemap: https://www.example.com/sitemap_index.xmlThat example is illustrative, and the right setup depends on what the business wants from each AI company. How to set up robots.txt for AI crawlers works through three configurations for three goals. Adoption is still low: the 2025 Web Almanac, from HTTP Archive’s crawl of millions of sites, found GPTBot named in 4.5% of desktop robots.txt files and ClaudeBot in 3.6%, both up on 2024. It also found that 84.9% of sites returned a 200 for robots.txt and around 13% returned a 404.17
Meta robots and X-Robots-Tag
Indexing rules go in a robots meta tag in the HTML head or in an X-Robots-Tag HTTP header. Any rule that works in the meta tag also works in the header.18 The common rules are noindex (keep the page out of results), nofollow (do not follow its links), none (both), nosnippet, max-snippet, max-image-preview and unavailable_after.18 A meta tag named for one crawler, such as googlebot, applies only to that crawler, and when rules conflict the more restrictive one applies.18
X-Robots-Tag is not part of any formal specification, but the major search engines support it, and it is the only way to put indexing rules on non-HTML files such as PDFs and images.19 It is also easy to forget, because it does not appear in the page source. In the 2025 Web Almanac, 3.5% of desktop pages carried a noindex in meta robots and only 0.6% used X-Robots-Tag at all.17 When a page is mysteriously not indexed and the HTML looks clean, check the headers.
# Keep a PDF out of search results
X-Robots-Tag: noindex
# Rules for one named crawler only
X-Robots-Tag: googlebot: nofollow
# The equivalent in HTML, inside <head>
<meta name="robots" content="noindex">The combination that does not work
If robots.txt blocks a URL, any indexing rule on it is never seen and is ignored.18 Blocking does not remove the URL from the index either: Search Console has a status for exactly this, "Indexed, though blocked by robots.txt", for URLs that are blocked but linked from elsewhere.2 To get a page out of results, let it be crawled and serve noindex. Once it has dropped out, you can block it in robots.txt if crawling it is a real cost. To keep something private, use authentication. Neither directive protects it.
What each control does
| Stops crawling | Keeps out of results | Works on PDFs and images | Binds badly behaved bots | |
|---|---|---|---|---|
| robots.txt disallowA blocked URL can still be indexed from links, without its content. | ||||
| meta robots noindexOnly works if the page can be crawled and the tag seen. | ||||
| X-Robots-Tag noindexSame rule as the meta tag, sent as a header. | ||||
| 404 or 410Crawling of the URL slows over time, but does not stop. | ||||
| Password or IP allow-listThe only option that keeps everyone out. Use it for staging. |
No Yes Best in the row
Canonicalisation and duplicate URLs
Most sites serve the same content at more URLs than their owners realise. The usual sources are http and https, www and non-www, trailing and non-trailing slashes, upper and lower case paths, tracking and session parameters, sort orders, filter combinations, print versions and the same product in several categories. Search engines group these duplicates and pick one URL to represent the group.1 Canonicalisation is the work of making sure they pick the one you want, and of reducing the number of duplicates they have to deal with in the first place.
What a canonical is allowed to point at
The canonical link relation is a standard of its own, RFC 6596, published in 2012. It says the canonical target should hold content that is duplicative of, or a superset of, the content at the referring URL.20 It also lists what not to do: give a page more than one canonical, point a canonical at a URL that redirects or returns an error such as 404, point at a page that itself declares a different canonical, or point page 2 of a paginated series at page 1, which risks the content on later pages being lost.20 Those rules come from the standard, and they hold for every search engine.
Canonical signals are hints
Canonicalisation methods are signals of different strength. A redirect is a strong signal that its target should be canonical. A rel="canonical" annotation is a strong signal. Inclusion in a sitemap is a weak one. The methods stack, so they work best when they agree.21 HTTPS URLs are preferred over HTTP equivalents as canonical, canonical annotations should use absolute URLs, and a rel="canonical" HTTP header works for non-HTML files such as PDFs.21 Do not use robots.txt or URL removal tools to canonicalise, and do not give conflicting canonicals through different methods.21
<!-- In the <head> of every duplicate, and of the canonical itself -->
<link rel="canonical" href="https://www.example.com/shoes/trail-runner">
# For a PDF, as an HTTP response header
Link: <https://www.example.com/guides/sizing.pdf>; rel="canonical"Signals that should all agree on one URL
rel="canonical"
https://www.example.com/shoes/trail-runner
Matches.XML sitemap
https://www.example.com/shoes/trail-runner
Matches.Main navigation link
/shoes/trail-runner/
Trailing slash version, which redirects. Internal links should use the final URL.
Does not match.
The canonical URL
https://www.example.com/shoes/trail-runner
The URL you want indexed
http:// version
301 to the https URL
Matches.Category page link
/running/trail-runner?ref=cat
A parameter creates a duplicate URL with its own internal links.
Does not match.hreflang alternate
https://www.example.com/shoes/trail-runner
Matches.
4 of 6 records agree
When the search engine disagrees with you
The Page indexing report shows the outcome. "Alternate page with proper canonical tag" means the duplicate points at an indexed canonical and nothing needs doing. "Duplicate without user-selected canonical" means Google picked a canonical because you did not. "Duplicate, Google chose different canonical than user" means it overrode yours.2 The last one is the useful signal. It almost always means your other signals contradict your canonical tag, or the two pages are not as similar as you think. A canonical pointing at a page with different content tends to be ignored, which is consistent with the standard’s requirement that the target be a duplicate or superset.20
Canonical tags are now on most of the web. The 2025 Web Almanac found them on 68% of desktop pages and 67% of mobile pages, up from 65% a year earlier. It also compared the raw HTML with the rendered page: only in around 2% of cases was a canonical missing from the raw HTML but present after rendering.17 If your site is in that 2%, move the canonical into the server response, for the reasons in the JavaScript chapter.
Find canonical conflicts in a crawl export
Fill in the parts in brackets before you send it
XML sitemaps and IndexNow
Links are the main way crawlers find pages. Sitemaps and push protocols such as IndexNow are the ways a site can tell them directly. They help most on large sites, new sites with few links, and pages that change often.
The sitemap protocol
The format is defined at sitemaps.org. Each entry is a <url> element with a required <loc>, and optional <lastmod>, <changefreq> and <priority>.22 A sitemap can hold no more than 50,000 URLs and be no larger than 50MB uncompressed; larger sites split into several files listed in a sitemap index.22 Files must be UTF-8, with characters such as the ampersand escaped.22 All URLs in a sitemap must be on one host, and a sitemap’s location limits which URLs it may list: a file at /catalog/sitemap.xml can only list URLs under /catalog/.22 Hosting the file at the root avoids that problem.23
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/shoes/trail-runner</loc>
<lastmod>2026-09-30</lastmod>
</url>
<url>
<loc>https://www.example.com/guides/fitting?size=uk&width=wide</loc>
<lastmod>2026-08-14</lastmod>
</url>
</urlset>What the search engines use from it
Google ignores <priority> and <changefreq>, and uses <lastmod> only if it is consistently and verifiably accurate. It should change when the content, structured data or links change, not when a copyright year ticks over.23 Bing’s product team made the same point about lastmod in a July 2025 post: it should reflect when the page content last changed, not when the sitemap file was generated, and Bing uses it to prioritise crawling. The same post says Bing typically revisits a submitted sitemap at least once a day.24 Both engines want fully qualified, absolute URLs.23
So a good sitemap lists only canonical, indexable URLs that return 200, with honest lastmod dates. A sitemap full of redirects, noindexed pages and duplicates is a weak canonical signal pointing the wrong way, and it trains crawlers to distrust the file. Submit it in Search Console and Bing Webmaster Tools, and reference it with a Sitemap: line in robots.txt so every other crawler can find it.23
IndexNow
IndexNow is a push protocol. Instead of waiting for a crawler to notice a change, the site tells participating search engines that a URL has been added, updated or deleted.25 Ownership is proven with a key: a text file on the host containing the key, which must be between 8 and 128 characters.26 A single URL can be submitted with a GET request, and up to 10,000 URLs at once with a JSON POST.26 A submission to one participating engine is shared with all of them.26
POST /indexnow HTTP/1.1
Host: api.indexnow.org
Content-Type: application/json; charset=utf-8
{
"host": "www.example.com",
"key": "your-key-here",
"keyLocation": "https://www.example.com/your-key-here.txt",
"urlList": [
"https://www.example.com/shoes/trail-runner",
"https://www.example.com/shoes/discontinued-model"
]
}Read the responses carefully. A 200 only means the URL was received, not that it will be crawled or indexed. A 403 means the key is invalid, a 422 means the URLs do not match the host or key, and a 429 means you are sending too much.26 The participants listed in the IndexNow FAQ are Amazon, Bing, Naver, Seznam.cz, Yandex and Yep. Google is not among them.25 IndexNow is worth setting up on any site that changes often and cares about Bing or the other participants, but it does nothing for Google.
IndexNow complements sitemaps; it does not replace them. The protocol’s own guidance is to push high-priority or frequently changing URLs through IndexNow and keep sitemaps as the complete inventory, and not to resubmit the same URL many times a day without a real change.25 Many platforms can do it for you. Cloudflare’s Crawler Hints, for example, uses cache signals to tell IndexNow engines when content has probably changed.27
JavaScript rendering and what non-rendering crawlers see
A page built in the browser by JavaScript reaches different crawlers in different states. Googlebot sees it eventually. Most AI crawlers see only what the server sent. Which state a crawler sees decides whether your content exists for it at all.
How Google renders
Google processes JavaScript pages in three phases: crawling, rendering and indexing.8 Every page that returns 200 is queued for rendering, and a page may wait in that queue for a few seconds or longer.8 The rules that follow from that design are worth knowing by heart:
- Links are only discovered if they are
<a>elements with anhrefattribute.8 A button or a div with a click handler is not a link to a crawler. - Client-side routing should use the History API. Loading different content on different URL fragments does not create separate pages.8
- If Google sees noindex in the initial HTML, it may skip rendering, so JavaScript that removes the tag may never run.8
- JavaScript may set the canonical, but should not change it to a different URL from the one in the HTML.8
- A single-page application’s error views need a real 404: either redirect to a URL the server answers with 404, or add a noindex to the error view.8
- Fingerprint JavaScript file names, such as main.2bb85551.js, so Google does not render with an outdated cached file.8
What AI crawlers see
The best public evidence on AI crawlers and JavaScript is still Vercel and MERJ’s December 2024 analysis of traffic on Vercel’s network. It found that none of the major AI crawlers it measured rendered JavaScript, naming OpenAI, Anthropic, Meta, ByteDance and Perplexity. ChatGPT’s crawler fetched JavaScript files in 11.50% of its requests and Claude’s in 23.84%, but neither executed them. The exceptions were Gemini, which uses Googlebot’s infrastructure, and Applebot, which renders with a browser-based crawler.10
That is one analysis from one network, and it is close to two years old. No AI company publishes a rendering specification, so nobody outside them knows for certain what each crawler does today. I treat it as the working assumption until someone shows otherwise: if the copy, internal links, canonical or structured data of a page only exist after JavaScript runs, assume ChatGPT search, Claude and Perplexity cannot use them.
Who reads what on a client-rendered page
- Fetches the HTML
- Queues the page for rendering
- Runs the JavaScript in Chrome
- Indexes the rendered page
- Fetches the HTML
- Can render JavaScript
- Not every framework is supported
- Fetch the HTML
- Some download JavaScript files
- Do not execute them
- Use only the first response
Bing, dynamic rendering, and advice that has aged
Bing’s position is older and worth reading with its date in mind. In October 2018, two Bing program managers wrote that Bingbot can generally render JavaScript but does not support every framework a modern browser does, and that processing JavaScript at scale is hard. They recommended dynamic rendering, serving prerendered HTML to Bingbot and client-side rendering to people, and said it was not cloaking as long as the content was the same.9 Google now calls dynamic rendering a workaround and not a long-term solution, and points to server-side rendering, static rendering or hydration instead, while agreeing that dynamic rendering is not cloaking if the content matches.28 The two positions are less opposed than they look. Both want crawlers to receive complete HTML. They differ on whether a separate path for bots is an acceptable way to get there, and the server-rendered options give every crawler the same complete HTML without one.
| Approach | How it works | What the raw HTML contains | My view for public pages |
|---|---|---|---|
| Client-side rendering | The server sends a shell; JavaScript fetches data and builds the page | Little or nothing beyond the shell | Avoid for any page you want found |
| Server-side rendering | The server builds the full HTML on each request | The complete page | Good default for dynamic pages |
| Static generation | Pages are built to HTML at deploy time | The complete page | Best for marketing pages and articles |
| Hydration | Server or static HTML, then JavaScript attaches interactivity | The complete page | The normal pattern in modern frameworks |
| Dynamic rendering | Bots get prerendered HTML, people get client-side rendering | Complete for bots you detect, a shell for those you miss | A stopgap, documented as a workaround28 |
JavaScript weight is a performance problem as well as a crawling one. The 2024 Web Almanac found a median of 558 KB of JavaScript on mobile pages and 613 KB on desktop, and estimated that about 44% of the JavaScript bytes delivered at the median on mobile went unused during page load.29
How to test what crawlers see
The test is a comparison between the raw HTML and the rendered page. Fetch the raw response with curl, save the rendered DOM with headless Chrome, and check that the title, main copy, internal links, canonical, meta robots and structured data appear in both. At scale, a crawler with a rendering mode does the comparison for you: Sitebulb’s Response vs Render report, for example, compares meta robots, canonical, title, meta description and internal and external links between the response HTML and the rendered HTML, and needs its Chrome Crawler switched on.30 For Google’s own view, use the URL Inspection tool’s live test in Search Console.
# What a non-rendering crawler gets
curl -s -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" https://www.example.com/pricing > raw.html
# What a rendering crawler gets (Chrome or Chromium installed)
chrome --headless --dump-dom https://www.example.com/pricing > rendered.html
# Does a key sentence survive without JavaScript?
grep -c "Billed annually" raw.html rendered.htmlThe full method, including what to change in common frameworks, is in how to check what crawlers see on a JavaScript site. If you build with an AI coding agent, the Reach skill runs the same check: it fetches real pages, reads the served HTML before any JavaScript runs, checks the gatekeepers, and moves content into the server response where it is missing.
Site architecture and internal linking
Internal links do three jobs at once. They are the main route by which crawlers discover pages, they tell search engines which pages the site itself considers important, and their anchor text describes the target. Every page you care about should have a link from at least one other page on your site, and that link has to be an <a> element with an href for Google to crawl it.31 Anchor text should be descriptive, reasonably concise and relevant to the page it points to.31
Depth, hubs and orphans
Google’s own example of discovery is a hub page, such as a category, linking to a new article.1 That is the pattern to build on. A clear hierarchy of hub pages (home, sections, categories, then detail pages) gives crawlers short paths to every page and gives people the same. There is no published click-depth limit. My working rule is that any page you want to rank should be reachable within a few clicks of the homepage through ordinary navigation, without relying on search boxes, filters or pagination alone.
An orphan page is one that exists but has no internal links. A crawl from the homepage cannot find it by definition, so you find orphans by comparing the crawl with other lists of URLs: the XML sitemaps, analytics landing pages, Search Console and server logs. Log tools can do the matching: Screaming Frog’s Log File Analyser, for example, can import a crawl and match it against the URLs in the logs to surface orphans.32 An orphan that gets traffic needs links. An orphan nobody visits may not need to exist.
URL structure
Readable, stable URLs make every other chapter easier. Pick one form for case, trailing slashes and hostname, and redirect the others to it. Keep parameters for things that change the content, and use the standard & separator, since commas, semicolons and brackets are hard for crawlers to detect as parameter separators.33 Keep filter parameters in a consistent order, so the same combination always produces the same URL.33 A URL that changes whenever a menu is reorganised will need redirects every time it does.
Pagination
Paginated series are where canonical mistakes are most common. Each page in a series lists different items, so page 2 is not a duplicate of page 1, and the canonical standard specifically warns against pointing later pages at the first.20 Give each page a self-referencing canonical, link the pages to each other with ordinary <a href> links, and make sure the items on deep pages are also reachable some other way, through subcategories or related links, so their discovery does not depend on a crawler paging through dozens of listings.
Crawl budget on large sites
Crawl budget is the amount of crawling a search engine is willing and able to do on a site. For most sites it is not a constraint, and time spent on it is time not spent on content. Google’s crawl budget guide is written for three groups: large sites with over a million pages whose content changes moderately often, medium or larger sites with over 10,000 pages whose content changes very quickly, and sites with a large share of URLs stuck in "Discovered - currently not indexed".15 If you are not in one of those groups, the rest of this chapter is mostly about hygiene.
Capacity and demand
Two forces set the budget. The crawl capacity limit is how much crawling the site can take without strain, and it rises and falls with how quickly and reliably the server responds. Crawl demand is how much the search engine wants to crawl, driven by how many URLs it thinks exist, how popular they are and how stale its copies are.15 A fast, healthy server raises the ceiling, but only demand fills it.
What sets how much of a site gets crawled
01Capacity
How much crawling the server can take. Fast, error-free responses raise it; slow responses and 5xx errors lower it.
02Demand
How much the crawler wants: driven by the URLs it knows about, their popularity, and how stale its copies are.
03Waste
Duplicates, endless parameter combinations, soft 404s and redirect chains that use up crawling without adding anything.
Where crawl budget goes to waste
The documented causes of waste are duplicate URLs, infinite scrolling or listing pages that duplicate linked content, soft 404s, long redirect chains, and URLs you do not want crawled but have not blocked.15 The fixes follow: consolidate duplicates, block genuinely useless URL patterns such as filter combinations in robots.txt, return 404 or 410 for pages removed for good, keep lastmod accurate in sitemaps, speed up responses, and support conditional requests so unchanged pages can be answered with a 304.15
Faceted navigation
Faceted navigation is the most common source of runaway crawling. Each filter combination creates a URL that looks new, and a crawler cannot know it is useless without fetching it.33 The primary fixes are to disallow the facet patterns in robots.txt, or to implement filters with URL fragments, which Google generally does not crawl.33 rel="canonical" and nofollow on filter links can reduce crawling over time but are described as less effective in the long run.33 Filter combinations with no results should return 404.33 A few facet combinations do deserve indexing because people search for them. Which faceted navigation pages to index sets out a method for choosing them, and the ecommerce guide covers catalogues more widely.
AI crawlers and server load
AI crawlers add load that search crawl budget models were not built for. Vercel and MERJ counted 569 million GPTBot requests and 370 million from Claude across Vercel’s network in a month, together about 20% of Googlebot’s 4.5 billion over the same period.10 Cloudflare’s 2025 Radar Year in Review, which counts Googlebot as an AI crawler because it crawls for both search and AI training, found that Googlebot alone accounted for 4.5% of HTML request traffic on its network and all other AI bots together for 4.2%.34 If AI crawling strains your servers, robots.txt can opt out of crawlers you get nothing from, and Anthropic’s crawlers honour Crawl-delay.4 For Google, the short-term lever is temporary 500, 503 or 429 responses, with the risks covered in the status code chapter.13
Measuring it
Search Console’s Crawl Stats report shows total crawl requests, download size and average response time, a breakdown by response code, file type, purpose and Googlebot type, and host status for robots.txt fetching, DNS resolution and server connectivity. It is only available for root-level properties.35 It is a sample of Google’s view. For everything else, including Bing and AI crawlers, you need logs.
Log file analysis: what crawlers actually request
Every other tool in this guide shows what a crawler could do or what one search engine chooses to report. Server logs show what every crawler actually requested and what your server sent back. They are the only source for AI crawler behaviour on your site, the most reliable way to find crawl waste, and the evidence you need before arguing for crawl budget work.
What a log line contains
A typical access log line records the client IP address, timestamp, request method and path, status code, response size, referrer and user agent. That is enough to answer most questions. Logs can come from the origin server, a load balancer or the CDN. If a CDN serves cached pages, the origin logs will be missing those requests, so take logs from the CDN where you can.
66.249.66.1 - - [06/Oct/2026:08:14:02 +0000] "GET /shoes/trail-runner HTTP/1.1" 200 48211 "-" "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.0.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"First, separate real crawlers from impostors
Anyone can send a Googlebot user agent string, and scrapers often do. For Google, verify a request either by a reverse DNS lookup on the IP, checking the hostname ends in googlebot.com, google.com or googleusercontent.com, and then a forward lookup that resolves back to the same IP; or by matching the IP against the ranges Google publishes in JSON files for its common crawlers, special crawlers and user-triggered fetchers.36 OpenAI publishes IP lists for OAI-SearchBot, GPTBot and ChatGPT-User,3 Perplexity for both its crawlers,5 and Anthropic publishes the IP ranges its bots use.4 Analysis that skips this step counts fake bots as real ones.
# Reverse lookup, then forward lookup to confirm
host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1
# Status codes served to anything calling itself Googlebot (verify the IPs separately)
grep "Googlebot" access.log | awk '{print $9}' | sort | uniq -c | sort -rn
# The paths OAI-SearchBot requests most often
grep "OAI-SearchBot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20The questions logs answer
- What share of each crawler’s requests go to URLs you do not want crawled: parameters, filters, internal search, redirects and 404s? On a large site, that number is the case for crawl budget work.
- How often are your most valuable templates crawled, and how long after publishing does a new page get its first visit from each crawler?
- Which status codes do crawlers receive, and does any crawler get a different answer from the one users get, such as 403s from bot protection?
- Which URLs are crawled but not linked internally (orphans), and which important URLs are never crawled at all?
- Which AI crawlers visit, which pages they fetch, and whether search and user-triggered fetchers such as OAI-SearchBot, ChatGPT-User and Perplexity-User reach the pages you want cited.
- Did a change work? After a robots.txt update, a redirect cleanup or a migration, the logs show whether crawler behaviour changed.
User-triggered fetches deserve a note. ChatGPT-User, Claude-User and Perplexity-User visit because a person asked an assistant something that led it to your page.345 They are as close as logs get to a record of your pages being used in AI answers. They are not a complete one, since assistants can also answer from an index built earlier.
Tools
For a one-off look, command-line tools on a log sample are enough. For regular work, a dedicated tool saves time: Screaming Frog’s Log File Analyser can verify search engine bots automatically, report crawl frequency and response codes per URL, and import crawl data to compare against the logs. Its free version is limited to 1,000 log lines.32 On large sites, logs usually end up in a data warehouse or an observability platform where they can be queried by template. The log analysis skill in this site’s skill pack walks an AI agent through the same order: get the logs, identify the real crawlers, then analyse.
Summarise crawler activity from a log sample
Fill in the parts in brackets before you send it
Page experience and Core Web Vitals, in proportion
Core Web Vitals are three field metrics for the experience of loading and using a page. Largest Contentful Paint (LCP) measures loading, and a good score is within 2.5 seconds. Interaction to Next Paint (INP) measures responsiveness, and good is under 200 milliseconds. Cumulative Layout Shift (CLS) measures visual stability, and good is under 0.1.37 INP replaced First Input Delay as a Core Web Vital in 2024.38 Each is assessed at the 75th percentile of page loads, separately for mobile and desktop, so a page passes only if three in four real visits meet the threshold.38
How much they matter for ranking
Core Web Vitals are used by Google’s ranking systems, but there is no single page experience signal, and Google says it will still show the most relevant content even when the page experience is poor.39 Read together, that means speed can separate pages that are otherwise similar, but it will not lift a page that is less relevant than its competitors. In my judgement, fixing a failing template is worth doing for users and conversions in its own right, and for ranking it belongs after crawlability, indexing and content problems, not before.
Most of the web is not there yet, which gives some sense of proportion. The 2025 Web Almanac, using Chrome UX Report data from July 2025, found that 48% of origins passed all three Core Web Vitals on mobile and 56% on desktop. On mobile, 62% had good LCP, 77% good INP and 81% good CLS.40 LCP is the metric most sites fail, and it is usually the one to start with.
Field data first, lab data second
Field data from real Chrome users (the Chrome UX Report, shown in PageSpeed Insights and in Search Console’s Core Web Vitals report) is what the thresholds are measured against. Lab tools such as Lighthouse and Chrome DevTools reproduce a problem under fixed conditions so you can find its cause, but lab measurement is not a substitute for field data.38 Field data is collected over weeks of real visits, so a fix takes weeks to show up in it. Judge a fix in the lab first, then confirm it in the field.
Server speed also feeds back into crawling. A faster, more reliable server raises Google’s crawl capacity limit, and slow or erroring responses lower it.15 On a large site that is often a better argument for performance work than ranking. The performance skill diagnoses LCP by sub-part, INP by phase and CLS by cause, starting from field data.
HTTPS, hosting, CDNs and bot protection
Everything in front of your application can change what a crawler receives: DNS, TLS, the CDN, the firewall and the bot management rules. These layers are usually owned by a different team from SEO, and changes to them are rarely tested against crawlers.
The layers a crawler request passes through
- 01
DNS
Resolves the hostname. Failures show in Search Console’s Crawl Stats host status.
- Resolves
- Times out
- Wrong records
- 02
TLS and HTTPS
Certificate and protocol. Expired or mismatched certificates stop crawlers as well as browsers.
- Valid certificate
- HSTS
- http to https redirect
- 03
CDN and cache
May serve a cached copy, a cached error or a different variant from the origin.
- Cache hit
- Cache miss
- Cached error
- 04
Firewall and bot management
Can allow, challenge, rate limit or block a request based on IP, user agent and behaviour, whatever robots.txt says.
- Allow
- Challenge
- Block (403)
- Rate limit (429)
- 05
Origin server
The application. Returns the status code, headers and HTML you think you are serving.
- 200
- 3xx
- 4xx
- 5xx
HTTPS
HTTPS is now the norm: the 2025 Web Almanac found 91.7% of desktop pages served over it.17 For SEO the work is consolidation. Every http URL should permanently redirect to its https equivalent in a single hop, and canonicals, sitemaps and internal links should all use https. When both versions exist, the HTTPS URL is the preferred canonical.21 The Strict-Transport-Security header tells browsers to use HTTPS for the host on every future visit, and only works once a browser has received it over a secure connection.41 Be careful with its includeSubDomains and preload options, which apply the policy to every subdomain, and preload requires a max-age of at least a year.41 Check that no subdomain still needs plain http before turning them on.
Strict-Transport-Security: max-age=63072000; includeSubDomainsHosting and DNS
Search Console’s Crawl Stats report flags host problems with robots.txt fetching, DNS resolution and server connectivity. For DNS, for example, failures on more than 5% of requests on a given day count as a problem.35 Persistent server errors slow Google’s crawling and eventually drop URLs from the index,11 and a robots.txt that returns 5xx stops crawling altogether, as covered in chapter 4.16 When moving hosts, compare response times and error rates in the logs before and after, and keep the old host running until crawler traffic to it stops.
CDNs and bot protection
Bot protection is now the most common way a site blocks crawlers it meant to allow. A firewall rule that challenges or blocks a crawler returns a 403, a challenge page or a 429 before robots.txt is consulted, and a crawler that cannot pass a challenge receives nothing. Cloudflare, for example, keeps a list of verified bots: bots that identify themselves honestly through cryptographic signatures, published IP lists or reverse DNS, and that behave well, including respecting robots.txt.42 Firewall rules can let those bots through while still challenging others. When your CDN blocks AI crawlers covers the settings and how to test them across providers.
Testing is harder than it looks. Protection that checks IP addresses will treat a curl request with Googlebot’s user agent from your laptop as an impostor and block it, which tells you nothing about what the real Googlebot gets. Use the URL Inspection live test for Google. For other crawlers, look in the CDN’s firewall event log and your server logs for requests from verified IP ranges, and check the status codes they received.
Allowing AI crawlers is a business decision
Whether to let AI crawlers in is not a purely technical question, and the traffic numbers explain why owners hesitate. Cloudflare measured crawl-to-referral ratios, the number of pages crawled for every visit referred back. For July 2025 it reported 38,065 crawls per referral for Anthropic, 1,091 for OpenAI, 194 for Perplexity and 5.4 for Google.6 Its 2025 Year in Review reported Anthropic’s ratio reaching as much as 500,000 to 1 at points in the year.34 These are measured on Cloudflare’s network, and referrals undercount AI influence, because an assistant can name a business without sending a click. Still, they show that training crawlers send very little back.
My default for most businesses that want to be found is to allow the search and user-triggered crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User), because those are what put a site into AI answers, and to decide on training crawlers separately. Publishers whose content is the product may reasonably decide otherwise. The AI search guide covers what happens after the crawl.
A practical technical SEO audit, in order
An audit should follow the order of the gates, because a problem early in the pipeline hides every problem after it. There is no point tuning canonicals on pages a firewall blocks, or tuning Core Web Vitals on pages that are noindexed. Work through the checks below in order, and fix blockers as you find them.
For a site about to launch, the launch QA skill runs the gatekeeper checks against the production URL and splits the results into launch blockers and fix-this-week items. If you are auditing because traffic has already fallen, start with the traffic drop tutorial, which separates technical causes from algorithm updates and demand changes before you spend time on fixes.
Turn audit findings into a prioritised fix list
Fill in the parts in brackets before you send it
Common questions
01What is the difference between crawling and indexing?
02Should I use robots.txt or noindex to keep a page out of Google?
03Do AI crawlers like GPTBot and ClaudeBot run JavaScript?
04Does Google use IndexNow?
05Is a 301 or a 308 redirect better for SEO?
06Do I need to worry about crawl budget?
07Why did Google ignore my canonical tag?
08How much do Core Web Vitals affect rankings?
09Can my CDN block Googlebot or AI crawlers even if robots.txt allows them?
10How do I know a request really came from Googlebot?
Sources
- Platform docs32
- Regulator4
- Industry study6
- 01In-depth guide to how Google Search worksGoogle Search CentralPlatform docs
- 02Page indexing reportSearch Console HelpPlatform docs
- 03Overview of OpenAI CrawlersOpenAIPlatform docs
- 04Does Anthropic crawl data from the web, and how can site owners block the crawler?AnthropicPlatform docs
- 05Perplexity CrawlersPerplexityPlatform docs
- 06The crawl-to-click gap: Cloudflare data on AI bots, training, and referralsCloudflareIndustry study
- 07Google's common crawlersGoogle for DevelopersPlatform docs
- 08Understand JavaScript SEO BasicsGoogle Search CentralPlatform docs
- 09bingbot Series: JavaScript, Dynamic Rendering, and Cloaking. Oh My!Bing Webmaster BlogPlatform docs
- 10The rise of the AI crawlerVercelIndustry study
- 11How HTTP Status Codes Affect Google's CrawlersGoogle for DevelopersPlatform docs
- 12RFC 9110: HTTP SemanticsIETFRegulator
- 13Reduce the Google crawl rateGoogle for DevelopersPlatform docs
- 14RFC 9309: Robots Exclusion ProtocolIETFRegulator
- 15Crawl Budget Management For Large SitesGoogle for DevelopersPlatform docs
- 16How Google interprets the robots.txt specificationGoogle for DevelopersPlatform docs
- 17SEO: 2025 Web AlmanacHTTP ArchiveIndustry study
- 18Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsGoogle Search CentralPlatform docs
- 19X-Robots-Tag headerMDN Web DocsPlatform docs
- 20RFC 6596: The Canonical Link RelationIETFRegulator
- 21How to specify a canonical with rel="canonical" and other methodsGoogle Search CentralPlatform docs
- 22Sitemaps XML formatsitemaps.orgRegulator
- 23Build and Submit a SitemapGoogle Search CentralPlatform docs
- 24Keeping Content Discoverable with Sitemaps in AI Powered SearchBing Webmaster BlogPlatform docs
- 25IndexNow FAQIndexNowPlatform docs
- 26IndexNow documentationIndexNowPlatform docs
- 27Crawler HintsCloudflare DocsPlatform docs
- 28Dynamic rendering as a workaroundGoogle Search CentralPlatform docs
- 29JavaScript: 2024 Web AlmanacHTTP ArchiveIndustry study
- 30Response vs Render ReportSitebulbPlatform docs
- 31SEO Link Best Practices for GoogleGoogle Search CentralPlatform docs
- 32Log File AnalyserScreaming FrogPlatform docs
- 33Managing crawling of faceted navigation URLsGoogle for DevelopersPlatform docs
- 34The 2025 Cloudflare Radar Year in Review: the rise of AI, post-quantum, and record-breaking DDoS attacksCloudflareIndustry study
- 35Crawl Stats reportSearch Console HelpPlatform docs
- 36Verify requests from Google crawlers and fetchersGoogle for DevelopersPlatform docs
- 37Understanding Core Web Vitals and Google search resultsGoogle Search CentralPlatform docs
- 38Web Vitalsweb.dev (Google)Platform docs
- 39Understanding page experience in Google Search resultsGoogle Search CentralPlatform docs
- 40Performance: 2025 Web AlmanacHTTP ArchiveIndustry study
- 41Strict-Transport-Security headerMDN Web DocsPlatform docs
- 42Verified botsCloudflare DocsPlatform docs