Skip to content
← All writing
SaaS17 September 2026Updated 7 October 202610 min read13 sources

How to check what crawlers see on a JavaScript site

Googlebot renders JavaScript later, most AI crawlers never do, and a few commands show you exactly which parts of your SaaS site each of them can read.

By Mani Bharij, SEO & AI Search Consultant

In this postHover or tap a point
01020304050607

Chapter 012 min

How Googlebot renders a JavaScript page

A SaaS marketing site built with React can look perfect in a browser and arrive almost empty at a crawler. The browser runs the JavaScript and builds the page. Googlebot does the same, but later and in a separate step. Many AI crawlers do not run it at all, so they read only the HTML the server sent. The test that matters is what sits in the HTML before any script runs, and what changes after. This post shows how to run that test with curl and PowerShell, how to read Google's own view of the rendered page, and which rendering choices fix what you find.

The wider technical picture for software companies, from app subdomains to staging environments, is in the SaaS SEO guide. This post takes its section on JavaScript rendering and turns it into a method you can run in an afternoon.

Chapter 01 / 072 min

How Googlebot renders a JavaScript page

JavaScript pages go through three phases at Google: crawling, rendering and indexing.1 Googlebot fetches the URL, checks robots.txt first, and parses the HTML it gets back. The page then joins a render queue. It may stay there for a few seconds, or for longer.1 Rendering happens in an evergreen version of Chromium, which executes the JavaScript, and the rendered HTML is what gets indexed.1

You will still hear this called the two-wave model. At I/O in 2018, Google told developers that indexing and rendering do not happen at the same time, and that it may defer rendering to a later point.2 SEOs summarised that as two waves: index the raw HTML first, then render and index again. Google's current documentation describes phases and a queue instead of waves. The practical lesson is the same either way. Anything that exists only after JavaScript runs depends on an extra step that happens later, and anything in the raw HTML does not.

What the Web Rendering Service does not do

The rendering component is the Web Rendering Service, or WRS. It does not keep state between page loads: local storage, session storage and cookies are cleared each time.3 So content that only appears after a visitor accepts a cookie banner, picks a region or logs in is unlikely to be in what Google renders. Three more rules from the same JavaScript guidance catch SaaS sites often:

  • A noindex in the initial HTML may stop Google rendering the page at all, so JavaScript that removes it may never run.1
  • Set the canonical in the HTML. If JavaScript has to touch it, it should set the same value the HTML already has.1
  • Google only discovers links that are <a> elements with an href attribute.1 A button with a click handler that changes the route is not a link to a crawler.

The canonical rule is one most sites already follow. HTTP Archive's 2025 Web Almanac found that in only around 2% of cases was a canonical missing from the raw HTML but present in the rendered DOM.4 If your site is in that group, the fix is usually a small template change.

Seeing the page as Google rendered it

The URL Inspection tool in Search Console and the Rich Results Test both show the loaded resources, JavaScript console output and the rendered DOM.3 In URL Inspection, run Test live URL and open View tested page. A screenshot of the rendered page is only available in a live test, and the live result can differ from the indexed version.5 Use the live test to check a fix, and the indexed result to see what Google stored.

Chapter 02 / 071 min

What AI crawlers read when they skip JavaScript

The most detailed public evidence comes from a log analysis Vercel and MERJ published in December 2024. It found that none of the major AI crawlers it measured rendered JavaScript, naming OpenAI's OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic's ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot.6 ChatGPT's and Claude's crawlers did fetch JavaScript files, in 11.50% and 23.84% of their requests, but did not execute them.6 The exceptions were Gemini, which uses Googlebot's infrastructure, and Applebot, which renders pages in a browser-based crawler.6

Read that with its limits in mind. It is one analysis of traffic on one network, from late 2024, and crawler behaviour can change without notice. I still treat it as the working assumption for any page you want an assistant to read: if the copy is not in the server's HTML, ChatGPT search, Claude and Perplexity probably cannot see it. Vercel's own recommendation was to server-render important content for this reason.6 MDN's web docs list the same benefit for server-side rendering: search engines, social media crawlers and other bots can read the content without executing JavaScript, and while major search engines can run JavaScript, social media crawlers usually cannot.7

Diagram

Who reads what on a client-rendered pricing page

One URL, content built in the browser
  • Fetches the HTML and checks robots.txt
  • Queues the page for rendering
  • Runs the JavaScript in the Web Rendering Service
  • Indexes the rendered HTML
  • Fetch the HTML the server returns
  • Some fetch JavaScript files too
  • Do not execute them
  • Read only what was in the first response
  • Render JavaScript, per the Vercel analysis
  • Gemini through Googlebot's infrastructure
  • Applebot through a browser-based crawler
Based on Google's documentation and the Vercel and MERJ analysis from December 2024. Behaviour can change, so test your own pages.

Rendering is the second gate. The first is whether you let these crawlers in at all, which is covered in how to set up robots.txt for AI crawlers. A crawler you allow but that receives an empty shell is no better off than one you block. For how that missing copy affects which products an assistant names, see how an assistant decides which businesses to name.

Chapter 03 / 071 min

A worked example: one pricing page, two versions

Illustrative. Rotawise is a fictional staff scheduling product. Its marketing site is a single-page React app: the server returns the same small HTML file for every route, and the browser fetches plan data from an API and draws the page. In a browser the pricing page shows three plans, a comparison table, an FAQ and links to integrations. This is what each version contains.

ElementRaw HTML from the serverRendered DOMWho misses it
Title"Rotawise" on every route, from the shared index file"Rotawise pricing: plans for care homes and clinics", set by a scriptNon-rendering crawlers see the same title on every page
H1 and body copyNone. An empty <div id="root"></div>H1, three plans, prices, FAQ answersMost AI crawlers
Navigation and internal linksNoneHeader, footer and integration linksMost AI crawlers; Googlebot finds them only after rendering
CanonicalMissingAdded by a head library at runtimeNon-rendering crawlers; Google prefers it in the HTML
Structured dataMissingInjected JSON-LD for the productNon-rendering crawlers
Illustrative: the Rotawise pricing page before and after rendering

Google can read JSON-LD that a script adds to the DOM, because it processes the structured data present when it renders the page.8 A crawler that never renders gets none of it. If Rotawise moved its marketing routes to static generation, the raw HTML would carry the same content as the rendered DOM, and the only thing left to JavaScript would be the monthly and annual price toggle.

Chapter 04 / 071 min

Checking the raw HTML with curl and PowerShell

Start with a handful of URLs that matter: the homepage, pricing, one feature page, one integration page, one comparison page and one blog post. Fetch each one exactly as a non-rendering crawler would, save it, and look for the things that should be there. Pick a sentence from the visible copy for each page, such as a price line, so you can search for it.

Bash, on macOS, Linux or WSL

Checklist0 of 8 done

PowerShell, on Windows

These work in Windows PowerShell 5.1 and PowerShell 7. Fetch once and keep the response in a variable.

Checklist0 of 8 done

Run the fetch a second time with a crawler's user agent string in place of Mozilla/5.0, for example GPTBot. If the two responses differ, something on the server or CDN is treating bots differently, which is worth knowing before you trust either result. A spoofed user agent does not reproduce rules based on IP address, so a bot management layer can still treat the real crawler differently; server logs are the place to confirm that.

Chapter 05 / 071 min

Comparing the raw HTML with the rendered DOM

The raw HTML tells you what a non-rendering crawler gets. The rendered DOM tells you what a browser, and eventually Googlebot, builds from it. The gap between the two is the part of the page that depends on JavaScript.

  • Headless Chrome can save the rendered DOM. In bash: google-chrome --headless --dump-dom https://www.example.com/pricing > rendered.html. In PowerShell: & "C:\Program Files\Google\Chrome\Application\chrome.exe" --headless --dump-dom "https://www.example.com/pricing" | Out-File rendered.html -Encoding utf8. Then run the same counts on rendered.html.
  • To list links that exist only after rendering: diff <(grep -o 'href="[^"]*"' raw.html | sort -u) <(grep -o 'href="[^"]*"' rendered.html | sort -u)
  • For a quick visual check, open DevTools, press Ctrl+Shift+P (Cmd+Shift+P on a Mac), and run Disable JavaScript. It stays off in that tab while DevTools is open.9 Reload and look at what is left.
  • For Google's own view, use URL Inspection's live test, and the Rich Results Test for structured data. Prefer the URL input to pasting code, because the code input has JavaScript limitations.8
What you findWhat it meansWhat I would do
Raw and rendered match on headings, copy, links, canonical and JSON-LDEvery crawler gets the full pageNothing. Add the check to the release process.
Headings and copy in the raw HTML, but links or JSON-LD only after renderingPartly server-rendered, with client-side additionsMove links and structured data into the server output.
An empty root element in the raw HTMLClient-side renderingPrerender or server-render the marketing routes.
Canonical or robots meta differs between the twoJavaScript is changing directivesSet them in the HTML and stop scripts touching them.
The GPTBot fetch differs from the browser fetchUser agent based serving or bot rulesCheck the CDN and bot management settings.
Reading the comparison
Chapter 06 / 072 min

Choosing between SSR, SSG and ISR

The fix for an empty raw HTML is to produce the HTML before it reaches the crawler. The vocabulary comes from web.dev: server-side rendering sends HTML from the server for each request, static rendering builds HTML for each URL at build time, client-side rendering builds the page in the browser with JavaScript, and hydration attaches client-side scripts to server-rendered HTML to make it interactive.10

In the Next.js App Router, layouts and pages are Server Components by default, and Client Components are also prerendered to HTML on the server for the first page load.11 Marking a component 'use client' does not on its own hide it from crawlers. What does is data fetched in the browser after load, such as plans pulled from an API inside an effect, or content that waits for a browser-only API. That data is not in the prerendered HTML, so move the fetch into a Server Component.

Incremental Static Regeneration sits between static and server rendering. It updates static pages without rebuilding the whole site, using a revalidate time or on-demand revalidation.12 Next.js documents an x-nextjs-cache response header with the values HIT, STALE, MISS and REVALIDATED,12 so curl -sI https://www.example.com/pricing | grep -i x-nextjs-cache shows whether a page came from the cache, where your hosting passes the header through.

Page typeWhat I would useWhy
Homepage, features, use casesStatic generationChanges only on release, and needs to be complete for every crawler
PricingStatic generation or ISRPrices must be in the HTML; ISR lets finance update them without a deploy
Integration and template directoriesISRHundreds of pages from a dataset that changes weekly
Blog and docsStatic generation or ISRHigh volume, mostly unchanged after publication
Logged-in appClient-side rendering is fineNot meant for search; keep it on noindex or behind a login
Rendering choices for a SaaS site (my judgement)
Diagram

How much of the page a non-rendering crawler receives

Empty shellComplete HTML
  • Client-side renderingClient-side rendering: 5 percent of the way from Empty shell to Complete HTML.
  • Server shell, data fetched in the browserServer shell, data fetched in the browser: 40 percent of the way from Empty shell to Complete HTML.
  • Dynamic renderingDynamic rendering: 80 percent of the way from Empty shell to Complete HTML.
  • Server-side renderingServer-side rendering: 95 percent of the way from Empty shell to Complete HTML.
  • Incremental Static RegenerationIncremental Static Regeneration: 95 percent of the way from Empty shell to Complete HTML.
  • Static generationStatic generation: 100 percent of the way from Empty shell to Complete HTML.

Hover a row for the main thing that goes wrong.

Practitioner judgement, assuming the rendering is set up correctly. ISR and static generation differ in freshness, not in completeness.

Avoid dynamic rendering for a new build. Google describes it as a workaround and not a long-term solution, and recommends server-side rendering, static rendering or hydration instead.13 If the site is a plain client-side React app, the usual fix is a framework or build step that prerenders the marketing routes. If that change also alters URLs, treat it as a website migration and plan the redirects before launch.

Chapter 07 / 071 min

Making the check part of every release

Rendering regressions arrive quietly. A developer moves a pricing table into a client component, a new head library sets the canonical at runtime, and nothing looks different in the browser. A short check on each release catches it before a crawler does.

Checklist0 of 4 done

The Reach skill in my open-source pack runs this check: it fetches real pages, reads the HTML served before any JavaScript runs, and moves content into the server response where it is missing. Whatever runs it, the habit matters more than the tool. Compare what the server sends with what the browser builds, and keep the gap small.

Questions5 answered

Common questions

01Does Google index content that is rendered with JavaScript?
Yes. Google renders pages with an evergreen version of Chromium and indexes the rendered HTML, but pages wait in a render queue first, which can take a few seconds or longer.1 Content in the server's HTML does not depend on that step.
02Do ChatGPT, Claude and Perplexity crawlers run JavaScript?
Not according to the Vercel and MERJ analysis from December 2024, which found that OpenAI's crawlers, ClaudeBot and PerplexityBot did not execute JavaScript, although ChatGPT's and Claude's crawlers fetched some JavaScript files.6 Crawler behaviour can change, so check your own pages and logs.
03How do I see the rendered HTML Google used?
Use the URL Inspection tool in Search Console, or the Rich Results Test. Both show the rendered DOM, loaded resources and console output.3 The screenshot is only available in a live test.5
04Is a Next.js 'use client' component invisible to crawlers?
No. Next.js prerenders Client Components to HTML on the server for the first page load.11 Content goes missing when the component fetches its data in the browser after the page loads, so that data never reaches the server HTML.
05Should I use dynamic rendering for bots?
Not for a new build. Dynamic rendering is a workaround and not a long-term solution, and server-side rendering, static rendering or hydration are the documented alternatives.13
References13 sources

Sources

  • Platform docs11
  • Industry study2
  1. 01Understand the JavaScript SEO basicsGoogle Search CentralPlatform docs
  2. 02Google Search at I/O 2018Google Search Central Blog (June 2018)Platform docs
  3. 03Fix Search-related JavaScript problemsGoogle Search CentralPlatform docs
  4. 04SEO: 2025 Web AlmanacHTTP ArchiveIndustry study
  5. 05URL Inspection toolSearch Console HelpPlatform docs
  6. 06The rise of the AI crawlerVercel and MERJ (December 2024)Industry study
  7. 07Server-side rendering (SSR)MDN Web DocsPlatform docs
  8. 08Generate structured data with JavaScriptGoogle Search CentralPlatform docs
  9. 09Disable JavaScriptChrome for DevelopersPlatform docs
  10. 10Rendering on the Webweb.devPlatform docs
  11. 11Server and Client ComponentsNext.js DocsPlatform docs
  12. 12How to implement Incremental Static Regeneration (ISR)Next.js DocsPlatform docs
  13. 13Dynamic Rendering as a workaroundGoogle Search CentralPlatform docs

Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the SaaS SEO guide.

Keep reading

Questions about any of this?

LinkedIn is the easiest place to reach me. Send a message or connect, and I will reply there.