A SaaS marketing site built with React can look perfect in a browser and arrive almost empty at a crawler. The browser runs the JavaScript and builds the page. Googlebot does the same, but later and in a separate step. Many AI crawlers do not run it at all, so they read only the HTML the server sent. The test that matters is what sits in the HTML before any script runs, and what changes after. This post shows how to run that test with curl and PowerShell, how to read Google's own view of the rendered page, and which rendering choices fix what you find.
The wider technical picture for software companies, from app subdomains to staging environments, is in the SaaS SEO guide. This post takes its section on JavaScript rendering and turns it into a method you can run in an afternoon.
How Googlebot renders a JavaScript page
JavaScript pages go through three phases at Google: crawling, rendering and indexing.1 Googlebot fetches the URL, checks robots.txt first, and parses the HTML it gets back. The page then joins a render queue. It may stay there for a few seconds, or for longer.1 Rendering happens in an evergreen version of Chromium, which executes the JavaScript, and the rendered HTML is what gets indexed.1
You will still hear this called the two-wave model. At I/O in 2018, Google told developers that indexing and rendering do not happen at the same time, and that it may defer rendering to a later point.2 SEOs summarised that as two waves: index the raw HTML first, then render and index again. Google's current documentation describes phases and a queue instead of waves. The practical lesson is the same either way. Anything that exists only after JavaScript runs depends on an extra step that happens later, and anything in the raw HTML does not.
What the Web Rendering Service does not do
The rendering component is the Web Rendering Service, or WRS. It does not keep state between page loads: local storage, session storage and cookies are cleared each time.3 So content that only appears after a visitor accepts a cookie banner, picks a region or logs in is unlikely to be in what Google renders. Three more rules from the same JavaScript guidance catch SaaS sites often:
- A
noindexin the initial HTML may stop Google rendering the page at all, so JavaScript that removes it may never run.1 - Set the canonical in the HTML. If JavaScript has to touch it, it should set the same value the HTML already has.1
- Google only discovers links that are
<a>elements with anhrefattribute.1 A button with a click handler that changes the route is not a link to a crawler.
The canonical rule is one most sites already follow. HTTP Archive's 2025 Web Almanac found that in only around 2% of cases was a canonical missing from the raw HTML but present in the rendered DOM.4 If your site is in that group, the fix is usually a small template change.
Seeing the page as Google rendered it
The URL Inspection tool in Search Console and the Rich Results Test both show the loaded resources, JavaScript console output and the rendered DOM.3 In URL Inspection, run Test live URL and open View tested page. A screenshot of the rendered page is only available in a live test, and the live result can differ from the indexed version.5 Use the live test to check a fix, and the indexed result to see what Google stored.
What AI crawlers read when they skip JavaScript
The most detailed public evidence comes from a log analysis Vercel and MERJ published in December 2024. It found that none of the major AI crawlers it measured rendered JavaScript, naming OpenAI's OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic's ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot.6 ChatGPT's and Claude's crawlers did fetch JavaScript files, in 11.50% and 23.84% of their requests, but did not execute them.6 The exceptions were Gemini, which uses Googlebot's infrastructure, and Applebot, which renders pages in a browser-based crawler.6
Read that with its limits in mind. It is one analysis of traffic on one network, from late 2024, and crawler behaviour can change without notice. I still treat it as the working assumption for any page you want an assistant to read: if the copy is not in the server's HTML, ChatGPT search, Claude and Perplexity probably cannot see it. Vercel's own recommendation was to server-render important content for this reason.6 MDN's web docs list the same benefit for server-side rendering: search engines, social media crawlers and other bots can read the content without executing JavaScript, and while major search engines can run JavaScript, social media crawlers usually cannot.7
Who reads what on a client-rendered pricing page
- Fetches the HTML and checks robots.txt
- Queues the page for rendering
- Runs the JavaScript in the Web Rendering Service
- Indexes the rendered HTML
- Fetch the HTML the server returns
- Some fetch JavaScript files too
- Do not execute them
- Read only what was in the first response
- Render JavaScript, per the Vercel analysis
- Gemini through Googlebot's infrastructure
- Applebot through a browser-based crawler
Rendering is the second gate. The first is whether you let these crawlers in at all, which is covered in how to set up robots.txt for AI crawlers. A crawler you allow but that receives an empty shell is no better off than one you block. For how that missing copy affects which products an assistant names, see how an assistant decides which businesses to name.
A worked example: one pricing page, two versions
Illustrative. Rotawise is a fictional staff scheduling product. Its marketing site is a single-page React app: the server returns the same small HTML file for every route, and the browser fetches plan data from an API and draws the page. In a browser the pricing page shows three plans, a comparison table, an FAQ and links to integrations. This is what each version contains.
| Element | Raw HTML from the server | Rendered DOM | Who misses it |
|---|---|---|---|
| Title | "Rotawise" on every route, from the shared index file | "Rotawise pricing: plans for care homes and clinics", set by a script | Non-rendering crawlers see the same title on every page |
| H1 and body copy | None. An empty <div id="root"></div> | H1, three plans, prices, FAQ answers | Most AI crawlers |
| Navigation and internal links | None | Header, footer and integration links | Most AI crawlers; Googlebot finds them only after rendering |
| Canonical | Missing | Added by a head library at runtime | Non-rendering crawlers; Google prefers it in the HTML |
| Structured data | Missing | Injected JSON-LD for the product | Non-rendering crawlers |
Google can read JSON-LD that a script adds to the DOM, because it processes the structured data present when it renders the page.8 A crawler that never renders gets none of it. If Rotawise moved its marketing routes to static generation, the raw HTML would carry the same content as the rendered DOM, and the only thing left to JavaScript would be the monthly and annual price toggle.
Checking the raw HTML with curl and PowerShell
Start with a handful of URLs that matter: the homepage, pricing, one feature page, one integration page, one comparison page and one blog post. Fetch each one exactly as a non-rendering crawler would, save it, and look for the things that should be there. Pick a sentence from the visible copy for each page, such as a price line, so you can search for it.
Bash, on macOS, Linux or WSL
PowerShell, on Windows
These work in Windows PowerShell 5.1 and PowerShell 7. Fetch once and keep the response in a variable.
Run the fetch a second time with a crawler's user agent string in place of Mozilla/5.0, for example GPTBot. If the two responses differ, something on the server or CDN is treating bots differently, which is worth knowing before you trust either result. A spoofed user agent does not reproduce rules based on IP address, so a bot management layer can still treat the real crawler differently; server logs are the place to confirm that.
Comparing the raw HTML with the rendered DOM
The raw HTML tells you what a non-rendering crawler gets. The rendered DOM tells you what a browser, and eventually Googlebot, builds from it. The gap between the two is the part of the page that depends on JavaScript.
- Headless Chrome can save the rendered DOM. In bash:
google-chrome --headless --dump-dom https://www.example.com/pricing > rendered.html. In PowerShell:& "C:\Program Files\Google\Chrome\Application\chrome.exe" --headless --dump-dom "https://www.example.com/pricing" | Out-File rendered.html -Encoding utf8. Then run the same counts onrendered.html. - To list links that exist only after rendering:
diff <(grep -o 'href="[^"]*"' raw.html | sort -u) <(grep -o 'href="[^"]*"' rendered.html | sort -u) - For a quick visual check, open DevTools, press Ctrl+Shift+P (Cmd+Shift+P on a Mac), and run Disable JavaScript. It stays off in that tab while DevTools is open.9 Reload and look at what is left.
- For Google's own view, use URL Inspection's live test, and the Rich Results Test for structured data. Prefer the URL input to pasting code, because the code input has JavaScript limitations.8
| What you find | What it means | What I would do |
|---|---|---|
| Raw and rendered match on headings, copy, links, canonical and JSON-LD | Every crawler gets the full page | Nothing. Add the check to the release process. |
| Headings and copy in the raw HTML, but links or JSON-LD only after rendering | Partly server-rendered, with client-side additions | Move links and structured data into the server output. |
| An empty root element in the raw HTML | Client-side rendering | Prerender or server-render the marketing routes. |
| Canonical or robots meta differs between the two | JavaScript is changing directives | Set them in the HTML and stop scripts touching them. |
The GPTBot fetch differs from the browser fetch | User agent based serving or bot rules | Check the CDN and bot management settings. |
Choosing between SSR, SSG and ISR
The fix for an empty raw HTML is to produce the HTML before it reaches the crawler. The vocabulary comes from web.dev: server-side rendering sends HTML from the server for each request, static rendering builds HTML for each URL at build time, client-side rendering builds the page in the browser with JavaScript, and hydration attaches client-side scripts to server-rendered HTML to make it interactive.10
In the Next.js App Router, layouts and pages are Server Components by default, and Client Components are also prerendered to HTML on the server for the first page load.11 Marking a component 'use client' does not on its own hide it from crawlers. What does is data fetched in the browser after load, such as plans pulled from an API inside an effect, or content that waits for a browser-only API. That data is not in the prerendered HTML, so move the fetch into a Server Component.
Incremental Static Regeneration sits between static and server rendering. It updates static pages without rebuilding the whole site, using a revalidate time or on-demand revalidation.12 Next.js documents an x-nextjs-cache response header with the values HIT, STALE, MISS and REVALIDATED,12 so curl -sI https://www.example.com/pricing | grep -i x-nextjs-cache shows whether a page came from the cache, where your hosting passes the header through.
| Page type | What I would use | Why |
|---|---|---|
| Homepage, features, use cases | Static generation | Changes only on release, and needs to be complete for every crawler |
| Pricing | Static generation or ISR | Prices must be in the HTML; ISR lets finance update them without a deploy |
| Integration and template directories | ISR | Hundreds of pages from a dataset that changes weekly |
| Blog and docs | Static generation or ISR | High volume, mostly unchanged after publication |
| Logged-in app | Client-side rendering is fine | Not meant for search; keep it on noindex or behind a login |
How much of the page a non-rendering crawler receives
- Client-side renderingClient-side rendering: 5 percent of the way from Empty shell to Complete HTML.
- Server shell, data fetched in the browserServer shell, data fetched in the browser: 40 percent of the way from Empty shell to Complete HTML.
- Dynamic renderingDynamic rendering: 80 percent of the way from Empty shell to Complete HTML.
- Server-side renderingServer-side rendering: 95 percent of the way from Empty shell to Complete HTML.
- Incremental Static RegenerationIncremental Static Regeneration: 95 percent of the way from Empty shell to Complete HTML.
- Static generationStatic generation: 100 percent of the way from Empty shell to Complete HTML.
Hover a row for the main thing that goes wrong.
Avoid dynamic rendering for a new build. Google describes it as a workaround and not a long-term solution, and recommends server-side rendering, static rendering or hydration instead.13 If the site is a plain client-side React app, the usual fix is a framework or build step that prerenders the marketing routes. If that change also alters URLs, treat it as a website migration and plan the redirects before launch.
Making the check part of every release
Rendering regressions arrive quietly. A developer moves a pricing table into a client component, a new head library sets the canonical at runtime, and nothing looks different in the browser. A short check on each release catches it before a crawler does.
The Reach skill in my open-source pack runs this check: it fetches real pages, reads the HTML served before any JavaScript runs, and moves content into the server response where it is missing. Whatever runs it, the habit matters more than the tool. Compare what the server sends with what the browser builds, and keep the gap small.
Common questions
01Does Google index content that is rendered with JavaScript?
02Do ChatGPT, Claude and Perplexity crawlers run JavaScript?
03How do I see the rendered HTML Google used?
04Is a Next.js 'use client' component invisible to crawlers?
05Should I use dynamic rendering for bots?
Sources
- Platform docs11
- Industry study2
- 01Understand the JavaScript SEO basicsGoogle Search CentralPlatform docs
- 02Google Search at I/O 2018Google Search Central Blog (June 2018)Platform docs
- 03Fix Search-related JavaScript problemsGoogle Search CentralPlatform docs
- 04SEO: 2025 Web AlmanacHTTP ArchiveIndustry study
- 05URL Inspection toolSearch Console HelpPlatform docs
- 06The rise of the AI crawlerVercel and MERJ (December 2024)Industry study
- 07Server-side rendering (SSR)MDN Web DocsPlatform docs
- 08Generate structured data with JavaScriptGoogle Search CentralPlatform docs
- 09Disable JavaScriptChrome for DevelopersPlatform docs
- 10Rendering on the Webweb.devPlatform docs
- 11Server and Client ComponentsNext.js DocsPlatform docs
- 12How to implement Incremental Static Regeneration (ISR)Next.js DocsPlatform docs
- 13Dynamic Rendering as a workaroundGoogle Search CentralPlatform docs
Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the SaaS SEO guide.