Skip to content
← All topics
06AI search / Study guide

AI search: how assistants find, read and cite websites, and what it changes about SEO

How AI Overviews, AI Mode, ChatGPT, Perplexity, Copilot and Claude turn a prompt into a cited answer, which crawlers and controls decide whether your pages are in it, how to measure it, and where the work sits for SaaS, ecommerce, travel, B2B, publishing and local sites.

chapters
13chapters
to read
32 minto read
sources
30sources
questions
9questions

By Mani Bharij · Updated 7 October 2026

Chapter mapHover or tap a point
01020304050607080910111213

Chapter 011 min

How your customers prompt AI search, and what each prompt needs from your site

AI search is the name that has stuck for answers written by a language model on top of a search system: Google AI Overviews and AI Mode, ChatGPT search, Perplexity, Microsoft Copilot, Claude with web search, and others. Someone types or speaks a prompt, the system decides whether it needs the web, fetches and reads pages, and returns prose with a handful of links. This guide covers the mechanics, what each provider has documented, and where I think the work sits for different kinds of business.

This part of search attracts the most confident claims and has the least evidence behind them. I tie each claim to the provider's own documentation, to independent studies or to published research where they exist, and mark general practice and my own judgement where it does not. The products change monthly, so read every dated statement as true in October 2026.

Chapter 01 / 131 min

How your customers prompt AI search, and what each prompt needs from your site

Your customers use AI search in three ways. They ask it things, they tell it to do things, and increasingly they hand it whole tasks to finish. Each of those prompts pulls in different pages from your site, and most content plans are only built for the first.

Kind of promptExampleWhat the system has to find
A question“Is a heat pump worth it for a 1930s semi?” or “What does SOC 2 Type II actually test?”Explanations, guides and opinion. The classic informational page.
An instruction“Draft a migration plan for moving a Shopify store to a new domain.” “Compare these three CRMs in a table with price per seat and API limits.”Hard facts in a usable form: pricing, specifications, limits, documentation, process steps. The answer is a new artefact built from many sources.
A task for an agent“Find a waterproof jacket under £150 in medium and put it in my basket.” “Book a table for four near the venue on Friday.”Live availability, prices, policies and a page or protocol an agent can act on. Some providers now complete purchases and bookings inside the chat.
Three ways customers prompt AI search, and the pages each one needs

A question can be answered from a good explainer. An instruction such as “compare these three CRMs in a table” cannot: the system needs each vendor's pricing, plan limits and integration list, and it will take them from wherever they are stated most clearly, which may be a review site and not the vendor. A task needs the stock level, the returns window and the checkout to be machine-readable. In Google's own words, AI Mode suits queries needing “further exploration, reasoning, or complex comparisons”, and people can ask nuanced questions that might previously have taken several searches.1 OpenAI documents ChatGPT finding restaurant availability and, depending on the reservation provider, completing the booking flow inside the chat.2

Being “in the answer” can mean being cited as a source, being named as an option, or being the page whose data an agent uses to act. They overlap, and all of them start with retrieval.

Chapter 02 / 131 min

Two routes into an answer: training data and live retrieval

A model either already knows something because it was in its training data, or it looks it up while answering. The research literature calls these parametric and non-parametric memory. The paper that named retrieval-augmented generation described combining a model's built-in knowledge with a separate searchable index of documents, and found the combination produced more specific and factual answers than the model alone.3

Parametric knowledge is what the model absorbed when it was trained. It is frozen at a training cutoff, it cannot be updated for your business on request, and the influence of any single page on it is unknowable. Live retrieval happens at answer time: the system runs searches, fetches pages, and writes from what it read. For anything current, priced, local, or specific to one product, retrieval is where the answer comes from.

The two routes are fed by different crawlers. That clears up most of the confusion about robots.txt and AI. Blocking a training crawler changes what future models may learn from your site, and has no documented effect on whether today's assistants can find and cite you. Blocking a search crawler does the opposite.

RouteCrawlers and fetchersWhat blocking it does
Training data (parametric)GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (a robots.txt token for Gemini model training)Signals that your content should not be used to train future models. OpenAI says this is independent of its search crawler.4 Anthropic says blocking ClaudeBot excludes future material from training datasets.5
Live search indexOAI-SearchBot, Claude-SearchBot, PerplexityBot, GooglebotRemoves or limits your pages in that provider's search answers. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.4
Live fetch on behalf of a userChatGPT-User, Claude-User, Perplexity-UserVaries. Anthropic says disabling Claude-User stops it retrieving your content for a user's query.5 OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person initiated the request.46
Which crawler feeds which route, as documented by each operator in October 2026
Diagram

How much control you have over each input to an answer

Little controlMost control
  • What a model learned in trainingWhat a model learned in training: 10 percent of the way from Little control to Most control.
  • Third-party reviews, comparisons and forumsThird-party reviews, comparisons and forums: 30 percent of the way from Little control to Most control.
  • Profiles and listings you manageProfiles and listings you manage: 60 percent of the way from Little control to Most control.
  • Product feeds and structured dataProduct feeds and structured data: 80 percent of the way from Little control to Most control.
  • Your own pages in a live search indexYour own pages in a live search index: 90 percent of the way from Little control to Most control.

Hover a row for the main thing that goes wrong.

My judgement, matching the text. Control falls as you move from your own live pages to what a model learned in training.

Use the scale to set priorities. The inputs you control most are also the ones retrieval reads. The ones you control least, training data and third-party opinion, get most of the commentary.

Chapter 03 / 133 min

How retrieval-augmented generation turns a prompt into a cited answer

The system searches first and writes second. For Google this is documented directly: retrieval-augmented generation, also called grounding, relies on its core Search ranking systems to retrieve relevant, up-to-date pages from the Search index, and the model then reviews specific information from those pages to generate a response with links to them.7 Microsoft's transparency note for Copilot describes the same idea in general terms: for conversations where people seek information, Copilot is grounded in web search results, centres its response on high-ranking content and links to it.8 Other providers describe pieces of the same pipeline. The steps below mix documented behaviour with general industry practice, and the text says which is which.

Diagram

From prompt to cited answer

01

Prompt

A question, an instruction or a task, often long and specific, sometimes with an image. The system first decides whether it needs the web at all.

Step 1 of 6
A general model of the pipeline. Google and OpenAI document parts of it; the reading and reranking steps are general retrieval practice, and no provider publishes its exact internals.

Understanding and rewriting the prompt

OpenAI says that when ChatGPT search uses partner search providers it typically rewrites the prompt into one or more targeted queries, and may send further, more specific queries after reviewing the first results.2 It may add general location inferred from the IP address, and if memory is on, saved details such as dietary preferences can shape the rewritten query.2 The term “query fan-out” came from Google at the launch of AI Mode, for issuing multiple related searches concurrently across subtopics and multiple data sources, then bringing the results together.9 In a later Google example, a prompt about a lawn full of weeds fans out into searches about herbicides, chemical-free removal and prevention.7

So the searches that decide whether you are retrieved are not the prompt the person typed. They are narrower, more literal sub-queries. You can see real examples of them: Bing Webmaster Tools now reports the “grounding queries” its AI used when retrieving content it later cited.10

Diagram

One prompt, several searches, one answer

Prompt

Compare three project management tools for a 40-person agency that bills by the hour, in a table with price per seat and built-in time tracking.

Searches it might run

  1. 01project management software with time tracking for agenciesFinds [1] [4]
  2. 02project management tool pricing per user annual billingFinds [2]
  3. 03project management tool invoicing integration accounting softwareFinds [3]
  4. 04agency project management tool reviews billable hoursFinds [1] [4]

Pages it could retrieve

  1. [1]A comparison page on a software review siteLists candidate tools and their headline features side by side.
  2. [2]Each vendor's pricing pageStates price per seat, billing terms and which plan includes time tracking.
  3. [3]The vendor's integrations docsConfirms which accounting tools it connects to, by name.
  4. [4]A forum thread from agency ownersGives first-hand views on how time tracking works in practice.
Answer

A three-row table with price per seat, whether time tracking is native or an add-on, and invoicing integrations, citing the pricing pages and docs for the figures and the review site for the shortlist. 1234

Illustrative only. No provider publishes the real fan-out for a given prompt; these show the shape of it, using invented prompts and generic page types.

Retrieving candidates from an index

Each provider documents this differently. Google retrieves from its own Search index using its core ranking systems.7 OpenAI says ChatGPT search sometimes partners with other search providers, and its help article points readers to Microsoft's and Shopify's privacy policies for how those providers process queries.2 Perplexity says its public answer engine runs on its own index covering hundreds of billions of web pages.11 Anthropic's help page for Claude's web search says Claude processes multiple sources and that every response includes citations, but does not name the index behind text results.12 The detail is in the table further down this page.

Ranking, reranking and reading passages

A first-stage retriever is built to be fast and returns many candidates. Retrieval systems commonly add a second, more expensive ranking pass over a short list, and then pass only the best sections to the model, because a model's working context is limited and every extra page costs time and money. That two-stage pattern is general industry practice; no AI search provider has published its exact reranking. Perplexity has said the most about the passage part: its indexing divides documents into “fine-grained units” that are individually scored against the query, so relevant snippets come back already ranked.11 Google's systems likewise review “specific information” from retrieved pages.7

Writing the answer and attaching citations

Providers describe this step only in outline. The model writes from the retrieved material and links to pages that support what it said. Google's models also identify further supporting pages while the response is being generated, which is why AI features can show a wider set of links than a classic results page.1 OpenAI notes that search results and citations can be incomplete, outdated or wrong, and tells users to open the source and check it.2 Citation is not proof that a page was the main input, and not being cited is not proof it was unused. It is the visible end of the pipeline.

Chapter 04 / 133 min

How an AI system reads a page: passages, chunking and meaning

A retrieval system usually works with sections of a page, so each section needs to make sense on its own. Perplexity's description of sub-document units is the clearest primary statement of this.11 Research systems work the same way: the dense passage retrieval paper that much later work builds on split each Wikipedia article into disjoint 100-word passages, prefixed each one with the article title, and retrieved those passages as its basic units.13 The title prefix matters. A passage that has lost its context is harder to match, which is one reason to name the product, plan or place inside each section.

These writing rules are my reasoning from how passage retrieval works. No provider documents them as ranking rules.

  • State the claim near its heading. A section titled “Pricing” that opens with the price is easy to match and quote. One that opens with brand story is not.
  • Make each section self-contained. Name the product, plan or service in the section itself, so a passage does not depend on “this” or “as mentioned above”.
  • Prefer specifics a sub-query can match: numbers, units, model names, certifications, place names, dates, limits.
  • Use tables for comparable facts. A row such as “Business plan, £18 per user per month, SSO included” survives being lifted out of the page. Microsoft's advice to publishers on AI answers names clear headings, tables and FAQ sections for the same reason.10
  • Cut filler. A sentence with no fact in it gives a system nothing to retrieve and nothing to cite.

The closest thing to a controlled test is academic. The GEO paper (Aggarwal and others, KDD 2024) rewrote source pages in different ways and measured how much of a generated answer drew on each. Adding citations to other sources, quotations or statistics gave a relative improvement of 30 to 40% on its main visibility measure, and keyword stuffing gave little to no improvement.14 The tests ran on GPT-3.5 fed with search results and on Perplexity, using a benchmark the authors built, so read it as evidence for the direction. It does not show that any current assistant behaves the same way.

Diagram

Which passages on a pricing page could answer the prompt

Prompt

Which plan do I need for SSO and 50 users, and what will it cost per year?

A SaaS pricing page
  1. 01

    Business plan: £18 per user per month, billed annually. Includes SAML single sign-on, SCIM provisioning and audit logs.

  2. 02

    Powerful features for teams of every size. Start your journey today.

  3. 03

    Users: Starter up to 10, Team up to 100, Business unlimited.

  4. 04

    As mentioned above, this is included in all the higher tiers.

  5. 05

    Need advanced security? Talk to our sales team.

  6. 06

    Prices exclude VAT. Annual billing is 20% cheaper than paying monthly.

Switch to the retrieval view to see which passages could be lifted

An invented SaaS pricing page judged against one prompt. Usable passages are specific and make sense alone; unusable ones are filler or depend on another section.

Lexical and semantic matching

Older search matched words. Newer retrieval also matches meaning. This is general information-retrieval background; no provider documents its own mix. Lexical methods such as TF-IDF and BM25 score a passage by the query words it contains and how rare those words are. They are fast and precise when the words match, and blind to synonyms. Semantic, or dense, retrieval turns the query and each passage into embeddings, lists of numbers that place similar meanings close together, so “cancel my plan” can match a passage about “ending a subscription”. In 2020 the dense passage retrieval paper reported that a learned dense retriever beat a strong BM25 baseline by 9 to 19 percentage points on top-20 passage retrieval accuracy across open-domain question answering benchmarks.13 Production systems commonly combine both.

Google's AI systems understand synonyms and general meaning, so for Google you do not need to cover every long-tail variant of a phrase.7 But exact terms still matter where the exact term is the fact: a model number, a certification such as ISO 27001, an integration name, a fare class. Write the way your customers and their prompts do, and spell out the specific names.

Chapter 05 / 131 min

Where each assistant's web results come from

Providers vary in how much they say about their index. The table records only what each has documented, and says so where something is undocumented.

ProductIndex or search sourceWhat else is documented
Google AI Overviews and AI ModeGoogle's own Search index, retrieved with its core ranking systems.7Both may use query fan-out across subtopics and data sources. They may use different models and techniques, so their responses and links vary.1
ChatGPT searchOpenAI's OAI-SearchBot crawls for ChatGPT search, and ChatGPT search sometimes partners with other search providers.42Prompts are rewritten into targeted queries. OpenAI says results are ranked on multiple factors, placement is not guaranteed, and eligibility requires allowing OAI-SearchBot, including at the host or CDN.2
PerplexityIts own index covering hundreds of billions of pages, which it says processes tens of thousands of updates a second.11PerplexityBot surfaces and links sites in its results and is not used to crawl content for foundation models.6
Microsoft CopilotWeb search results, for conversations where people seek information. Microsoft's transparency note does not name the index.8The AI Performance report covers Copilot, AI summaries in Bing and select partner integrations, and shows grounding queries.10
ClaudeAnthropic has not documented which index supplies text results. Its help page says image results are powered by Bing.12Every web search response includes citations.12 Anthropic runs Claude-SearchBot to improve search result quality and Claude-User to fetch pages for a user's prompt.5
What each provider has documented about its source of web results, October 2026

Whichever assistant you care about, a page the system cannot fetch, or that its index does not hold, cannot be used. The answer is also assembled from several sources at once, so your site is one input among many. I have written up how an assistant decides which businesses to name as a working model of how that plays out.

Chapter 06 / 132 min

Google's guidance on AI Overviews, AI Mode and your website

Google's position is more direct than most commentary about it: the best practices for SEO remain relevant, and there are no additional requirements to appear in AI Overviews or AI Mode and no other special optimisations necessary.1 To be eligible as a supporting link, a page must be indexed and eligible to be shown in Google Search with a snippet.1 A newer guide to generative AI features, last updated in July 2026, adds one condition: the site must also be included in Search generative AI features in Search Console, which is the default.7

The recommended practices are ordinary ones:1

  • Allow crawling in robots.txt, and by any CDN or hosting infrastructure.
  • Make content findable through internal links.
  • Provide a good page experience.
  • Make sure important content is available as text.
  • Support text with high-quality images and video where it helps.
  • Make sure structured data matches the visible text on the page.
  • Keep Merchant Center and Business Profile information up to date.

The guide goes further than the AI features page on two points. Creating content for every variation of how people search, including fan-out queries, primarily to manipulate rankings or AI responses, violates Google's scaled content abuse policy.7 And it addresses the AEO and GEO labels head on: from Google Search's perspective, optimising for generative AI search is optimising for search, and so still SEO.7

I think Google is right here. If a site is not appearing in AI Overviews, the first places to look are whether it is indexed, whether its snippets are restricted, whether it has been excluded in Search Console, and whether it ranks for the narrower searches a fan-out is likely to run. A separate AI strategy is rarely the missing piece.

What independent data says about clicks

Google's AI features page makes one claim about traffic: “when people click from search results pages with AI Overviews, these clicks are higher quality (meaning, users are more likely to spend more time on the site)”.1 It gives no figure for how many clicks there are. Independent studies have measured that, and they point the same way.

  • Pew Research Center tracked the Google searches of 900 US adults in March 2025. On pages with an AI summary, people clicked a traditional result in 8% of visits, against 15% on pages without one, and clicked a link inside the summary in 1% of visits.15
  • Ahrefs compared 300,000 keywords using aggregated Search Console data, desktop only. With December 2025 data, the presence of an AI Overview correlated with a 58% lower average click-through rate for the top-ranking page. Its April 2025 version of the study had found 34.5%.16
  • Seer Interactive, an agency, measured 3,119 informational queries across 42 client organisations from June 2024 to September 2025. Organic click-through rate on queries with an AI Overview fell from 1.76% to 0.61%, and on queries without one from 2.74% to 1.62%. Brands cited in the AI Overview had a 35% higher organic click-through rate than brands not cited.17

The methods differ, so the percentages cannot be compared with each other, and Ahrefs sells SEO software while Seer sells SEO services. The cited-versus-uncited figure is a correlation; cited brands may differ in other ways. Google's quality claim and the volume data can both be true. For reporting, the practical point is that AI Overview impressions are a poor proxy for traffic, and being cited is the part worth working towards.

Chapter 07 / 133 min

AI crawlers and robots.txt: which bots to allow and what each controls

Most providers now split their crawlers by purpose, so you can make separate decisions about search visibility and model training. The table below follows each operator's own documentation.

CrawlerOperatorPurposeRouterobots.txt token and behaviour
GooglebotGoogleCrawling for Google Search, which includes AI Overviews and AI Mode. Robots.txt rules for Googlebot are the control for how sites are crawled for Search.1LiveGooglebot. Respects robots.txt.
Google-ExtendedGoogleA product token, with no separate user agent. Covers training future Gemini models for Gemini Apps and the Vertex AI API, and grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It does not affect inclusion or ranking in Google Search.18Training, plus grounding outside SearchGoogle-Extended. Read from robots.txt only.
OAI-SearchBotOpenAISurfaces websites in ChatGPT's search features.4LiveOAI-SearchBot. Respects robots.txt; changes take about 24 hours.
GPTBotOpenAICrawls content that may be used to train OpenAI's foundation models.4TrainingGPTBot. Respects robots.txt.
ChatGPT-UserOpenAIVisits pages for certain user actions in ChatGPT and Custom GPTs.4Live, user-triggeredChatGPT-User. OpenAI says robots.txt rules may not apply.
Claude-SearchBotAnthropicCrawls to improve search result quality for Claude users.5LiveClaude-SearchBot. Respects robots.txt.
Claude-UserAnthropicFetches pages when a person asks Claude something.5Live, user-triggeredClaude-User. Anthropic says disabling it stops retrieval for user queries.
ClaudeBotAnthropicCollects content that could contribute to model training.5TrainingClaudeBot. Respects robots.txt and Crawl-delay.
PerplexityBotPerplexitySurfaces and links websites in Perplexity results. Not used for foundation model training.6LivePerplexityBot. Respects robots.txt; changes take up to 24 hours.
Perplexity-UserPerplexityVisits a page when a user's prompt needs it, and may link to it. Not used for training.6Live, user-triggeredPerplexity-User. Perplexity says it generally ignores robots.txt.
AI-related crawlers and fetchers, as documented by each operator in October 2026

OpenAI says each of its settings is independent, so a site can allow OAI-SearchBot and disallow GPTBot.4 Anthropic says its bots honour robots.txt, asks for rules to be set per subdomain, and advises against blocking by IP address, because that can stop it reading your robots.txt at all.5

My default for a business that wants to be recommended is to allow the search crawlers and the user-triggered fetchers. Whether to block training crawlers is a separate decision about how you feel about your content being used to train models. A publisher with a licensing position will reach a different answer from a SaaS company that wants its documentation known, and that is a commercial call more than an SEO one. Cloudflare's network data puts numbers on the trade. In July 2025 Anthropic crawled 38,065.7 pages for every visitor it referred, OpenAI 1,091.4 and Perplexity 194.8, against 5.4 for Google, and over the preceding 12 months 80% of AI crawling was for training, 18% for search and 2% for user actions.19 Those ratios moved a lot within 2025, so check Cloudflare Radar for current figures. I have set out worked robots.txt configurations for each goal.

JavaScript rendering: what AI crawlers can see

If your content only appears after JavaScript runs in a browser, most AI crawlers will not see it. The providers do not document their rendering. The best public evidence is a December 2024 log analysis by Vercel and MERJ, which found that none of the major AI crawlers it measured rendered JavaScript, including OpenAI's OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic's ClaudeBot and PerplexityBot. Google's Gemini and Applebot were the exceptions.20 ChatGPT's and Claude's crawlers did fetch JavaScript files, around 11.5% and 23.8% of their requests respectively, but did not execute them.20

Google itself can process JavaScript as long as it is not blocked, though JavaScript-heavy sites are generally more complex to work on.7 The safe position for every other assistant is server-rendered or statically generated HTML for anything you want quoted: prices, specifications, plan limits, policies, availability. A single-page app that shows a pricing table only after a client-side API call is, as far as the 2024 evidence goes, a blank page to most AI crawlers. The study is nearly two years old, so test your own logs before assuming it still holds for a given crawler.

Chapter 08 / 131 min

Controlling how Google's AI features use your content

There are now two documented sets of controls for AI Overviews and AI Mode. Neither is Google-Extended, the control people most often misunderstand. Google-Extended governs Gemini training and grounding outside Search, and does not affect Search inclusion.18

Page-level: snippet controls

ControlWhere it goesEffect on AI Overviews and AI Mode
nosnippetRobots meta tag or X-Robots-Tag headerNo text snippet in search results, and the content is not used as a direct input for AI Overviews and AI Mode.21
max-snippet:[number]Robots meta tag or X-Robots-Tag headerCaps snippet length and limits how much of the content may be used as a direct input for AI Overviews and AI Mode.21
data-nosnippetHTML attribute on span, div and section elements onlyExcludes the marked text from snippets.21 It is one of the controls that limit what is shown in Search, including AI features.1
noindexRobots meta tag or X-Robots-Tag headerRemoves the page from Google Search entirely, and so from AI features too.1
Google snippet controls and their documented effect on AI features

Site-level: the Search generative AI control

Since 31 August 2026 every Search Console property has a Search generative AI control under Settings.22 Including the site is the default. Excluding it removes your links and content from AI Overviews, AI Mode and generative AI features in Discover, and stops content crawled from your site being used as input for those responses. A change generally takes a few days to apply, and some content takes longer because of caching.22 The control is not used as a ranking or inclusion signal for other parts of Search, it does not affect AI training (Google-Extended covers that), and your content may still be used to help Search understand language generally.22

These are blunt instruments for most sites. nosnippet costs the ordinary snippet as well, and a site-wide exclusion gives up a surface where other sites' content will still appear, possibly similar to yours.22 I would use them only for a specific reason: paywalled or licensed content, or a block of text you do not want quoted out of context, where data-nosnippet on that block is the proportionate choice.

Chapter 09 / 132 min

Agents and commerce: shopping, checkout and product data

Assistants are starting to act as well as answer, and for shops that means product data is read by machines that may complete the purchase. Everything in this section is dated, because it is changing quickly.

ChatGPT

OpenAI's shopping help article says that when a prompt suggests shopping intent, ChatGPT shows products with images, details and links to merchants, and that product results are selected independently and are not ads.23 It lists what ChatGPT considers: structured metadata such as price and description from first-party and third-party providers, other third-party content, and the responses the model generated before it looked at new search results.23 When several merchants sell the same item, they are ranked on factors such as availability, price, quality and whether the merchant is the maker or primary seller.23 Merchants on Shopify are already integrated through Shopify Catalog, and other merchants can apply to send OpenAI a direct product feed.23

In September 2025 OpenAI launched Instant Checkout, built on the Agentic Commerce Protocol it co-developed with Stripe and released as an open standard. The merchant stays the merchant of record and handles payment, fulfilment and returns in its own systems.24 OpenAI said Instant Checkout items are not preferred in product results, but that whether Instant Checkout is enabled is one factor when ranking merchants selling the same product.24 In March 2026 OpenAI said the first version of Instant Checkout lacked the flexibility it wanted, and that merchants can use their own checkout while it concentrates on product discovery. I have set out what that changes for product data.

Google AI Mode and Gemini

Shopping in AI Mode is powered by Google's Shopping Graph, which Google puts at more than 50 billion product listings with 2 billion updated every hour, and AI Mode can show comparison tables with insights drawn from reviews.25 In November 2025 it began rolling out agentic checkout in the US: a shopper tracks a price, and when it falls within budget Google can buy the item on an eligible merchant's site with Google Pay, after the shopper confirms.25 In January 2026 Google announced the Universal Commerce Protocol, an open standard for agentic checkout in AI Mode and the Gemini app with the retailer as seller of record, along with new Merchant Center attributes for things like answers to common product questions, compatible accessories and substitutes.26 Browser agents may read a site through screenshots, the DOM and the accessibility tree.7

My view is that for a shop the product feed is now a search asset in its own right, and it has to agree with the product page. Price, availability, variants, shipping and returns need to be accurate in the feed, in the structured data and in the visible text, because each surface may read a different one. An agent that finds a price mismatch or an out-of-stock variant will move on. The ecommerce SEO pillar covers feeds and product data in more depth. For services and SaaS, the equivalent is a clean, crawlable signup or booking path and accessible forms, because a browser agent works from the same page a person sees.

Measurement is still partial, but 2026 brought the first proper first-party reports. Each source answers a different question.

Google Search Console

Search Console now has a Generative AI performance report, rolled out to all sites on 31 August 2026. It shows impressions in AI Overviews and AI Mode, grouped by page, country, date and device, with a filter for text and image-based searches.27 It reports impressions; Google's report documentation does not list clicks or queries as metrics. AI feature traffic also stays inside the main Performance report under the Web search type.1 The counting rules matter when you read it: an AI Overview occupies one position and every link in it gets that position, a click on an external link in an AI Overview or AI Mode counts as a click, and a follow-up prompt in AI Mode is counted as a new query.28

Bing Webmaster Tools

Microsoft added an AI Performance report in February 2026, in public preview. It shows how often your pages are cited in AI answers across Copilot, Bing's AI summaries and select partner integrations, which pages are cited, and the grounding queries behind those citations.10 The grounding queries are the most useful part for anyone working on content, because they show the narrower searches an AI system ran. Check it even if Bing is a small share of your traffic.

Analytics referrals

Clicks from ChatGPT, Perplexity, Copilot, Claude and Gemini usually arrive as referrals from the assistant's domain, so a channel grouping for those referrers is the first thing to set up. It misses a lot. Clicks from AI Overviews and AI Mode arrive as ordinary Google organic traffic, some apps strip the referrer, and a recommendation that leads to a branded search or a phone call days later leaves no referral at all. Treat the referral figure as a floor.

Server logs

Logs are the only direct evidence of whether these crawlers reach your pages and what status codes they get. Filter by the user agents in the crawler table. User agent strings are easy to fake, so verify anything important against the IP lists the operators publish: Anthropic links to its list from its crawler page.5 A search crawler that only ever receives 403s or redirects is a problem you can fix this week.

Sampling prompts

Answers vary between runs, models, users and places. AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary,1 and OpenAI says location and saved memories can change the queries ChatGPT runs.2 That variation also differs by country and language, which matters for anyone running international SEO. A single check tells you very little. My own method, with no documented standard behind it, is to write a fixed set of prompts that mirror your buyers' questions, instructions and tasks, run each several times per assistant on a regular schedule, and record how often you are named or cited. Report the share across runs and watch the trend over months. I have written up how to build and sample the prompt set in more detail. Be wary of any tool that reports a single visibility score without publishing its prompts and sampling method; no third-party tool has access to Google's internal ranking or AI systems.7

Chapter 11 / 131 min

Common myths about AI search, checked against the sources

ClaimWhat the sources say
“You need special AI schema to appear in AI answers.”There is no special schema.org structured data to add, and structured data is not required for Google's generative AI features, though it is still worth using for rich results.7
“llms.txt is a ranking factor.”llms.txt is a proposal by Jeremy Howard for a Markdown summary file at /llms.txt, now in a second version.29 Google says Google Search ignores such files, and that having one neither helps nor harms visibility in Search.7 No other major provider has documented using it to choose sources, in the sources I checked.
“You can pay to be the recommended answer.”OpenAI says ChatGPT product results are selected independently and are not ads.23 Since it began testing ads in ChatGPT in 2026, it has said ads are separate and clearly labelled and do not influence the answers.30 Paid formats do exist on some AI surfaces: Google is piloting Direct Offers, discounts shown in AI Mode and labelled as a sponsored deal.26 Those are labelled ads. Neither provider documents a way to buy the organic answer or a citation.
“Blocking Google-Extended keeps you out of AI Overviews.”Google-Extended does not affect inclusion in Google Search.18 The documented levers are snippet controls and the Search generative AI control.22
“Chunk every page into short blocks for AI.”Chunking is on Google's list of things you can ignore; its systems find the relevant part of a page.7 Clear, self-contained sections still help readers and passage retrieval.
“Publish a page for every fan-out query.”Doing that primarily to manipulate rankings or AI responses violates Google's scaled content abuse policy.7 The GEO paper also found keyword stuffing gave little to no improvement in generated answers.14
“Get your brand mentioned everywhere, by any means.”Seeking inauthentic mentions is not as helpful as it seems, because Google's AI features rely on core ranking systems and spam systems.7
Claims seen in the industry and what the documentation says, October 2026
Chapter 12 / 131 min

Where the work sits for different kinds of business

These priorities are my judgement, based on the mechanics above. No provider documents them.

Business typeWork that matters mostWhat I would stop doing
SaaSPublic pricing with plan limits, integration pages that name each tool, crawlable docs, security and compliance pages, honest comparison pages.Hiding pricing behind “contact sales” where it is not bespoke. Shipping docs as a client-rendered app.
EcommerceAccurate feeds that match the page, specification tables in text, clear returns and delivery policies, availability by variant.Copying manufacturer descriptions used by every other retailer. Thin category pages with no facts.
TravelFares, timetables, inclusions and cancellation terms stated plainly, location detail on every property page, guides with first-hand detail.Generic destination listicles that restate what every other site says.
B2B servicesNamed certifications with scope, sector case studies, team and process pages, a clear statement of who you serve and who you do not.Vague capability pages that could describe any competitor.
Publishing and mediaOriginal reporting and analysis, clear bylines and dates, a deliberate decision on training crawlers and snippet controls.Rewriting other outlets' stories with nothing added, which gives an assistant no reason to cite you over the original.
Local servicesConsistent details across Business Profile and directories, service and area pages with specifics, reviews that describe the work.Mass-producing near-identical location pages.
Judgement: what matters most per business type, and what to stop doing

The SaaS SEO pillar covers pricing, docs and comparison pages in detail. For local businesses the work is existing local SEO done more strictly, and I have set out what AI search changes for a local business separately.

Across all of them, consistency is the thread. An answer assembled from several sources is safer for the system when those sources agree. My working assumption is that a product or business described the same way on its own site, its feeds and profiles, and on the review and comparison sites in its sector, is a lower-risk thing to name than one with conflicting details. That is my reasoning about how a careful system would behave, with no documentation behind it, and it is the same consistency that brand and local search have always rewarded.

This is the order I would follow on most sites. The early steps are cheap and catch most of the problems.

Checklist0 of 7 done
Questions9 answered

Common questions

01Do I need to optimise separately for Google AI Overviews and AI Mode?
Not according to Google. There are no additional requirements or special optimisations: a page needs to be indexed and eligible to show with a snippet,1 and the site must be included in Search generative AI features in Search Console, which is the default.7
02Does blocking Google-Extended remove my site from AI Overviews?
No. Google-Extended covers Gemini model training and grounding in Gemini Apps and Vertex AI, and it does not affect inclusion in Google Search.18 To limit AI Overviews and AI Mode, use snippet controls such as nosnippet and max-snippet,21 or the Search generative AI control in Search Console.22
03Can I block GPTBot and still appear in ChatGPT search?
According to OpenAI, yes. GPTBot is for model training and OAI-SearchBot is for ChatGPT search, and each robots.txt setting is independent. Sites that block OAI-SearchBot will not be shown in ChatGPT search answers.4
04What is query fan-out?
It is Google's term for an AI system splitting one prompt into several related searches run at the same time across subtopics, then combining the results into one answer.9 It means the searches that decide whether your page is retrieved are narrower than the prompt someone typed.
05Can I see AI Overviews and AI Mode data in Search Console?
Yes. The Generative AI performance report shows impressions in AI Overviews and AI Mode by page, country, date and device.27 Clicks from those features are counted in the main Performance report under the Web search type,1 where all links in an AI Overview share its single position and a click on one counts as a click.28
06Do AI crawlers see content loaded with JavaScript?
Mostly not, on the best public evidence. A December 2024 analysis by Vercel and MERJ found that the major AI crawlers it measured, including those from OpenAI, Anthropic and Perplexity, did not render JavaScript, while Gemini and Applebot did.20 Put anything you want quoted in the server-rendered HTML.
07Should I add an llms.txt file?
It is optional. llms.txt is a proposal for a Markdown summary at /llms.txt,29 and Google Search ignores such files.7 It makes most sense for developer documentation that coding tools and agents may fetch directly.
08How do product feeds affect shopping answers in ChatGPT?
OpenAI says ChatGPT considers structured product metadata such as price and description from first-party and third-party providers, ranks merchants on factors such as availability and price, and lets merchants apply to send a direct product feed.23
09Is AI-generated content penalised by Google?
Using AI tools is not against Google's guidance in itself, but content must meet its spam policies, and producing pages for every query variation primarily to manipulate rankings or AI responses violates the scaled content abuse policy.7
References30 sources

Sources

  • Platform docs21
  • Industry study5
  • Research3
  • Practitioner1
  1. 01AI features and your websiteGoogle Search CentralPlatform docs
  2. 02Searching the web with ChatGPTOpenAI Help CenterPlatform docs
  3. 03Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLewis et al., NeurIPS 2020 (arXiv)Research
  4. 04Overview of OpenAI CrawlersOpenAIPlatform docs
  5. 05Does Anthropic crawl data from the web, and how can site owners block the crawler?Claude Help CenterPlatform docs
  6. 06Perplexity CrawlersPerplexityPlatform docs
  7. 07Optimizing your website for generative AI features on Google SearchGoogle Search CentralPlatform docs
  8. 08Transparency Note for Microsoft Copilot (for individuals)Microsoft SupportPlatform docs
  9. 09Expanding AI Overviews and introducing AI ModeGoogle (March 2025)Platform docs
  10. 10Introducing AI Performance in Bing Webmaster Tools Public PreviewBing Webmaster Blog (February 2026)Platform docs
  11. 11Introducing the Perplexity Search APIPerplexity (September 2025)Platform docs
  12. 13Dense Passage Retrieval for Open-Domain Question AnsweringKarpukhin et al., EMNLP 2020 (arXiv)Research
  13. 14GEO: Generative Engine OptimizationAggarwal et al., KDD 2024 (arXiv)Research
  14. 15Google users are less likely to click on links when an AI summary appears in the resultsPew Research Center (July 2025)Industry study
  15. 16Update: AI Overviews Reduce Clicks by 58%Ahrefs (February 2026)Industry study
  16. 17AIO Impact on Google CTR: September 2025 UpdateSeer Interactive (November 2025)Industry study
  17. 18Google's common crawlersGoogle Crawling InfrastructurePlatform docs
  18. 19The crawl-to-click gap: Cloudflare data on AI bots, training, and referralsCloudflare (August 2025)Industry study
  19. 20The rise of the AI crawlerVercel and MERJ (December 2024)Industry study
  20. 21Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsGoogle Search CentralPlatform docs
  21. 22Search generative AI controlSearch Console HelpPlatform docs
  22. 23Shopping with ChatGPT SearchOpenAI Help CenterPlatform docs
  23. 24Buy it in ChatGPT: Instant Checkout and the Agentic Commerce ProtocolOpenAI (September 2025)Platform docs
  24. 25Let AI do the hard parts of your holiday shoppingGoogle (November 2025)Platform docs
  25. 26New tech and tools for retailers to succeed in an agentic shopping eraGoogle (January 2026)Platform docs
  26. 27Generative AI performance report (Search)Search Console HelpPlatform docs
  27. 28What are impressions, position, and clicks?Search Console HelpPlatform docs
  28. 29The /llms.txt file, v2llms-txt (Jeremy Howard)Practitioner
  29. 30Our approach to advertising and expanding access to ChatGPTOpenAI (January 2026)Platform docs
Going deeper on AI search

Shorter pieces on one part of this subject

Questions about any of this?

LinkedIn is the easiest place to reach me. Send a message or connect, and I will reply there.