Most AI search measurement I see is one person typing a question into ChatGPT, seeing whether the brand comes up, and taking a screenshot. That tells you what one model said to one account in one place at one moment. Run it again and the answer may change. The first-party reports have arrived, but each covers one provider and one kind of event, so a team can easily read a rise in one number as a win everywhere. A method that holds up needs a fixed prompt set, sampled repeatedly, with citations recorded and joined to Search Console, Bing Webmaster Tools and analytics.
The mechanics behind why answers vary, from query fan-out to the crawlers involved, are covered in the AI search guide.
What each source can and cannot tell you
Start by being clear about what each source measures. They are not interchangeable, and most reporting mistakes come from treating them as if they were.
| Source | What it measures | What it misses |
|---|---|---|
| Your own prompt sampling | Whether you are named or cited for prompts you chose, in any assistant you test | Only covers the prompts you wrote. Real users ask things you did not think of. |
| Search Console, Generative AI performance report | Impressions of links to your site in AI Overviews and AI Mode, by page, country, date and device1 | Clicks and queries are not listed as metrics. ChatGPT, Copilot, Claude and Perplexity are outside it. |
| Search Console, main Performance report | Clicks from AI features, counted under the Web search type alongside ordinary results2 | AI feature clicks are mixed in with classic results, so you cannot isolate them there. |
| Bing Webmaster Tools, AI Performance report | Citations in Copilot, Bing AI summaries and some partner integrations, cited pages, and a sample of grounding queries3 | Clicks and placement within the answer3 |
| Analytics referrals | Visits that arrive with an assistant as the referrer or a tagged source | AI Overview and AI Mode clicks look like Google organic. Visits that lose the referrer, and buyers who come back later by another route. |
Build a prompt set from questions, instructions and tasks
Your prompt set is the only part of this you design, so it decides how useful everything else is. Write it the way your buyers prompt, which is not the way they used to search. People ask questions, give instructions and hand over tasks, and each kind pulls in different pages. A set made only of short questions will tell you about your blog and almost nothing about your pricing, product data or policies.
Turn keywords into needs, and needs into stages
A list of prompts hides its own gaps. A better structure starts from what customers want. Group keywords that ask the same thing into one need, so you never pay to ask the same question three ways. Then split every need by how close the customer is to choosing: exploring, comparing or ready to buy. Laid out that way, an empty slice shows at a glance what nobody is measuring yet. This is how the prompt research in SearchOps, the platform I built, organises a business's prompts.4
A prompt coverage map for a local vet
Vaccinations3 tracked · 1 suggested
Ready to buy
- “Book puppy vaccinations near Clifton this week”
Comparing
- “Cheapest vet for cat vaccinations in Bristol”
- “Which vets in Bristol offer a vaccination plan?”
Exploring
- “What vaccinations does a puppy need in the UK?”
The structure works for any business. For a SaaS product the needs are the jobs buyers hire it for, such as rota planning or payroll export. For a shop they are product families and the problems behind them. The three stages stay the same, because every buyer moves from exploring to comparing to ready to buy.
Where good prompts come from
The best prompts come from evidence of what customers actually ask, not from a brainstorm. SearchOps suggests prompts from Search Console queries, the searches that find a business on Maps, its reviews, the questions Google shows, and the searches answer engines ran about it, and it accepts call notes and enquiry forms as input.4 Whatever tool you use, or none, those sources apply:
- Search Console queries, especially long ones, which read like prompts already.
- The searches that find you on Maps, for anything local.
- Your reviews, which name the service and the situation in the customer's own words.
- The questions Google shows in its results for your main terms.
- The searches an assistant ran while answering. ChatGPT discloses its searches, so a tracker can collect them.5
- Call notes, enquiry forms and sales calls, the richest and least used source of all.
Keep three states for every prompt. Tracked prompts are measured on a schedule. Suggested prompts wait for a person to approve them. Saved prompts are kept for later and cost nothing to hold.4 The discipline matters more than the tool: a prompt gets tracked because it maps to a need and a stage, not because someone thought of it in a meeting.
| Business | Question | Instruction | Task |
|---|---|---|---|
| SaaS: rota and scheduling software for care homes | “What should a care home look for in staff scheduling software?” | “Compare three rota tools for a 60-bed care home in a table with price per staff member and whether they handle agency staff.” | “Find a scheduling tool with a free trial that integrates with our payroll provider and start the signup.” |
| Ecommerce: outdoor clothing retailer | “Are synthetic insulated jackets better than down in wet weather?” | “List five insulated jackets under £200 that pack small, with weight and fill type.” | “Find a medium synthetic jacket under £180 in stock with free returns and add it to my basket.” |
| B2B: industrial cleaning contractor | “How often should a food factory have a deep clean?” | “Shortlist contractors in the West Midlands who do overnight food factory cleaning and hold BRCGS experience. Say which publish prices.” | “Request quotes from three of those contractors for a monthly deep clean.” |
These are my working rules for the set. Nobody has published a standard for this yet.
- Keep most prompts unbranded. A prompt that names you measures whether the assistant can describe you, which is worth knowing, but it says nothing about whether you are recommended. Keep the two groups separate in the reporting.
- Write them long. Real prompts carry constraints such as size, budget, location and integrations, and those constraints change which sources are retrieved.
- Avoid leading prompts. “Is our brand the best option for X?” will usually get you named.
- Cover the buying journey, from early questions to shortlist instructions to tasks. The task prompts often show whether an agent can use your site at all.
- Include location where it matters. For anything local or delivered, write the same prompt for two or three towns.
- Fix the set. Add prompts when you find new ones, but never edit an existing prompt in place, or the trend line breaks.
Size depends on how many products and markets you have. I would start with 30 to 50 prompts for a single-product business and resist going much higher until the process runs smoothly, because each prompt multiplies across assistants and repeat runs.
Sample repeatedly and control the conditions
An assistant does not return a fixed result for a prompt. AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary.2 OpenAI says ChatGPT may infer a general location from the IP address, and may use saved memories when it rewrites a prompt into search queries.6 So the same prompt can produce different shortlists for two colleagues, or for one person on two days.
The answer is to treat every result as a sample and control what you can. This is the routine I would follow.
Personalisation: measure the customer asking cold
Assistants personalise. A signed-in person's history, memories and location shape what they see, and nobody outside can reproduce that answer exactly.5 Chasing it is a trap. The case worth measuring is the customer asking cold, with nothing to personalise on, because that is the answer you can repeat and the one a new customer most likely gets.
- Put the personalisation into the prompt, where it is visible and repeatable. Write the town or city into every local prompt, as SearchOps does,4 and write the buyer's constraints into the rest: budget, team size, the integration they need.
- Treat a persona as a set of constraints, not an account. A 60-bed care home manager and a two-site franchise owner write different prompts, so give each persona its own prompts in the same needs map.
- Report each engine separately. SearchOps runs every prompt on ChatGPT and Gemini and never blends the two,5 because they draw on different sources and an average of the two describes neither.
Record citations, not just mentions
Being named and being cited are different events. An assistant can recommend you while citing a review site, a directory or a competitor's comparison page as its source. It can also cite your page for a fact without recommending you. So record every answer as structured data, the same fields every time, and keep the full text so you can re-read it later.
One answer, recorded as data
PromptWhich vets near Clifton, Bristol are good with nervous dogs? I need one open on Saturdays.
Here are three options near Clifton that are known for handling anxious dogs and open on Saturdays. Harbourside Animal Clinic offers longer, quieter appointments and has a Saturday morning surgery. Elmbrook Vets is mentioned in several reviews for patient handling of nervous pets. Downs Veterinary Centre also opens on Saturdays. Sources: a local pet directory, a review site.
- You
- named, position 2 of 3, no stated reason
- Businesses named
- Harbourside Animal Clinic, Elmbrook Vets, Downs Veterinary Centre
- Named with a reason
- Harbourside: yes · Elmbrook: no · Downs: no
- Cited sources, by type
- directory (cites rivals), review site
- Searches the engine ran
- vets clifton bristol nervous dogs; vets open saturday bristol
- Outcome
- Mentioned (shortlisted would need a reason)
- Since last check
- new: absent in the previous check
- Date, assistant, model or mode, logged-in state, location and run number.
- Whether you were named, and whether the answer gave a reason for naming you. Being named with a reason is a stronger result than a bare listing.
- Every cited URL, marked as yours, a competitor's or a third party's, and the kind of source: directory, review site, publication, official body.
- The searches the assistant ran, where it discloses them.
- How the business was matched. For local businesses, match on website, phone number or name and address, not the name alone, which is often ambiguous.
- An outcome label, such as shortlisted, mentioned, cited only or absent, and how it compares with the previous check: new, consistent or lost.
- Which competitors were named.
- Any wrong fact about you: an old price, a discontinued product, a service you do not offer. For prompts that name you, also record the verdict: favourable, mixed, neutral or critical. Grounding does not prevent this: Microsoft's transparency note for Copilot warns that a response might misrepresent the content it was grounded in.7
The third-party citations are where the work is. If a care home software prompt keeps citing the same independent comparison article, the questions become whether you are in it and whether it describes you accurately. If an outdoor retailer's jacket prompts cite a rival's size guide, the gap may be a missing size guide on your own site. The mentions tell you how you are doing. The citations tell you what to do next. The most useful single view is a list of sources that cite your competitors but never you, grouped by type.4
An assistant writes a new answer every time, in a new order, so a single "you are number 2" is not a number anyone can defend.4 The figures that hold up are pooled across every response in a period:
| Metric | What it means | How to calculate it |
|---|---|---|
| Visibility rate | How often the answers include you | Responses that name you, divided by all responses, pooled over every prompt and check in the period |
| Share of voice | Your share of all the businesses named | Times you were named, divided by all business mentions in those responses |
| Named with reasons | How often the assistant explains why it chose you | Responses that give a reason for naming you, divided by responses that name you |
| Consistency | Whether you are named reliably or occasionally | The share of checks in which a prompt names you, per prompt |
| Engine agreement | Whether ChatGPT and Gemini agree about you | Prompts where both name you, or neither does, divided by all prompts |
Pool the counts; never average averages. Five prompts at 100% and one at 0% is not 83% if the zero prompt ran four times as often. Read the trend over 30 and 90 days, and compare the first half of a window with the second to see the direction, which is how SearchOps reports movement.4 A rise inside one week is usually noise.
Join it up with Search Console, Bing and analytics
Search Console
The Generative AI performance report counts impressions, defined as how many times links to your site were shown in a generative AI feature on Google Search, and groups them by the final URL after redirects.1 Standard Performance report limits such as the 1,000 row limit apply, and the newest data is preliminary.1 Compare its top pages with the URLs your prompt set sees cited in AI Mode. Where they disagree, your prompts are probably missing a kind of query that real users run.
When you read clicks in the main Performance report, remember the counting rules. All links in an AI Overview share its single position, a click on an external link in an AI Overview or AI Mode counts as a click, and a follow-up question is treated as a new query.8 Position figures for pages that appear mostly in AI Overviews will look better than the experience feels. Do not read AI Overview impressions as a stand-in for clicks either. Pew Research Center tracked the Google searches of 900 US adults in March 2025 and found that people clicked a link inside an AI summary in 1% of visits to pages that showed one, and clicked a traditional result in 8% of those visits, against 15% on pages without a summary.9 Seer Interactive, an SEO agency, found that across 3,119 informational queries for 42 client organisations, brands cited in an AI Overview had a 35% higher organic click-through rate than brands that were not, which is a correlation and supports recording citations separately from mentions.10
Bing Webmaster Tools
The AI Performance report shows total citations, average cited pages, page-level citation activity and grounding queries, the phrases the AI used when it retrieved content it went on to cite.3 Microsoft says the grounding query data is a sample.3 Use those queries to improve your prompt set. A B2B contractor might find grounding queries about accreditation or site induction rules that nobody on the team would have written as a prompt.
Analytics
OpenAI says ChatGPT adds utm_source=chatgpt.com to referral URLs, so those visits can be tracked in analytics.11 For other assistants, check your referral report for their domains and confirm what actually arrives before you write any rules. In GA4, a custom channel group can collect these sources into one channel using conditions such as “source matches regex”. Traffic goes to the first channel it matches, so place the AI channel above Referral, and custom groups apply to reports retroactively.12
Treat this channel as a floor. A SaaS buyer who sees you in a comparison table and comes back a week later through a branded search shows up as organic or direct, with no trace of the assistant. For that reason I would also watch branded search impressions in Search Console and add an “assistant” option to “how did you hear about us” fields on demo and quote forms.
Where a tracking tool fits, and what to ask of one
A spreadsheet works until the prompt count, the engines and the schedule multiply. At that point a tool earns its keep, provided it shows its method. Google says no third-party tool has access to its internal ranking or AI systems,13 so every tracker is sampling, as you are. Ask any vendor to show you the prompts, the run counts, the locations and how they count.
SearchOps automates the method on this page for local businesses. It runs every tracked prompt on ChatGPT and Gemini on a weekly, fortnightly or monthly schedule and reports the engines separately.4 It stores the full response, the businesses named, the searches ChatGPT ran and the citation sources, matches your business by website, phone number or name and address, and compares each check with the last.4 Its reports group cited sources by type and flag the ones that cite rivals but never you, score the verdict on branded prompts,4 and show AI Overview and AI Mode citations beside the assistant results.5 You can see the details on the ChatGPT Visibility page.
Outside local, the same method still applies. A SaaS team or an ecommerce manager can run it in a spreadsheet, or ask any tool they consider to show the same evidence.
Mistakes that make the numbers lie
- Reporting a single run as a position. One answer is an anecdote.
- Testing from your own logged-in account, where memory and history pull the answer towards things you have asked before.
- Adding branded prompts to the headline figure, which inflates it with prompts you were always going to win.
- Comparing Search Console impressions with Bing citations as if they measured the same event. One counts links shown; the other counts sources cited.
- Reading Bing's grounding queries as a complete list. Microsoft describes them as a sample.3
- Buying a single visibility score without knowing the prompts and sampling behind it. Every tool is sampling, as you are. Ask to see the prompts.
- Changing the prompt set and the site in the same month, then crediting the site for the difference.
None of this needs special software to start. A needs map, a fixed prompt list, a run log and a monthly summary will tell you more than a dashboard built on prompts you cannot see. Once the routine works, automate the running and keep the method.
Common questions
01Can Search Console show which queries trigger AI Overviews for my site?
02How do I track traffic from ChatGPT in Google Analytics?
03How many times should I run each prompt?
04Does the Bing AI Performance report show clicks?
05Can I track what ChatGPT shows a signed-in customer?
06Are AI visibility tracking tools accurate?
Sources
- Platform docs11
- Industry study2
- 01Generative AI performance report (Search)Search Console HelpPlatform docs
- 02AI features and your websiteGoogle Search CentralPlatform docs
- 03Introducing AI Performance in Bing Webmaster Tools Public PreviewBing Webmaster Blog (February 2026)Platform docs
- 04ChatGPT VisibilitySearchOps (the author's company)Platform docs
- 05Gemini VisibilitySearchOps (the author's company)Platform docs
- 06Searching the web with ChatGPTOpenAI Help CenterPlatform docs
- 07Transparency Note for Microsoft Copilot (for individuals)Microsoft SupportPlatform docs
- 08What are impressions, position, and clicks?Search Console HelpPlatform docs
- 09Google users are less likely to click on links when an AI summary appears in the resultsPew Research Center (July 2025)Industry study
- 10AIO Impact on Google CTR: September 2025 UpdateSeer Interactive (November 2025)Industry study
- 11Publishers and Developers - FAQOpenAI Help CenterPlatform docs
- 12Custom channel groupsAnalytics HelpPlatform docs
- 13Optimizing your website for generative AI features on Google SearchGoogle Search CentralPlatform docs
Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the AI search guide.