AI search visibility measurement framework for technical teams
One panel gives you a snapshot. A report compares two periods, and that is where most AI visibility changes turn out to be artifacts of the panel, the sample size or an unlogged model swap. The design that makes a trend checkable.
When I measured Beetroot's visibility in AI search results across 359 answers, the same product came out at 0, 9, 13 or 100 percent depending on the surface and the prompt, and one prompt out of eight carried every non-branded mention. That piece is about taking one measurement correctly. This one is about what happens next, when someone runs the panel again a month later and puts the two numbers side by side.
My claim is narrow and, I think, uncomfortable for a lot of dashboards: most reported changes in AI search visibility are artifacts of how they were measured. The prompt panel changed, the sample was too small to see the difference, or the model underneath changed and nobody logged it. A visibility number that ships without its panel version, its sample count, and its model record cannot be checked by anyone, including the team that produced it.
What follows is a design for running that measurement as a recurring program, for the people who have to build the pipeline and defend the trend line: engineering on collection, analytics on definitions and statistics. The definitions of mention, citation and answer inclusion, and the one-day workflow, are in the companion piece; here they reappear only where tracking them over time adds a requirement. The framework has four objects, queries, mentions, citations, and competitors, plus the sampling rules that sit underneath all of them. The limits come near the end, and they are the part I would read first if someone handed me a report built this way.
Why can't you reuse your SEO dashboard for this?
Because the unit of observation changed. A search results page also varies with personalization and location, but it is stable enough that one check per query per day is a defensible measurement, and a rank tracker can store a position. An assistant's answer is generated fresh each time. The wording, the brand list, and the cited sources can all change between consecutive runs of the same prompt, and when the assistant searches the web first, the retrieval step adds its own variance.
So position becomes a rate. One check becomes many runs. "We rank third" becomes "we appeared in 7 of 20 runs, with a 95% interval of roughly 18 to 57 percent." Nobody wants that on a slide, and it is the only version of the number that survives a second run.
The other difference is the click. SEO reporting has impressions without sessions too, but most of its outcomes eventually land in analytics. A large share of assistant answers end with the user satisfied and gone, which is the dynamic I traced for publishers in The Quality Paradox. For measurement, that means you have to observe the answer itself instead of waiting for traffic it may never send.
How many runs do you actually need?
More than intuition suggests. And the question that matters is whether you can detect a change, which is harder than getting one interval. All intervals here are Wilson score intervals for a binomial proportion. For a single rate near 50 percent, 10 runs gives roughly 24 to 76 percent, 30 runs gives 33 to 67, and 100 runs gives 40 to 60.
But a report compares two periods, and the difference between two rates is noisier than either rate. For a two-sided test at the 5 percent level with 80 percent power, near a 50 percent baseline, the smallest change you can reliably detect is roughly:
| Runs per period | Detectable change |
|---|---|
| 30 | about 36 points |
| 100 | about 20 points |
| 400 | about 10 points |
| 1,000 | about 6 points |
Those numbers assume independent runs, and runs are not fully independent. Runs of the same prompt on the same day share a retrieval index, caches, and whatever state the model is in that week. Pooling every prompt and every run into one big binomial produces intervals that look precise and are too narrow. Three habits help here: spread runs across days within a period, treat the prompt as the unit when you compute panel-level uncertainty (a bootstrap that resamples prompts, not runs, is the simple version), and compare periods on the same prompts as paired data.
The companion panel shows how quickly this bites. Beetroot's mention rate on the web-search surface was 7 of 80 non-branded answers, about 9 percent. Near that baseline, 80 runs per period can reliably detect a change of roughly 13 points in either direction. A doubling to 18 percent next month would most likely not register as significant, and neither would a halving to 4 percent. The single snapshot was informative. A month-over-month comparison at the same size would mostly measure noise.
Then do the budget arithmetic before the design is fixed. A panel of 100 prompts, 20 runs each, on 3 surfaces is 6,000 collected answers per period, before any accuracy grading. At that size you can see panel-level shifts of several points, but no single prompt carries enough runs to detect a 10-point move. That trade-off is the real design decision: fewer prompts with more runs, fewer surfaces, or accepting that per-prompt numbers are diagnostic only. Write the answer down before the first report, as "a change smaller than X is noise at our sample sizes."
One more trap. With 100 prompts tested at the 5 percent level, about five will show a "significant" move every period by chance alone, and those five are exactly the ones that end up in the deck. Per-prompt movement is for finding things to look at, not for reporting.
Layer 1: which queries should you measure?
A fixed, versioned panel of prompts, stratified by intent, where every prompt records how it was made. The panel is the denominator of every metric that follows, so it is the most consequential choice in the framework and the easiest one to bias.
Build three strata and report them separately:
- Unbranded discovery prompts. Start from real queries (Search Console, keyword tools) and keep the original query next to the assistant-style version you derive from it. Rewriting changes intent easily: "best open source clipboard manager for Windows" and "a free clipboard manager that keeps images" are different requests, and the second silently drops the open-source constraint. Label every derived prompt as synthetic and store the transformation.
- Comparison prompts. "Alternatives to X", "X vs Y for a team of 20". Balance them across competitors, so your brand is not the named anchor in most of them.
- Entity prompts. "What is [product]?", "Who makes [product]?". These name the brand in the question, so a mention is nearly guaranteed. They measure whether the description is correct, not whether you are recommended, and they must never be blended into a discovery score.
Treat the panel like code: an ID, a version, a freeze for each measurement period. Adding or retiring prompts bumps the version, and numbers across versions are only compared after restating both panels. A score that rose because someone added five branded prompts looks exactly like a win. The composition effect does not even need branded prompts. In the companion run, one prompt about AI features produced 17 of Beetroot's 17 non-branded mentions across the two search surfaces; adding a second prompt like it would have nearly doubled the headline rate with nothing changing in the world.
Layer 2: what counts as a mention?
Your entity named in the generated answer text, detected by a matcher you can audit and validate. Raw answers abbreviate names, misspell them, describe "the Rust-based one" without naming it, or attach your name to a competitor's feature.
Record per answer:
- mentioned: named via exact match, a known alias, or a reviewed fuzzy match, from a versioned alias dictionary kept next to the panel;
- stance: neutral mention, recommendation, or explicit rejection. "X is an option, but most users prefer Y" is a mention and a loss;
- position: ordinal in the answer's list, when the answer is a list;
- accuracy: whether what the answer says about you is true, graded against a dated fact sheet (license, platforms, pricing model, core features), with "uncertain" as an allowed outcome. An answer that calls your open-source tool proprietary is a liability that a naive dashboard counts as a win.
Validate the matcher against a human-labeled sample that includes answers the matcher marked as not mentioning you, because reviewing only the detected mentions will never reveal the misses. If a model does the grading, measure its agreement with the human labels and version the grader and the rubric.
The companion run found both failure modes a matcher has to survive: an answer that folded Beetroot into a catch-all line of "similar indie apps", which a regex scores as a full mention, and a model without web access that confidently gave the wrong license and linked to unrelated repositories. Over time, the matcher and the grader also become versioned inputs. Change either one and old answers have to be re-scored before the trend line means anything.
The core metric is the mention rate: the share of completed runs of a prompt that mention you. Aggregate across the panel with fixed weights over prompts (equal weights are a reasonable default), not by pooling runs, or a surface that timed out more often this month quietly reweights your score.
Layer 3: what counts as a citation?
A link or source attribution to one of your URLs attached to the answer. It is a separate object from a mention, and the two can diverge in both directions.
You can be mentioned without being cited, and you can be cited without being mentioned: your comparison page or documentation used as a source for an answer that recommends someone else. The first case is tempting to over-read. An uncited mention may come from the model's training data, but it may also come from retrieved third-party pages the interface did not display. Retrieval sources and displayed citations are not the same set; OpenAI's own web search documentation distinguishes them. Treat the pattern as a hypothesis to test, not a diagnosis.
The citation rate is also the metric most exposed to plumbing. In the companion run, counting inline Markdown links as well as structured annotations moved the rate 2.5x on the same answers, and one surface's citations never reached the collector at all. For a trend line, that means the parsing rule has a version like everything else, and a silent parser change is indistinguishable from a real shift.
Record every cited URL raw and normalized (redirects resolved, tracking parameters stripped), its owner class (yours, a competitor's, a third party's) from a versioned ownership list, and whether the citation is inline or in a separate source panel. Keep structured API responses or rendered evidence where you can, because a flat list of URLs loses which claim a citation was attached to. From that record:
- citation rate: share of completed runs citing at least one of your URLs;
- cited pages: which of your URLs get cited, the only part of this framework that maps directly onto pages your team can change;
- third-party citation share: the share of observed citations pointing to review sites, forums, and directories rather than to you or a tracked competitor. It tells you who the displayed sources are, not how much each one contributed to the text.
Server logs add an independent view, with two classes of bots to keep apart. Search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) index pages for later use. User-triggered fetchers (ChatGPT-User, Perplexity-User, Claude-User) retrieve a page because someone asked something right then. The second class is closer to evidence of live use; neither proves a citation. Google's AI Overviews and AI Mode draw on Google's regular index, so logs show nothing specific to them. Verify crawler identity against published IP ranges or reverse DNS before trusting a user agent string.
Layer 4: how do you measure competitors fairly?
Same panel, same runs, same matcher, and share of voice rather than your own rate in isolation. Your mention rate is informative on its own, but it reads very differently next to a category leader at 35 percent than next to one at 90.
Define share of voice on presence: for each answer, each tracked brand counts once if it appears, however many times it is named. Share of voice for a prompt is your presence count divided by the total presence count of all tracked brands. When no tracked brand appears, the value is undefined, not zero. Report those answers in two separate buckets: answers that name no brand at all, and answers that name only brands you do not track. A growing second bucket suggests your competitor set is out of date.
Compute it per intent before any category total. In the companion run, Ditto and CopyQ dominated most prompts, while on the AI-features prompt one surface named Microsoft's PowerToys Advanced Paste in all 10 answers, with Beetroot second each time, and neither Ditto nor CopyQ appeared at all. A category-wide share would average two different competitive situations into one that exists nowhere. Citations belong here too: in comparison prompts, it is worth tracking whose page is cited when the answer frames the choice.
Fix the competitor set together with the panel. Adding a strong competitor mid-period lowers everyone's share and looks like a decline.
What goes into the run record?
Every run, including the ones that failed. The companion piece lists the minimum record for a single measurement; a recurring program needs more, because a trend has to survive changes to the model, the collector and the definitions. The schema separates what you set from what you observed, because on most consumer surfaces you cannot set the model or force retrieval:
# set by the collector
run_id, panel_version, prompt_id, surface, requested_mode,
country, language, account_state, fresh_conversation, timestamp
# observed in the response
status (completed | refused | timeout | parse_error | no_ai_feature),
retry_of, model_reported, search_triggered, issued_queries[],
answer_text, citations[] (raw_url, normalized_url, placement)
# applied later, re-runnable
matcher_version, grader_version, ownership_list_versionstatus is what keeps the denominator honest. A timeout is not an answer without a mention, and for Google, "no AI Overview shown for this query" is its own outcome: report visibility both across all eligible queries and conditional on the feature appearing. Publish completion rates next to every metric.
model_reported is often empty. APIs expose model identifiers; consumer products mostly do not, and vendors change models behind an unchanged product name. On those surfaces a model change can only be inferred, from release notes or from an abrupt shift across many prompts at once. Record what you can, mark the rest as unknown, and annotate the time series with every known vendor release. I made the same argument for agent pipelines in A model is a dependency that won't hold still: you cannot separate the vendor's drift from your own changes unless you logged what you were measuring against.
API results and consumer interfaces are different surfaces, not proxies for each other; keep their series apart. If you collect from consumer interfaces through browser automation, read the product's terms first. Automated access to a consumer app can violate them and put the account at risk.
What can these numbers not tell you?
Put these in the body of every report built this way. A limit in a footnote is a limit nobody reads.
It is not the population of answers. The panel is a designed sample. Real users phrase things their own way, carry conversation history and memory, and see personalized results. The panel measures a controlled condition, which is what makes it comparable over time, and why it is not a traffic estimate. More runs reduce sampling noise; they do nothing about a biased panel.
Vendor tools are measurements too. Commercial trackers differ: some let you define your own prompts, others report on their own prompt databases and estimate reach by weighting prompts with search volume adjusted for each platform's usage. Record which tool, which mode, and which methodology version, and label its series clearly. Its numbers are not interchangeable with yours, even when the metric has the same name.
Referrals are observed, not impact. ChatGPT search adds utm_source=chatgpt.com to referral links (OpenAI publisher FAQ), and some assistants pass a referrer. Search Console has had a separate generative AI performance report since this summer, but it covers impressions for AI Overviews and AI Mode together, and clicks still land in the ordinary web totals. Call these "observed attributable sessions." They are neither a lower bound on the business effect nor a conversion rate for mentions.
A trend is not a cause. A stable panel shows that visibility moved. It does not show that the schema markup or the new comparison page moved it. Testing a specific change needs its own design: a treated set of pages or prompts, a comparable untreated set, and an observation window fixed in advance.
Implementation checklist for engineering and analytics teams
The body explains each item. This list is what has to exist, and who owns it, before the first number leaves the team.
Engineering
- Prompt panel, alias dictionary, ownership list, and fact sheet in version control, each with a version ID.
- One collector per surface (API where one exists, browser automation only where the terms allow), running every prompt N times per period, spread across days, from fixed regions and account states.
- The full run record for every attempt, including failures, raw answer text, and structured citations.
- Log parsing for search crawlers and user-triggered fetchers by URL, with identity verification.
- Reprocessing: any matcher, grader, or ownership change can be re-run against stored raw answers.
Analytics
- Panel built in three strata (discovery, comparison, entity), each derived prompt linked to its source query, frozen per period.
- Sample size chosen from the detectable-change table and the collection budget, and the resulting noise threshold written down.
- Matcher and grader validated against human labels, including negatives.
- Metrics reported per stratum and surface: mention rate, stance, accuracy, citation rate, cited pages, third-party citation share, share of voice, each with completion rate, run count, and interval.
- Time series annotated with panel versions, known model releases, and your own site releases.
- Referral data labeled as observed attributable sessions, never as the outcome metric.
Shared
- One page that defines every metric in plain words. When the numbers are challenged, that page is what you defend.
- A quarterly panel review: retire prompts nobody asks anymore, add new demand, bump the version, restate the comparison.
Where this fits
A framework is a method, not a finding. The single-panel measurement is the first data point; the findings come from running something like this framework long enough to see which of the popular recommendations for "getting into AI answers" hold up, and then testing the ones that do with a proper intervention design. The research section holds the original-data work on this site, each piece with its method and sources stated next to the result. That is the standard a visibility number should meet too: if nobody else can check how it was produced, it is an opinion with a chart.