How to measure brand visibility in AI search results
Mention, answer inclusion and citation come apart the moment you sample an assistant more than once. One prompt can carry a brand's entire presence, a parser can move its citation rate 2.5x, and a hallucinated license still counts as a mention. A workflow for technical teams, built on 359 answers.
On September 28 I ran Beetroot, the Windows clipboard manager I build, through ten prompts on four AI surfaces. Depending on which figure you pick, its visibility was 0 percent (any model without web access), 9 or 13 percent (the two search surfaces, averaged over the panel) or 100 percent (Sonar, on the one prompt about AI features). All four are correct. That is the first thing to understand about how to measure brand visibility in AI search results: none of those numbers answers "how visible are we?"
A visibility number is only meaningful next to the surface it came from, the prompts it was averaged over, the number of samples behind it, and the rule that decided what counted. Drop any of those and you are looking at a random variable dressed as a KPI.
Four terms get used as synonyms in AI visibility reports: visibility, mention, citation, answer inclusion. The data below pulls them apart, so I define each one first. After that comes a workflow a technical team can run, including the mistakes my own small run caught me making.
What do visibility, mention, citation and answer inclusion actually mean?
Three of them are outcomes you record per answer, and the fourth is the summary you build from them. Collapse them into one number and you can no longer tell whether the brand was named, recommended or credited.
Mention. The brand's name appears in the generated answer text. It is a string match (exact name plus a reviewed alias list), counted per answer. It says nothing about context: "Beetroot and similar indie apps" is a mention, "avoid Beetroot" is a mention, and so is any answer to a prompt that already contained the brand's name.
Answer inclusion. The brand is part of what the answer actually recommends for the user's need: its own list item, its own paragraph, a named pick. This is a judgment, not a string match, and it usually needs a grader (a person, or a second model that you audit). Record the position too. First of five and "also worth a look" at the bottom are different outcomes.
Citation. A URL you own is attached to the answer as a source: an annotation, a footnote, an inline link in an answer built from retrieval. A link a model writes from memory, with no retrieval behind it, is a different field (more on that below). Citation is independent of mention. A model can name you from memory with no source, and it can cite your comparison page while recommending a competitor.
Visibility. The umbrella summary: the rate at which a brand shows up in a declared set of answers, meaning a fixed prompt panel on a fixed surface over a fixed period. It is always a vector (mention rate, inclusion rate, citation rate, each with its sample count), never a single score, and it never exists without the panel it was computed over.
There is a fifth object upstream of all four, which people rarely measure because most tools hide it: retrieval, the pages the system fetched before writing. Retrieved is not cited, and the gap between them turned out to be large.
What did one small panel show?
I ran ten prompts through four surfaces via one API gateway on September 28, 2026. Eight prompts were non-branded, the kind a buyer asks ("What is the best clipboard manager for Windows?", "What are good alternatives to Ditto clipboard manager?", "Recommend a clipboard manager for Windows that has built-in AI features, like rewriting or translating copied text"). Two were entity prompts that name the product ("Is the Beetroot clipboard manager free and open source? What license does it use?"). Each non-branded prompt ran 10 times per surface, each entity prompt 5 times. 359 answers came back (one request hit a rate limit), and the whole run cost under three dollars in gateway credits.
Here are the non-branded results, 80 answers per surface (79 for Sonar). One caveat for the web-search row: in 9 of its 80 answers the model chose not to search at all, so "search available" and "search used" are not the same condition.
| Surface | Mention rate | 95% interval | Own-URL citation |
|---|---|---|---|
| gpt-5-mini, no tools | 0 / 80 | 0–5% | 0 |
| Gemini 2.5 Flash, no tools | 0 / 80 | 0–5% | 0 |
| gpt-5-mini, web search | 7 / 80 | 4–17% | 5 answers (2 by annotations only) |
| Perplexity Sonar | 10 / 79 | 7–22% | not observable (see below) |
The no-tools rows are not a surprise once you look at the dates. Beetroot went public in February 2026. OpenAI lists gpt-5-mini's knowledge cutoff as May 31, 2024, and in one run the model said so itself: no such product "in my training up through mid-2024". A product younger than a model's training data can still be mentioned from memory (the entity section below shows how), but it cannot be known. For Beetroot, on these two models, retrieval surfaces were the only place it appeared at all, which is why I would measure them first for any young brand.
Why did a single prompt carry all the visibility?
Because every mention came from the one prompt that matched what the product is actually different at. The AI-features prompt produced 7 of 10 mentions with web search and 10 of 10 on Sonar. The other seven non-branded prompts produced 0 of 70 with web search and 0 of 69 on Sonar.
So the panel-level "9 percent" and "13 percent" hide the only structure in the data. An average over a panel is a probability only under that panel's mix of questions, and nothing says real users ask in that mix. What the answers show is one intent where Beetroot appeared in 17 of 20 runs and seven where it appeared in none of 139. The average blends them into a number neither group of answers looks like.
Two consequences for measurement. First, report per prompt (or per intent cluster) before you report any panel average. Second, the panel's composition is the metric: add one more AI-feature prompt that behaves like this one and the headline rate nearly doubles; add three and it nearly triples, without anything changing in the world. That is why the panel has to be versioned and frozen like code. The same scrutiny applies to any commercial tracker you use: its panel is someone else's composition decision.
The competitor view makes the same point from the other side. On the AI-features prompt, Sonar named Microsoft's PowerToys Advanced Paste in all 10 answers, and Beetroot came second in every one of them. Ditto and CopyQ, which dominate every other prompt, did not appear there at all on Sonar. A category-wide share of voice would average this away; computed per intent, it shows who owns which question.
Why is a mention not the same as answer inclusion?
Because the string match counts answers the brand did not really win. Across the seven web-search answers that mentioned Beetroot, six gave it its own recommendation slot (at positions from first to fourth), and one folded it into a catch-all line: "PastePaw (and similar indie apps: Beetroot, Klip, Clipboard Genie, PastePaw)". A regex scores that as a mention. A reader would not call it a recommendation.
Sonar was the opposite case: 10 of 10 inclusions, always as a named pick, never first. Its phrasing barely moved across runs ("a good alternative", "a strong alternative", "a solid second choice"). A consistent second place is a different finding from an unstable first place, and a bare mention rate cannot tell them apart. Record inclusion and position separately from the mention.
How much does your parser decide the citation rate?
Enough to move the rate 2.5x on the same answers. Three findings, all from plumbing rather than from the models.
Annotations versus inline links. With web search, the API returns citations as structured annotations, but some answers also carry sources as plain Markdown links in the text. Counting annotations only, 2 of the 7 mentioning answers cited a Beetroot URL. Counting inline links as well, 5 did. Same answers, same day, a 2.5x difference decided by one line of parsing code. Pick a definition, write it down, and keep the raw text so you can re-score old answers when you change it.
Citations your pipeline silently drops. Sonar's answers contained footnote markers like [2][6] next to every Beetroot recommendation, and my collector recorded no source URLs for any of them. I had not kept the raw responses, so afterwards I sent one more request through the same gateway route and saved the complete body: footnote markers in the text, not a single URL anywhere in the response. Perplexity's own API returns citations; this route did not carry them. The answer is visibly cited, but the citation rate for that surface is unobservable in my data. A collector can lose an entire metric without raising a single error, so save a few complete raw responses per surface and check them for the fields you think you are storing.
Retrieved is not cited. Counted once per answer, the web-search runs retrieved 2,086 URLs across 80 non-branded answers and cited 385 of them (annotations plus inline links), about 18 percent. Reddit threads were retrieved in 36 of those 80 answers and cited in 12. So "the model read your page" and "the model credited your page" are separate events. In one answer Beetroot's own page was retrieved and the brand still ended up in the catch-all line. In another Beetroot was recommended with no Beetroot URL retrieved at all, carried by a forum thread about it on community.openai.com. If you only log citations, both of those cases are invisible.
One more trap: the no-tools models also produced URLs, in roughly 3 of every 10 answers and often several at once, recalled from memory rather than retrieved. One pointed to clipjump.codeplex.com, a host on a service Microsoft shut down years ago. A URL in a memory-only answer is a claim about a link, not a citation, and it belongs in a different column.
What changes when the prompt already names the brand?
The mention rate becomes meaningless (the name is in the prompt), and what matters is whether the answer is right, wrong or declines to answer. This is where the four surfaces split most sharply.
Asked for Beetroot's license, both web surfaces answered Apache 2.0 in 5 of 5 runs, which is correct, and asked who makes it, both named the right maker in 5 of 5. gpt-5-mini without tools declined every time: on the license prompt it asked which Beetroot I meant, on the maker prompt it said it did not recognize the product. That is not a correct answer, but it is the safer failure. Gemini 2.5 Flash without tools said MIT in 4 of 5 license runs and, in 3 of them, linked to GitHub repositories under accounts that have nothing to do with the product, a different one each time. Asked who makes it, it said no such product exists in 3 of 5 runs and named a wrong developer in the other 2.
A dashboard that counts "brand mentioned" scores that Gemini row as full visibility. It is a confident falsehood with a wrong source attached, and it is the answer a buyer checking your license would read. On entity prompts, grade what the answer claims against a short fact sheet (license, maker, price, platform) and report correct, wrong and declined as three separate counts. A model that says "I don't know" is doing better than one that hands out the wrong repository.
How do you run this measurement yourself?
Treat it as a small data pipeline with a versioned input, raw storage and a scoring layer you can re-run. It is what I would hand to a technical team.
-
Write the fact sheet and the alias list. Canonical name, known misspellings, owned domains (site, repository, store listings) and five to ten facts an answer can get wrong. Version both files.
-
Build the prompt panel by intent. Non-branded category prompts from real demand (search queries rewritten as questions), comparison prompts that name competitors, at least one prompt per product differentiator, and a separate set of entity prompts. Tag each prompt with its intent and never mix entity prompts into the mention rate.
-
Choose surfaces explicitly. At minimum one memory-only model and one retrieval surface, logged by exact model ID. API models are not the consumer apps (no personalization, no chat history, often a different model build), so name what you measured.
-
Sample repeatedly, and spread the samples out. 10 runs per prompt per surface is a floor. Mine were all collected within 49 minutes, and repeats that close together are correlated, so treat the intervals above as optimistic and spread real runs across days.
-
Store everything raw. One record per answer:
textrun_id, panel_version, prompt_id, intent, surface, model_id, settings, timestamp, sample, status, search_fired, answer_text, cited_urls[], retrieved_urls[], search_queries[], raw_responseRaw text and the raw response are the non-negotiable fields, and
statusandsearch_firedkeep "no data" apart from "no retrieval" or "failed request". Two of my own definitions changed while I analyzed this data (the citation rule and the license matcher), and each change was a re-score, not a re-run. -
Score four things separately, and audit the matcher. Mention (regex plus aliases, checked by reading a sample of both hits and misses: in this run "Ditto" in several AI-feature answers meant an unrelated product called Ditto, not the clipboard manager), inclusion and position (graded, with a hand-audited sample if a model grades), citation (a written rule for annotations versus inline links versus memory-recalled URLs), and entity accuracy against the fact sheet. Add retrieval-to-citation ratios where the surface exposes retrieval.
-
Report a vector per intent, with intervals. For each prompt or intent cluster: mention rate, inclusion rate and citation rate, each with n and a Wilson interval (these are yes/no outcomes per answer). Competitor share is not a yes/no outcome, since one answer names several brands, so report it as counts per answer rather than squeezing it into the same interval. Put the panel average last, labeled with the panel version.
-
Re-run on a schedule and annotate. Same panel, same surfaces, same time of day. Mark model releases, panel version bumps and your own site changes on the time series, because a model update and a content change look identical in the numbers.
The interval in step 7 is ten lines of code, which is no excuse to skip it:
from math import sqrt
def wilson(k, n, z=1.96):
if n == 0:
raise ValueError("no answers, no estimate")
p = k / n
d = 1 + z * z / n
c = p + z * z / (2 * n)
m = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n))
return ((c - m) / d, (c + m) / d)wilson(0, 80) gives an interval from 0 to about 4.6 percent. That is the honest version of "we are not visible": no mentions in 80 answers, and a rate above roughly 5 percent would have been unlikely to produce that.
What can a panel this small not tell you?
It cannot tell you what real users see. The runs went through APIs, not the consumer ChatGPT, Gemini or Perplexity apps, which add personalization, location, chat history and sometimes a different model. It is one day of data, so it says nothing about drift. It covers one product in one niche, and the product is mine, which is exactly why I can check the facts but also why you should read my grading of "inclusion" with that in mind.
What it does show is how much of an AI visibility number comes from measurement choices rather than from the brand: which prompts are in the panel, whether entity prompts are separated, which surfaces count, how citations are parsed, and whether hallucinated mentions score as wins. None of those choices are wrong in themselves. Leaving them unstated is.
This is the same pattern I found when I looked at what AI answers did to publisher traffic: the visible metric moved one way while the thing it was supposed to represent moved another. It also connects to a point I made about agent pipelines, that a model is a dependency that won't hold still. Log the model ID with every answer, or the next model release will look like your content strategy working or failing.
The panel is small on purpose: it is the smallest instrument that already breaks the idea of a single visibility score. The larger studies, with their methods and limits next to the findings, are in research.