Quick answer
Measuring AI visibility reliably needs four parts: a defined prompt set grouped by topic and intent, repeated sampling across platforms and locations, metrics with explicit denominators, and stored evidence that links every figure to the answers behind it. Without all four, numbers cannot be compared over time.
The four parts
| Part | What it fixes | Common failure |
|---|---|---|
| Prompt set | Which questions represent your buyers | Prompts chosen after seeing results |
| Sampling | Random variation between runs | Reading a single answer |
| Metrics | How answers become comparable numbers | Missing or shifting denominators |
| Evidence | Whether figures can be checked | Numbers with no stored answers |
How to apply it
- 1Group prompts by topic and intent, and record the location and language of each.
- 2Run every prompt repeatedly on each platform and store the full response and any cited sources.
- 3Calculate mention, recommendation and citation rates against eligible answers only.
- 4Track omissions, wrong facts and source gaps as separate findings.
- 5Report changes only when the prompt set is unchanged, and mark any change to it.
Limits
Designing the sample
The framework rests on sampling, so the design of the sample matters. Decide the unit first. The unit of observation is one prompt run on one platform at one time under stated conditions. Everything else, from mention rate to share of voice, is calculated from a set of such units. Stating the unit makes it possible to say precisely what a rate is a rate of.
| Choice | Options | Consequence |
|---|---|---|
| Prompt coverage | Few core prompts, or many varied ones | Few gives depth on key questions; many gives breadth but is slower to review |
| Repeats | Single run, or several per prompt | Repeats smooth out sampling variation and let you estimate noise |
| Platforms | One, or several | Several exposes platform differences that a single one hides |
| Conditions | Fixed location, language, search on or off | Fixed conditions make periods comparable; varied ones need separate reporting |
Eligibility and denominators
A rate is only as meaningful as its denominator. Define an eligible observation before collecting data: the run completed, the conditions match the design, and, for citation rates, the platform was in a mode that can show sources. Failed and ineligible runs are excluded from denominators and reported separately, so the reader can see how many there were.
Quantifying uncertainty
Every rate calculated from a sample carries uncertainty, and it shrinks with the square root of the number of observations. Practically, quadrupling the number of runs roughly halves the margin of error. The simplest honest practice is to repeat the whole measurement and report how much the rate moved between repeats. That gives readers a feel for the noise without any statistical machinery, and it is more useful than a confidence interval computed under assumptions that may not hold.
Guarding against drift
- Freeze the core prompt set and record every addition with a date.
- Record model, mode and any visible settings with each run.
- Compare periods only when design choices are unchanged, and mark any change.
- Keep raw answers, so a change in method can be re-analysed rather than lost.
Quality checks before publishing a number
- 1Recompute a handful of rates by hand from the stored answers.
- 2Read a random sample of answers for each headline figure.
- 3Check that no run appears twice and no failed run is counted as a miss.
- 4Ask whether a second person, following the written method, would get the same result.
What this framework does not do
Is this specific to BrandWater AI?
No. Any team can follow the method with a spreadsheet, though it is slow at scale.
How many repeats are enough?
Enough that the rate stops moving much when you repeat the whole set. Start with several per prompt and increase until the change between repeats is small for the decision you need to make.
Do I need statistical software?
Not to start. A disciplined spreadsheet and a habit of repeating the measurement gives most of the value. Software helps at scale.
Sources
- 1. BrandWater AI: Methodology (Current)
Cite this page
BrandWater AI Research. (2026, 9 September 2026). A framework for measuring AI visibility: prompts, sampling, metrics and evidence. https://brandwaterai.in/research/ai-visibility-measurement-framework
How we work
Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.