Skip to content
BrandWater AI
Research note

A framework for measuring AI visibility: prompts, sampling, metrics and evidence

9 September 2026 3 min read

Headline findings

  • Define prompts before you measure, and change them deliberately.
  • Sample many times, because answers vary between runs.
  • State denominators and sample sizes on every figure.
  • Keep the answers so results can be audited.

Quick answer

Measuring AI visibility reliably needs four parts: a defined prompt set grouped by topic and intent, repeated sampling across platforms and locations, metrics with explicit denominators, and stored evidence that links every figure to the answers behind it. Without all four, numbers cannot be compared over time.

The four parts

Measurement framework
PartWhat it fixesCommon failure
Prompt setWhich questions represent your buyersPrompts chosen after seeing results
SamplingRandom variation between runsReading a single answer
MetricsHow answers become comparable numbersMissing or shifting denominators
EvidenceWhether figures can be checkedNumbers with no stored answers

How to apply it

  1. 1Group prompts by topic and intent, and record the location and language of each.
  2. 2Run every prompt repeatedly on each platform and store the full response and any cited sources.
  3. 3Calculate mention, recommendation and citation rates against eligible answers only.
  4. 4Track omissions, wrong facts and source gaps as separate findings.
  5. 5Report changes only when the prompt set is unchanged, and mark any change to it.

Limits

This observes public answers only. It cannot see private model data, hidden reasoning or internal retrieval, and it does not predict future answers.

Designing the sample

The framework rests on sampling, so the design of the sample matters. Decide the unit first. The unit of observation is one prompt run on one platform at one time under stated conditions. Everything else, from mention rate to share of voice, is calculated from a set of such units. Stating the unit makes it possible to say precisely what a rate is a rate of.

Design choices and their consequences
ChoiceOptionsConsequence
Prompt coverageFew core prompts, or many varied onesFew gives depth on key questions; many gives breadth but is slower to review
RepeatsSingle run, or several per promptRepeats smooth out sampling variation and let you estimate noise
PlatformsOne, or severalSeveral exposes platform differences that a single one hides
ConditionsFixed location, language, search on or offFixed conditions make periods comparable; varied ones need separate reporting

Eligibility and denominators

A rate is only as meaningful as its denominator. Define an eligible observation before collecting data: the run completed, the conditions match the design, and, for citation rates, the platform was in a mode that can show sources. Failed and ineligible runs are excluded from denominators and reported separately, so the reader can see how many there were.

Quantifying uncertainty

Every rate calculated from a sample carries uncertainty, and it shrinks with the square root of the number of observations. Practically, quadrupling the number of runs roughly halves the margin of error. The simplest honest practice is to repeat the whole measurement and report how much the rate moved between repeats. That gives readers a feel for the noise without any statistical machinery, and it is more useful than a confidence interval computed under assumptions that may not hold.

Guarding against drift

  • Freeze the core prompt set and record every addition with a date.
  • Record model, mode and any visible settings with each run.
  • Compare periods only when design choices are unchanged, and mark any change.
  • Keep raw answers, so a change in method can be re-analysed rather than lost.

Quality checks before publishing a number

  1. 1Recompute a handful of rates by hand from the stored answers.
  2. 2Read a random sample of answers for each headline figure.
  3. 3Check that no run appears twice and no failed run is counted as a miss.
  4. 4Ask whether a second person, following the written method, would get the same result.

What this framework does not do

It observes public answers. It cannot see private model data, hidden reasoning or internal retrieval, and it does not predict future answers. Findings describe association, not proof of cause.
Is this specific to BrandWater AI?

No. Any team can follow the method with a spreadsheet, though it is slow at scale.

How many repeats are enough?

Enough that the rate stops moving much when you repeat the whole set. Start with several per prompt and increase until the change between repeats is small for the decision you need to make.

Do I need statistical software?

Not to start. A disciplined spreadsheet and a habit of repeating the measurement gives most of the value. Software helps at scale.

Sources

  1. 1. BrandWater AI: Methodology (Current)

Cite this page

BrandWater AI Research. (2026, 9 September 2026). A framework for measuring AI visibility: prompts, sampling, metrics and evidence. https://brandwaterai.in/research/ai-visibility-measurement-framework

How we work

Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.

Put the theory to work on your brand.

Start with a first audit. We show where you appear, who is named instead, and what to fix first.