AI search performance metrics: define the denominator before tracking the trend

Measure prompt coverage, mention rate, citations and human referrals without turning a sample of answers into a market-share claim.

Benjamin Tannenbaum, Founder and CEO, Aiso
By Benjamin Tannenbaum · Founder and CEO, Aiso · LinkedIn
4 min read

First published . Analysis updated September 8, 2026.

Mention rate across repeated runs

Define a fixed set of relevant prompts and a repeat schedule. For each successful answer, record whether the brand is named under a consistent matching rule. Divide brand-positive answers by valid answers. Keep valid answers that name no brands in the denominator; report errors and unavailable responses separately.

Keep model, date, language, market and relevant conversation context. Report the raw numerator and denominator next to the percentage. A changing prompt mix can change the score even when performance on each individual question is unchanged.

Prompt coverage is a different rate

Coverage asks which distinct questions or prompt families ever include the brand under the specified repeat schedule. A brand appearing once in every prompt family can have broad coverage but a low per-run mention rate. Do not use the two percentages interchangeably.

Choose a coverage rule before collecting results: at least one mention, a minimum repeat rate or another explicit condition. More repeats can mechanically increase “at least once” coverage, so compare like with like.

Separate citations, recommendations and accuracy

A cited page can support a factual statement without earning the publisher a place in a product shortlist. Save the exact URL and the surrounding answer. Score product recommendations separately from background mentions and citations.

For accuracy, compare quoted prices, capabilities and restrictions with a dated source of truth. A positive mention with the wrong price can be less useful than a neutral but correct description.

Measure human referrals in analytics

Use session-level source information for human referral traffic. OAI-SearchBot and ChatGPT-User requests are agent activity, not referral sessions or answer impressions. There is no justified AI click-through rate when the denominator is a crawler-request count.

Google reports traffic from AI features within Search Console's Web search type. That is not a standalone report of every AI citation across assistants. Keep its traffic measures separate from your controlled answer-monitoring results.

Use a fixed comparison before interpreting change

Run the same prompt families against the same engines before and after a content change, with repeated observations. Check similar unchanged pages or questions where practical. Record other releases, promotions and model changes that could affect the result.

Sampling error falls with more independent observations, but repeated answers from the same question can be correlated. A larger panel of near-duplicate prompts does not automatically provide independent evidence. See how to choose a repeat schedule.

Use business outcomes to assess value

Track qualified enquiries and completed sales as separate outcomes. Keep your contribution-margin assumptions and measurement period visible. Do not multiply a sampled mention rate by an unsupported estimate of global prompt volume to manufacture revenue.

A useful update names the page changed, the tested questions, the observed difference and the uncertainty. The next task should follow from that evidence, not from a dashboard colour.

Explore the related measurement tools

See Aiso’s prompt, fan-out and source-analysis workflow, with its sampling and coverage limits.

Explore Aiso