How often to run AI-search prompts: choose precision before cadence
Use repeated answers to estimate variation, without treating one, ten or forty daily runs as a universal rule.
First published . Analysis updated September 8, 2026.
Start with the change you need to detect
A mention rate is an estimate from a specified set of answers. The useful repeat schedule depends on the size of the change that would alter your decision, the observed variability and the cost of collecting another answer.
There is no universal rule that a category leader needs one run, a mid-sized brand ten and a challenger forty. A brand's size does not determine the sampling properties of every prompt. Start with a pilot and inspect the data.
Forty runs can still leave a wide interval
Suppose a brand appears in 16 of 40 independent runs of the same question. The observed rate is 40%. Using a 95% Wilson interval gives approximately 26.3% to 55.4%. That is an illustrative calculation, not a measured Aiso client result.
NIST describes the Wilson interval and why a symmetric approximation can behave poorly for small samples or rates near zero and one. Do not report a small-sample rate as if it had no uncertainty.
Repeating one question is not the same as sampling new demand
Runs of one question can share context, model behaviour and date effects. Treating every response as an independent new customer can understate uncertainty. For a category-level analysis, retain the prompt-family grouping and distinguish within-question variation from variation between questions.
Collect repeat observations across the periods you intend to compare. Do not pool a changed model, changed prompt wording and changed market into one trend without labelling those changes.
A practical repeat schedule
Choose the commercially relevant prompt families first. Run a pilot with repeated captures, logging the visible model, market, date and valid answers with no brand mention. Report failures separately. Inspect whether the observed variation is large enough to change the next content decision.
Add repeats where a noisy result would alter a meaningful action. Add new prompt families where a buying requirement is missing. Those are different uses of the testing budget.
Compare changes on the same basis
Keep the question, engine and scoring rule fixed when assessing a page update. Use similar unchanged questions where practical. Log platform releases and other site changes. A movement in a single day's answer does not establish an improvement.
Do not repeatedly check a significance threshold and stop as soon as it is favourable. Define the sampling window or an appropriate sequential method before collecting the comparison.
For context on actual response variability, see Aiso's 19-prompt, ten-run study. For coverage, see how to choose a prompt set.
Explore the related measurement tools
See Aiso’s prompt, fan-out and source-analysis workflow, with its sampling and coverage limits.
Explore Aiso