The conversation is the real AI search query
We published three papers on arXiv that ask one connected question: what do prompt-level measurements miss about AI search? Together, the studies show that one answer can compress many information needs, the request often develops after the first prompt, and removing the preceding conversation can materially change the response.
median opening-prompt coverage of eventual vocabulary
670 English multi-turn commercial conversations
added an intent dimension after the opening prompt
same 670-conversation sample
produced a materially different answer without full history
weighted 180-pair experiment
Search has traditionally been measured as a sequence of discrete queries. Conversational AI breaks that unit of analysis.
A person can begin with a broad question, reveal constraints through follow-ups, compare alternatives introduced by the model, and end with a short instruction such as “which one would you choose?” The final answer may resolve a dozen facts, but neither the opening nor the closing prompt contains the full request.
Our three papers examine a different part of that process. The first looks inside the answer. The second follows the request across turns. The third tests what happens when the model loses that history. Read together, they point to the same conclusion: the useful measurement unit is the conversation.
Three measurements, one story
Answer-Reconstruction Search Density
The median answer contained 11 retrievable units and needed 3 structural facets to cover 80% of them across 183 information-seeking conversations.
An AI answer can compress a bundle of information needs that query-level analytics treat as one event.
Read the paper on arXivDoes the First Prompt Represent the Conversation?
In 670 English multi-turn commercial conversations, the opening prompt covered a median 50.0% of eventual unique content terms and 41.0% added an intent dimension later.
The initial prompt is often only the beginning of the request, not a stable specification of it.
Read the paper on arXivBeyond the Final Prompt
In a weighted experiment with 180 paired conversations, the full-context and final-prompt-only answers differed materially in 44.7% of cases.
Conversation history is not merely descriptive. Removing it can change the answer a model produces.
Read the paper on arXivLayer one: an answer is compressed search work
The first paper introduces answer-reconstruction search density, a way to estimate how much query and source work is compressed into one conversational answer. Across 183 information-seeking conversations containing 1,994 retained answer units, the median answer held 11 retrievable units. A median of three structural facets covered 80% of those units.
Those three facets are a lexical diagnostic, not three observed Google searches. A separate six-task live-web calibration found a median of 1.5 queries and two source pages to reconstruct 80% of an answer. The pilot is a feasibility check, not a population estimate. The important point is structural: one visible interaction can settle several questions and draw on several evidence paths.
That is why keyword volume and raw prompt counts are weak proxies for demand in AI systems. They count the expression, not the information work performed. Our paper-specific analysis of search density covers the full method and its limits.
Layer two: the request develops after the first prompt
The second paper examines 9,212 human-LLM conversations across commercial, values, and controversial domains. In the direct cross-domain multi-turn analysis of 8,133 conversations, the opening turn proved to be an incomplete representation of what users eventually asked.
In the 670-conversation English commercial cohort, the opening prompt contained a median 50.0% of the unique content terms that appeared across the eventual conversation. In 41.0% of those conversations, at least one observable intent dimension appeared only after the opening turn. The public PRISM cohort showed the same general pattern: 36.4% median opening coverage and a 41.7% later-added dimension rate.
A prompt panel built only from polished, standalone questions removes the process through which people clarify what they want. That matters for research, evaluation, and commercial visibility. You can read the companion analysis, The prompt is not the query, for the turn-level distributions.
Layer three: context changes the answer
The first two papers describe information structure. The third tests an outcome. For 180 paired multi-turn conversations, we compared answers produced with the full conversation against answers produced from the final prompt alone. We also tested a compact reconstruction of the preceding context, capped at 160 words.
Full-context and isolated answers differed materially in a weighted 44.7% of cases, with a 95% confidence interval of 33.8% to 56.1%. Full context also improved satisfaction by 0.49 points on a 0-to-4 scale. Adding the compact reconstruction reduced the material-difference rate to 30.8% and nearly closed the satisfaction gap, but did not make the answers equivalent.
The context effect varied by cohort: 68.5% in the commercial cohort and 35.4% in PRISM. That difference is a warning against assuming one universal effect size. It is also evidence that real-world commercial conversations may be particularly sensitive to information accumulated across turns.
A better measurement stack for AI search
Answer reconstruction
Measure the facts, facets, queries, and sources compressed into the answer.
Request trajectory
Track constraints and intent dimensions as they emerge across turns.
Context fidelity
Compare full-context outputs with isolated and compressed-context conditions.
What this changes in practice
For AI visibility measurement
Do not treat a list of isolated prompts as a complete market model. Use prompt tracking as a controlled benchmark, then add multi-turn scenarios that preserve evolving constraints and evaluate the resulting answers at the claim and source level. Sampling discipline still matters, which is why we separately publish guidance on how many prompts to track and how often to run them.
For content and brand teams
Map the decision, not only the head question. Publish evidence for the criteria, constraints, comparisons, objections, and follow-up questions that emerge after discovery. A page can matter even when its language never appears in the opening prompt, because it may support a later facet of the answer.
For product and privacy teams
Context has utility, but retaining everything is not the only option. The compact reconstruction in the third study recovered much of the satisfaction benefit while using a bounded summary. That creates a practical design space between context-free answers and unlimited retention, provided teams measure what is lost.
What these studies do not prove
The results do not establish a universal conversion effect, causal persuasion effect, or a single search-query equivalent for every AI answer. The web reconstruction used only six synthetic tasks. The context experiment covered 180 paired conversations and did not test persistent memory across separate sessions. All three measures are policy- and sample-dependent.
What the series does establish is narrower and useful: a prompt is an incomplete observation of conversational demand. Measuring the opening line alone misses later intent. Measuring the final line alone misses prior state. Counting the interaction alone misses the work compressed into the answer.
If AI search is conversational, its measurement has to be conversational too.
About the research
Aiso studies consent-governed and public conversational data to develop better measurement for AI search. We report cohort sizes, uncertainty, and limitations so the figures can be interpreted in scope.
Talk with us about the research