Research paper

How much search is hidden in one AI answer?

Our first paper on arXiv introduces a way to measure the query and source work a single conversational answer compresses. The short version: the median information-seeking answer packs 11 retrievable facts into 3 distinct facets, and much of what looks like “deeper conversations are richer” is just longer answers. Read the full preprint on arXiv (2607.18904).

Ben Tannenbaum, Founder of Aiso
By Ben Tannenbaum · Founder, Aiso · LinkedIn
Updated July 2026 · 10 min read · Reviewed by the Aiso Research Team

Bottom line

A conventional web search shows its work: you type, open results, compare sources, and stop when you have enough. A conversational answer hides that sequence behind one response. This paper measures what is inside it. Across 183 information-seeking conversations (1,994 answer units), the median answer held 11 retrievable facts organised into a median of 3 lexical facets to cover 80% of it. A separate 6-case live-web pilot found the single opening query covered a median 70% of an answer, and rebuilding 80% took a median of 1.5 queries but 2 source pages. Query compression and source compression are not the same thing.

11
retrievable facts packed into the median information-seeking answer
3
distinct facets to cover 80% of that answer (a lexical diagnostic, not a live query count)
70%
of the answer captured by the single opening query (6-case web pilot)
~76%
of the raw 'longer chats are denser' gap is answer length, not dialogue depth (Aiso analysis)

Source: Tannenbaum, B. (2026), “Answer-Reconstruction Search Density”, arXiv:2607.18904 [cs.IR]. The facet count is a structural, lexical diagnostic, not a count of live web searches. Readings marked “Aiso analysis” are our own from the paper's figures.

Search density measures informational compression, not correctness, source quality, or persuasion. A compact false answer can be highly retrievable; a true local fact can be hard to retrieve. Those need separate measurement. The work sits behind Aiso's AI-search visibility platform.

The measurement gap

Query counts still anchor interactive information retrieval, search analytics, and marketing demand measurement. But a single conversational turn can express and resolve several queryable aspects at once, and a long answer can also just repeat one idea. Neither prompt count, answer length, nor dialogue depth tells you how much conventional search work the delivered information represents.

Answer-reconstruction search density reverses the usual direction of measurement. Instead of starting from what a user did and estimating effort, it starts from a finished answer and asks for the smallest set of queries, and separately the smallest set of source pages, that could rebuild a fixed share of its retrievable content. The problem is a partial set-cover, solved exactly rather than approximately, under a documented and dated policy so the number means something specific.

What one answer contains

MeasureValueNote
Retained retrievable units per answer11 (median)IQR 7 to 16; uncapped max 142
Distinct facets to cover 80%3 (median)IQR 2 to 4; mean 3.28
Retrievable units covered per facet3.25 (median)each facet is a dense mini-answer
Answers exceeding the 16-unit cap34.4%the distribution is heavy-tailed

Primary structural policy (80% coverage, TF-IDF cosine threshold 0.15, 16-unit cap) over 183 conversations. The scale is policy-sensitive, but case ordering is stable: adjacent-policy rank correlations run 0.85 to 0.93, and the median stays at 3 across 8, 12, and 16-unit caps.

The depth result: volume, not dialogue

The tempting story is that back-and-forth conversations produce richer, denser answers. The raw numbers seem to agree: multi-turn cases showed 0.90 more facets on average (incidence-rate ratio 1.32). But multi-turn cases also carried more than twice as much material, a median of 18 uncapped units versus 8 for single-turn.

Adjust for that, and the depth effect nearly vanishes. The multi-turn gap falls to 0.22 facets (95% CI −0.35 to 0.78), and the Poisson incidence-rate ratio drops from 1.32 to 1.06. Within answer-size bins the depth groups line up. About three-quarters of the apparent effect is answer volume, not the multi-turn form itself. Deeper dialogues often produce more answer material, and it is the material that carries the facets.

The multi-turn density gap, before and after adjusting for answer size

Unadjusted gap+0.90 facets
After adjusting for answer units+0.22 facets
Full model+0.07 facets
Multi-turn coefficient from the descriptive OLS models. Bars are scaled to the unadjusted effect. The gap collapses once answer volume is controlled for.

The live-web pilot: queries and pages diverge

To show the measure works on the open web, the paper runs a small, fully public calibration: 6 synthetic tasks, 36 fixed queries, one reviewer, on 20 July 2026. It is a feasibility check, not a population estimate, and every task, query, and support judgment is released. The pattern is still instructive.

Rebuilding 80% of an answer took a median of 1.5 queries but 2 source pages. The opening broad query alone covered a median 70% of the units. So one query can carry most of the answer, yet the evidence behind it is spread across more pages than queries. Query compression is not source compression.

Public taskDomainOpening coverageQueries to 80%Pages to 80%
Induction cooktopPurchase60%23
Heat pumpPurchase80%12
Ransomware preparationSecurity100%11
Election-claim checkCivic60%23
CRM selectionOrganization80%11
Career retrainingDecision60%22

Public manual calibration (paper, Table 2). Median across the 6 tasks: 1.5 queries, 2 pages, 70% opening coverage. Queries are answer-aware, so they understate unaided human search effort.

Aiso analysis: four readings that matter for demand

These are our readings of the paper's own figures, framed for anyone who measures search demand. They are directional, and they respect the paper's caution that the facet count is not a live query count.

ReadingFigureWhy it matters
Query compression vs source compression1.5 vs 2Rebuilding 80% of an answer took a median 1.5 queries but 2 source pages. Winning the query is not the same as being in the pages.
Opening-query capture70%The first broad query covered a median 70% of an answer's units, so first-answer presence captures most of the value.
Depth effect that is really answer length~76%The multi-turn density gap fell from 0.90 facets to 0.22 after adjusting for answer size, so most of it is volume, not dialogue depth.
Demand hidden per answer11 to 1One prompt can resolve a median of 11 retrievable facts a keyword tool would count, at most, as a single query.

Aiso analysis of figures reported in arXiv:2607.18904. The web-pilot numbers come from 6 synthetic tasks and should be read as feasibility, not scale.

A brand, policy position, or source can be absent from the opening prompt yet relevant to one reconstructed answer unit. Evaluation based on a single short keyword can therefore miss the source competition inside a synthesized response.
From the paper's implications for search and marketing.

What it means for measuring AI demand

1

Keyword volume undercounts demand

A single answer at the corpus median resolves 11 retrievable facts across 3 facets: criteria, alternatives, constraints, availability, maintenance, risk. A keyword tool sees one query, or none. If you size AI demand by head-term search volume, you miss almost everything the answer settled.

2

The audit unit is the answer, not the keyword

A brand or source can be absent from the opening prompt yet decisive for one reconstructed answer unit. The paper's point is that task-level bundles of answer units and evidence paths are a more faithful thing to track than a single short keyword.

3

Be in the pages, not just the query

In the web pilot, source pages needed (2) exceeded queries needed (1.5). One query can surface evidence spread across several pages, and those pages are where source competition actually happens. Ranking for the query is necessary, not sufficient.

4

First-answer presence is worth the most

The opening query covered a median 70% of the answer. Whatever the model reaches first does most of the work, so being present in that first pass matters more than being reachable somewhere in a long tail.

How it was measured

  • Private structural analysis: 1,180 populated rows to 808 unique conversations to 401 English-labelled to 183 eligible information-seeking cases (1,994 answer units).
  • Consent-governed corpus. No organisation identity is used as a feature; raw transcripts, record identifiers, and private queries are excluded from the paper and arXiv package.
  • Answer units are atomic, externally retrievable facts. Social and stylistic material is removed. Facets are found by exact partial set-cover over TF-IDF prototypes.
  • Public live-web calibration: 6 synthetic tasks, 36 fixed queries, one reviewer, 20 July 2026. Fully released.

What the numbers do not say

  • The facet count is lexical and structural. It is not “three web searches” and not a formal lower bound on live-web queries.
  • The corpus is selected for commercial relevance and is not representative of all AI use. Only 183 cases met the bar; 351 unique records had an unknown language label.
  • The web pilot is a 6-case feasibility test with one reviewer, not inter-rater reliability or a population magnitude.
  • Density is not correctness, reading time, decision quality, persuasion, or economic value. It should be paired with source-quality and outcome measures.

Frequently asked questions

What is answer-reconstruction search density?

Answer-reconstruction search density (ARSD) is the minimum number of distinct query actions needed, under a fixed and dated reconstruction policy, to support a target share of the atomic, retrievable facts in a completed conversational answer. A parallel page-density measure counts the minimum distinct source pages. Together they estimate how much conventional search work a single synthesized answer compresses. It is a policy-relative measure of informational compression, not of correctness, source quality, or persuasion.

How many facts are in a typical AI answer?

Across 183 consent-governed, information-seeking conversations (1,994 retained answer units), the median answer held 11 retrievable units, organised into a median of 3 lexical facets to cover 80% of the answer. Each facet covered about 3.25 units. The distribution is heavy-tailed: 34.4% of answers carried more than 16 units, and the maximum reached 142.

Does one AI answer equal three Google searches?

No, and the paper is explicit about this. The figure of 3 is a structural facet count based on lexical (TF-IDF) similarity between answer units. It is a diagnostic of how fragmented an answer is, not a count of live web queries, and it must not be reported as 'three web searches'. The separate live-web pilot, on 6 synthetic tasks, found a median of 1.5 queries and 2 source pages to reconstruct 80% of an answer, but that is a feasibility check, not a population estimate.

Does a longer conversation produce a denser answer?

Mostly it just produces more answer material. Multi-turn conversations looked denser in raw terms (0.90 more facets; incidence-rate ratio 1.32), but the gap fell to 0.22 facets (IRR 1.06) once you adjust for how many answer units the conversation contained. Within answer-size bins the depth groups nearly align. So answer volume, not dialogue depth on its own, accounts for most of the association.

What does search density mean for marketers measuring AI demand?

Query-centric analytics see expressions of demand, not the whole information need. When one answer resolves selection criteria, alternatives, constraints, and risk in a single private interaction, visible keyword volume can diverge sharply from real informational demand. The practical shift is to audit AI visibility at the level of answer units and the sources behind them, across the facets of a question, rather than tracking a single head keyword.

Measure demand at the answer, not the keyword

Aiso tracks the real prompts customers ask AI assistants, the facets each answer covers, and the sources behind them, across ChatGPT, Claude, Gemini, Perplexity, and Copilot. See where your brand is inside the answer, and where it is missing.