Aiso research · Paper 06 · Submitted preprint

Why AI mentions a brand: what 34,960 answers tell us

Benjamin Tannenbaum, Founder and CEO, Aiso
By Benjamin Tannenbaum · Founder and CEO, Aiso · LinkedIn
9 min read

First published .

75 anonymized projects. Repeated GPT and Gemini runs. A closer look at page relevance, search evidence and the brands that make it into an answer.

Read the submitted paper

49.0% vs 2.8%

GPT mention rate with an own-site citation versus neither signal. No branded fan-out in either group.

58.4% vs 3.8%

The corresponding Gemini rates. All user prompts in the comparison were unbranded.

GPT mentioned the target brand in 49.0% of answers when it cited the brand’s own website. In answers with neither an own-site citation nor a search query naming the brand, the mention rate was 2.8%. The first group also had no branded search query. The user’s prompt did not name the target brand in either group.

That is a large difference. It is not proof that getting a citation causes a recommendation. A citation and a brand mention can emerge together from the same answer-generation process. But it gives us a more useful place to investigate than a generic score for how “AI-ready” a page looks.

Our sixth paper, From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search, examines 34,960 GPT and Gemini observations across 75 anonymized projects. It covers 2,854 distinct monitored prompts, with repeated runs between June and September 2026. We submitted it to arXiv on September 19. Read the submitted manuscript.

What we counted

For each monitored request, we checked whether the answer named the target brand, whether its own domain appeared in the stored source list, and whether the engine named the brand in a search query. Those extra searches are often called fan-outs: the searches an AI system produces after receiving a user’s question.

We excluded prompts that already named the target brand. Otherwise, asking “Tell me about Brand X” would make brand inclusion an easy and largely uninformative result. A mention was the outcome, not necessarily a positive recommendation, a click or a purchase.

The large panel consists of monitoring runs, including manually curated and generated prompts. It is not 34,960 independent conversations with real buyers. Separate, smaller cohorts examine real-user requests and page relevance. They help explain different parts of the process; they are not one connected experiment. Methods and cohort boundaries, Section 4.

When the brand entered the search process, it appeared much more often

The clearest result comes from splitting answers into four groups. Every row below uses an unbranded user prompt. “Branded fan-out” means the engine itself included the target brand in a subsequent search query.

Target-brand mention rates in the large monitoring panel
Observed evidenceGPTGemini
No own-site citation; no branded fan-out2.8% (438 / 15,524)3.8% (524 / 13,801)
Own-site citation; no branded fan-out49.0% (866 / 1,769)58.4% (1,995 / 3,415)
Branded fan-out; no own-site citation64.4% (38 / 59)84.8% (28 / 33)
Both signals91.4% (117 / 128)100.0% (231 / 231)

Source: Table 3, Section 7.1. These are observed rates in this dataset, not forecasts for another brand. The 100% Gemini result describes 231 observations, not a guarantee. The branded-fan-out-only groups are particularly small.

There is also an important blind spot. A third-party page can supply information about a brand without the brand’s own site being cited. So “no own-site citation” does not mean “the engine saw no evidence about the brand.” Displayed citations do not reveal every page the system considered.

The difference remained when we held the prompt and brand fixed

One objection to that first table is obvious: perhaps well-known brands simply get both more citations and more mentions.

We therefore compared repeated runs for the same project, prompt and engine. We kept runs without branded fan-outs and looked at prompt groups where own-site citation changed between runs. The average within-group difference in mention rate was 40.2 percentage points for GPT, across 477 groups, and 49.0 points for Gemini, across 540 groups.

Average within-prompt difference in target-brand mention rate associated with own-domain citation: 40.2 percentage points for GPT and 49.0 for Gemini.
Repeated runs hold the project and wording fixed. They do not control every change in the engine or the web. Section 7.2.

This comparison reduces one kind of confusion: the difference is not just between different brands asking different questions. It still does not establish cause and effect. Engine updates, changing source material or other time-varying factors could influence both outcomes.

Past visibility added information that current citations did not

A brand that appeared before was more likely to appear again for the same monitored request. That history was useful even after accounting for current search evidence.

We fitted models on the earlier 70% of eligible sequential observations and evaluated them on the later 30%. One used previous mention history. Another used current own-site citation, branded fan-out and prompt intent. The third combined them.

How well each model distinguished mentions from non-mentions on the later observations
Model inputGPT AUCGemini AUC
Previous mention history0.9370.917
Current search evidence and intent0.8800.840
History plus current evidence0.9630.942

AUC measures how well a model ranks mention cases above non-mention cases. 0.963 AUC does not mean 96.3% accuracy. Combining history with current evidence also reduced the Brier probability-error score by 15.4% on GPT and 12.8% on Gemini relative to history alone. Table 4 and Section 7.4.

These are diagnostic models, not a tool that predicts an answer before the engine runs. Current citations are only known during or after the answer process. History was updated as earlier test-period observations became available, and the evaluation concerns monitored requests rather than an entirely unseen set of brands and prompts.

We call the historical component a “prior.” That does not prove the brand was stored in model memory. A consistent third-party footprint, repeated retrieval or other unobserved information could also explain persistent visibility.

A relevant page still helps. It does not settle the answer.

A separate case study compared 199 prompts with a 275-page website corpus. A transparent lexical relevance measure, BM25, estimated how well the best page matched each prompt.

That score was more informative about own-site citation on Gemini, with AUC 0.641, than on GPT, with AUC 0.545. The latter is close to the 0.5 chance benchmark. The website crawl also came after the visibility measurements, so this is a retrospective coverage check, not proof that editing those pages changed the answers. Section 5.2.

Relevance belongs in the analysis. What did not work was treating page relevance as a complete explanation of whether the brand appeared.

Real buyer questions often became several searches

In a smaller exploratory cohort of 80 unique real-user requests, 18 of 23 commercial requests triggered observable search fan-out, or 78.3%. Only 2 of 55 informational requests did, or 3.6%. The two remaining requests had other intent labels. The 20 requests that triggered search produced 42 queries, an average of 2.1 each. These small-sample rates should not be treated as estimates for all ChatGPT usage. Section 5.1.

The search queries often added evaluation criteria. A request about a new mascara produced searches about ingredients and how to choose, alongside a search for recommendations. That means a brand can be competing on questions the person never typed explicitly.

What is the practical AI-search equation?

People often summarize SEO as content plus backlinks. Our paper asks a different question: which steps can fail between a request and a brand mention?

The observed AI-search path runs from a request through page match, search fan-out, source exposure and answer selection. Historical brand visibility is a separate diagnostic input.
The paper separates page match, engine-mediated evidence and answer selection, while retaining information from earlier runs.

Match: does a page answer the request? Exposure: does the engine surface supporting evidence? Selection: does the brand make it into the answer? Prior: how likely has it been to appear for this request before?

The compact version is prior + (1 − prior) × match × exposure × selection. It is a mnemonic, not a discovered ranking algorithm. The fitted model in the paper uses historical mention propensity, current own-site citation, branded fan-out and intent. It does not independently estimate every term in that multiplication.

For a marketing team, the distinction changes the next task. When the right page does not exist, investigate the content gap. When a useful page exists but never appears among the sources, inspect the engine’s searches and cited third-party pages before writing another version. When the website is cited but the brand is omitted, read the answer and its alternatives. Those are different problems.

Measure the same requests repeatedly after a change. Keep GPT and Gemini separate. Count mentions, positive recommendations and resulting business outcomes separately too. This study does not show that raising any one of them automatically raises the others.

Read the research

Download the full submitted paper, with methods, tables and limitations. It builds on our five earlier papers, including the four-engine citation comparison, our work on conversation context and the study of observable buyer responses.

Disclosure: I am the founder and CEO of Aiso, which sells AI-search measurement software. The analysis uses proprietary Aiso research and monitoring data. Client identities are anonymized, and the published manuscript does not release raw conversations or client-level exports. The results are observational and tied to the engines, dates and samples studied.

Watch the short version

This 68-second excerpt was made with Substack’s built-in AI voiceover. Watch it on Aiso’s YouTube channel, or read the Substack edition.