Brand visibility
Check whether your brand is in the weights after every major AI update
A model can know your company and still leave it off the buying shortlist. Test both after an update.
Recognition can improve while recommendations fall
Illustrative example, not measured Aiso results. The same 20 buyer prompts before and after an update.
40%
Before: mentioned in 8 of 20 buyer prompts
20%
After: mentioned in 4 of 20 buyer prompts
A 20 percentage-point loss that a company-description test would miss.
Aiso insight: A model can describe your company correctly and still leave it out of every buying shortlist. After a major model update, measure the gap between recognition and recommendation. A better answer to “What is my company?” is encouraging. It does not tell you whether a potential customer will see you.
Here is an illustrative example. Your brand appears in 8 of 20 buyer prompts before an update and 4 of the same 20 afterwards. Its mention rate has fallen from 40% to 20%, even if the new model now gives a perfect company description. That is a 20 percentage-point loss in this test. Counting branded recognition alone would miss it.
This is why checking whether you are “in the weights” belongs in the release checklist for models your customers use.
A new model release gives you a reason to rerun the test. So does a change to the default model inside a major assistant. Record search-product changes separately: a different retrieval system can change visibility even when the underlying model stays the same.
The commercial question is whether the update changes your chances of being considered.
What “in the weights” actually means
In this context, it means a model can recall information about you without looking it up on the web. InTheWeights makes that idea easy to explore.
The site, created by Thomas Dimson and Joey Flynn, asks models to identify a name, groups similar descriptions and produces a strength score. Its stated method combines recognition across models with their reported confidence. It is primarily framed around people. Its methodology page also warns about hallucinated details, ambiguous names and uncalibrated confidence.
Use it as an initial check of personal recognition. For a company, build a separate test around the correct business and its category. A high score is not a buying-intent metric, and a fluent answer does not prove that a particular page appeared in a training dataset.
The published cutoff date cannot answer this for you
There is a clever way to investigate how recent a model's knowledge really is: ask it about events with verifiable dates and outcomes that could not reliably have been guessed beforehand.
In November 2023, Matt Mazur tested GPT-4 Turbo using 168 celebrity deaths. For January 2022, the model answered 596 of 650 trials correctly, or 91.7%. For March 2023, the reported count was 210 of 450, or 46.7%, despite an advertised April 2023 cutoff. That is about a 45 percentage-point difference between the two monthly groups. It is a historical experiment on one model, not a benchmark for today's assistants. Read Mazur's experiment.
The wider research supports treating cutoffs cautiously. In Dated Data: Tracing Knowledge Cutoffs in Large Language Models, Jeffrey Cheng and colleagues found that effective cutoffs can differ by resource and from the reported date. They identified older material inside newer CommonCrawl dumps and complications in deduplication as contributing factors. Read the paper.
My implication for marketers: publishing before the advertised cutoff is no guarantee that a model will know your company or its latest positioning. Knowing a recent public event does not establish that it knows recent facts about your industry either.
How to probe the cutoff without fooling yourself
Use a collection of independently verified events on both sides of the reported cutoff. Include older events as controls. Ask for the outcome without supplying it in the question, and record when the information became public, which can differ from when the event happened.
Run those questions with search and other retrieval tools disabled. A browser answer that finds a recent obituary tells you about retrieval, not stored knowledge. Simply typing “don't browse” into a normal chat is a weaker control than a setup in which tools are actually unavailable.
Repeat the questions and group the results by month. Look for a decline in accuracy rather than declaring an exact cutoff from one correct answer. An isolated success could be a guess; an isolated failure could reflect poor recall.
If a model repeatedly answers questions about events after its stated cutoff, document that discrepancy. Check for hidden retrieval and prompt leakage before interpreting it. You still cannot identify which stage of model development supplied the information from the answer alone.
For brand monitoring, this exercise gives context. It does not replace testing your brand directly.
The audit to rerun after a major update
Keep a small, fixed set of prompts based on actual customer needs. Include commercially relevant questions where competitors appear and you currently do not, provided your offer really fits. Preserve the wording across versions.
| Test | Example prompt | What to record |
|---|---|---|
| Brand recognition, search off | “What is [company]? If the name is ambiguous, say so. If you don't know, say so.” | Whether it identifies the right company; correct facts; invented details |
| Buyer shortlist, search off | “Which coworking operators should I consider for a six-person team in [city], with private offices and parking?” | Whether you appear without being named; which competitors appear; reasons given |
| Same buyer shortlist, search available | Use the identical buyer prompt | Whether search was actually used; sources cited; your inclusion; differences from the no-search answer |
| Branded factual accuracy | “Does [company] offer [service] in [location]?” | Whether the answer matches independently verified current information |
Use fresh sessions without your chat history, saved memory or uploaded brand documents. Keep language and location consistent. Where possible, use explicit model version IDs and record the date, settings and tool access. If a consumer product hides the underlying version, label the result as a product-level observation.
For a first diagnostic, 20 buyer prompts repeated five times in each condition gives 100 responses per model per condition. This is a proposed starting design, not a statistically validated minimum. Prompts are different buying situations, so those 100 responses are not 100 interchangeable observations. Examine changes prompt by prompt, especially when the overall difference is small.
If the previous model is still accessible, rerun it alongside the new one. Otherwise compare against archived answers and note that the comparison also spans time.
A before-and-after change is an observation. It does not, by itself, prove which training change, source or marketing action caused it.
What to do with the result
If recognition improves but buyer inclusion does not, inspect the explanation. Does the model place you in the wrong category? Does it know your offer but associate the relevant use case with competitors? That points you toward the specific positioning or evidence gap to investigate.
If a search-enabled answer includes you and a no-search answer does not, inspect the cited sources. They may already be helping customers discover you. Keep those pages accurate, and check whether other buyer prompts can retrieve the same evidence.
If you appear without search but the model repeats obsolete information, correct the public sources you control and the important third-party profiles. Those corrections can support retrieval when indexed. They do not rewrite an already deployed model's weights.
If you are missing in both conditions, inspect where the recommended competitors are mentioned and how those pages substantiate their fit. Work on relevant coverage and clear product facts. Nobody can guarantee that the next training run will include a given page.
Put an ROI threshold on the work
A recall score is not revenue. Treat the audit as a way to decide where to spend effort, then measure whether that effort produces qualified business.
Here is an illustrative break-even calculation. Suppose the audit and resulting fixes cost $600. A qualified opportunity has a 20% chance of becoming a customer, and a customer contributes $2,000 in gross profit over a defined period. The expected gross profit per opportunity is $400.
You need 1.5 additional qualified opportunities to cover the $600, or roughly two in practice. If the work generates three incremental opportunities at those assumptions, expected gross profit is $1,200 and expected ROI is 100%: ($1,200 − $600) ÷ $600.
These are example inputs, not Aiso results. Use your own close rate, contribution and measurement period. In particular, do not convert a 20-point increase in model mentions directly into a 20-point increase in leads.
Track identifiable AI referrals and ask prospects how they found you. Compare outcomes over a stated period, allowing for other marketing changes. The prompt audit tells you where visibility changed; your sales evidence has to establish whether that change mattered commercially.
Start with the buying questions you already care about. Save the answers before the next major update, so you have something concrete to compare afterwards.
At Aiso, we work on brand visibility in AI answers. This audit connects a simple recognition check to the customer questions worth monitoring.
Continue on LinkedIn: the shorter article or join the discussion.
