We gave ChatGPT a celiac memory. The searches stayed the same.
In our controlled New York hotel test, personalization was visible in the answer long before it was visible in the observed search layer. A gluten-free travel memory changed what ChatGPT said to the user, but not the fan-out queries we captured.
of distinct celiac-memory answers mentioned gluten or celiac
12 of 20 answer instances
of observed celiac-memory fan-outs contained the dietary constraint
0 / 21 captured fan-out records with celiac, gluten or dietary wording
matched June 3 fan-out records were text-identical with memory on and off
a paired subset, not a market-wide estimate
The useful finding is not that personalization "does nothing." It is that personalization can show up at one layer of an AI answer without showing up at another.
We originally ran this experiment with SEO researcher Molly Nogami while investigating why some New York hotels appeared in ChatGPT and others did not. The project led to our Search Engine Land study on Bing and ChatGPT visibility. Memory was supposed to be one of the variables in that study. At the time, it produced no discernible difference in hotel mentions, citations or query fan-outs, so we pooled the runs for the Bing analysis.
We went back to the raw worksheet for this article and looked at the personalized wording itself. That is where the effect becomes visible.
The experiment
The opening prompt stayed fixed: "What are the best hotels in New York City?" We ran it 68 times in GPT-5.2 Instant and manually captured the full answer, citations and the web-search fan-outs exposed by ChatGPT. Reference chat history was switched off so one run would not contaminate the next.
We tested three memory states:
- saved memories off;
- saved memories on with unrelated real-user memories;
- saved memories on with one explicit travel memory: the user has celiac disease and needs hotels and restaurants to accommodate gluten-free menus.
The important control is what we did not change. The NYC hotel prompt itself did not mention celiac disease or gluten. The file we recovered supports a memory-on versus memory-off comparison. It does not contain a separate NYC arm where the dietary constraint was typed directly into the prompt, so we are not treating that as part of this result.
The extra analysis
Personalization moved the response language by 54.4 percentage points
In the celiac-memory condition, 12 of 20 distinct answer instances mentioned gluten or celiac needs: 60.0%. With memory off, 1 of 18 answers did so anyway: 5.6%.
60.0% minus 5.6% gives a 54.4 percentage-point difference in dietary-language incidence. This is a wording effect in this sample, not an estimate of how often ChatGPT personalizes answers globally.
What changed: the answer often acknowledged the constraint
The personalized language was usually late in the answer. ChatGPT would produce a general hotel list, then add a closing offer such as asking whether the user wanted the hotels refined for gluten-free dining or celiac-friendly breakfast. One answer added a generic gluten-free dining tip before the final follow-up.
That distinction matters. Someone reading only the final response can reasonably say, "ChatGPT personalized this." Someone reading only the fan-out logs can reasonably say, "The search behavior did not change." Both observations can be true at the same time.
What did not change: the observed fan-outs
We found no celiac, gluten or other dietary wording in the fan-outs captured under the celiac-memory condition. ChatGPT still searched variants such as "best hotels in NYC top luxury and boutique hotels New York" and "top rated luxury hotels New York City reviews."
The strongest paired check came from June 3. We could align 15 fan-out records across the memory-on and memory-off runs. All 15 query strings were exactly identical. That does not prove retrieval is never personalized. It does show that, in this controlled slice, the stored dietary constraint did not propagate into the observed web searches.
Observed retrieval layer
Generic NYC hotel fan-outs. No dietary terms. In the paired June 3 subset, memory-on and memory-off query strings matched 15 out of 15 times.
Response layer
The same broad hotel task, but the stored celiac constraint appeared in 12 of 20 distinct personalized answers, most often as a final refinement offer.
We did not see a recommendation-list effect we could defend
The original study looked at hotel mentions and citations across all 68 runs. We expected memory to shift those distributions. It did not do so in a discernible way, which is why the published methodology treated all three memory states as one dataset.
Individual answers still varied because ChatGPT is variable. The careful claim is therefore aggregate: we could detect personalized language, but not a reliable memory-driven shift in which hotels won visibility. We would not describe every recommendation list or every rank order as identical.
This is also why the Bing finding survived the memory test
The same dataset produced a different, much stronger signal at the retrieval-source layer. We extracted 25 unique query fan-outs and compared the pages ranking for them in Google and Bing. That work led us to Bing as the more useful search surface for explaining hotel visibility in ChatGPT. The full case study is published on Search Engine Land.
For practitioners, the connection is useful. A personalized answer does not automatically imply a personalized competitive set of web pages. If the fan-out stays generic, brands are still competing on a relatively stable retrieval surface even when the last sentence is tailored to the user.
What marketers should measure separately
Do not put "personalization" into one binary column. At minimum, separate the search query, the sources retrieved, the brands mentioned and the language used to present them. A model can personalize one of those without moving the others.
This also changes how to interpret AI visibility tests. If a brand disappears from a personalized answer, check whether the retrieval layer changed before assuming memory caused the loss. Our fan-out guide explains how to observe that layer, while our ranking-factors analysis covers the broader retrieval and selection chain.
A necessary 2026 caveat: ChatGPT memory has changed since this test
Our final matched captures were on June 3, 2026. On June 4, OpenAI announced a more capable memory architecture built on its "dreaming" system. OpenAI now describes memory as a way to carry forward preferences and constraints, and says current memory can tailor future replies. See the June 4 memory release and the current Memory FAQ.
So this is a historical controlled snapshot of one mechanism on one recommendation task. It is not evidence that current ChatGPT can never use memory to change recommendations or retrieval. In fact, that is exactly why we are publishing the raw distinction rather than turning it into a universal rule.
Methodology and calculation notes
- Prompt: "What are the best hotels in New York City?"
- Published study: 68 manual GPT-5.2 Instant iterations across three memory states.
- Celiac-memory answer analysis: 20 distinct answer instances after collapsing one duplicated captured answer row. 12 mentioned gluten or celiac, so 12 / 20 = 60.0%.
- Memory-off comparison: 1 of 18 answers mentioned gluten or celiac without that stored memory, so 1 / 18 = 5.6%.
- Difference: 60.0% - 5.6% = 54.4 percentage points.
- Fan-out language: 0 observed celiac-memory fan-outs contained celiac, gluten or dietary wording.
- Paired query check: 15 matched June 3 fan-out records were text-identical across the celiac-memory and memory-off conditions.
- Limits: one prompt, one vertical, manual capture, a small historical sample, model and memory behavior can change over time.
The takeaway for AI visibility work
Personalized wording and personalized retrieval are different measurements. Track both. Aiso uses real conversation data and observed fan-outs to show where a brand actually enters the recommendation chain, instead of treating every change in prose as a ranking change.
Talk with us about the research