AI retrieval · September 2026
ChatGPT is searching websites differently. Onsite content matters more now.
First published .
On August 8, Promptwatch measured the share of ChatGPT fan-out searches using the site: operator jumping from 0.37% to 16.8%. ChatGPT was suddenly searching inside specific domains far more often.
46×
increase in site: fan-outs measured by Promptwatch
86%
drop in Reddit's visible citation share in Promptwatch's August window
For the last two years, a lot of AI-search advice has focused on getting mentioned on Reddit, Wikipedia, review sites and large publishers. That advice was not irrational. Those sites were often visible in ChatGPT answers, and third-party corroboration still matters.
But the retrieval layer is changing quickly. The latest evidence points to a more direct pattern: ChatGPT increasingly decides which domain it wants to inspect, then searches inside that domain. That makes your own website more important, not because third-party coverage stopped mattering, but because your site is increasingly part of the verification step.
The change is visible in the searches ChatGPT runs
Promptwatch tracks the fan-out searches ChatGPT runs behind an answer. On August 8, 2026, it measured site:-restricted searches rising from 0.37% to 16.8% of fan-outs. Average searches per response rose from 1.08 to 1.83 at the same time.
A separate analysis of 615 buying questions found an even clearer sequence: a broad opening search identified possible vendors, then later fan-outs used site:vendor.com to verify them. In that dataset, site: appeared in 0% of opening queries and 64.7% of queries by the fourth search position.
That is a useful mental model. The first search proposes candidates. The next searches check the candidates against their own sites.
Third-party platforms have become less dependable as visible citations
The most dramatic recent example is Reddit. Promptwatch measured Reddit at 3.83% of ChatGPT Search citations from July 18 through August 7. By August 14 to 17 it averaged 0.52%, an 86.4% relative drop.
This was not the first large swing. Semrush tracked more than 230,000 prompts in 2025 and saw Reddit fall from appearing in close to 60% of ChatGPT responses in early August to around 10% by mid-September. Wikipedia fell from roughly 55% of responses to under 20% over the same period.
The exact percentages should not be combined into one time series. They come from different tracking systems with different denominators. The important part is the volatility. A third-party platform can go from dominant to marginal in visible citations without anything on your side changing.
There is also evidence from narrower B2B SaaS prompt sets that documentation and first-party pages are taking a larger share. Discovered Labs reports docs pages moving from 36.9% to 55.1% of citations between its 2025 and 2026 captures, while listicles fell from 6.0% to 1.0%. I would treat that as directional because it is one vendor's prompt set, not a universal ChatGPT benchmark.
Do not confuse “not cited” with “not used”
This is the important nuance.
A page can enter the retrieval pipeline without appearing as a blue citation. Ahrefs analyzed 1.4 million prompts and found that ChatGPT ultimately cited only about half of the URLs that appeared in its retrieval data. It also found that the dedicated Reddit retrieval channel supplied huge numbers of candidate URLs while only 1.93% of those data points became citations.
Ahrefs is careful here too: seeing a URL in retrieval data does not prove ChatGPT opened and read that page. But it does prove that “retrieved” and “cited” are different layers. So when Reddit citations collapse, the correct conclusion is not “ChatGPT no longer uses Reddit.” The visible citation layer changed.
The same distinction matters for any publication. The websites involved in retrieval and the websites shown to the user as citations are not necessarily the same set.
Training data is another layer we cannot observe directly
Third-party mentions may also matter because they can enter future training or model-update data. We cannot inspect the full training corpus of a frontier model, so nobody can tell you that a specific Reddit thread, article or directory listing made it into the weights.
Published cutoff dates are also less clean than marketers often assume. In our earlier piece, we looked at how to test whether a brand is “in the weights”. Historical experiments and academic work show that effective knowledge cutoffs can vary by source and do not behave like one simple hard line.
That is why I would not abandon third-party distribution. I would change the order of operations.
Start with the website you control
If ChatGPT already has your company in mind and runs site:yourdomain.com pricing, site:yourdomain.com integrations or site:yourdomain.com enterprise, there needs to be a useful page on the other side.
The content has to answer the actual question. A generic 1,500-word SEO article is not automatically valuable. Product pages, comparison pages, original research, documentation, location pages, pricing explanations, FAQs and concrete case studies can all be better retrieval targets when they contain the fact ChatGPT is trying to verify.
We have seen the same broader relationship in our own research. In 34,960 GPT and Gemini answers, own-site citations and past visibility tracked brand mentions. That does not prove causation, but it is another reason not to treat the company website as a secondary surface.
The practical work is to find the questions where you should be eligible to appear, inspect what ChatGPT searches for behind those questions, and make sure the site contains a strong answer. Our fan-out work is useful here because the prompt itself is often not the query the model sends to search.
Then expand the surface area
Once the site is good, third-party distribution becomes much more useful. Reviews, Reddit discussions, Wikipedia where appropriate, directories, trade publications, YouTube and independent comparisons can corroborate the claims you make on your own domain and expose the brand in places the model may retrieve or learn from later.
But doing this in the reverse order is fragile. If the third-party citation mix changes tomorrow, you have little control. If ChatGPT decides to inspect your domain directly and the relevant information is missing, the opportunity is also gone.
The strategy we keep coming back to at Aiso is simple: get your own website in order first. Make it useful enough that an AI checking your claims can actually find evidence. Then expand your surface area across the rest of the web.
The harder question is what “good” onsite content means for your specific prompts. That is measurable. Start from the real questions customers ask, inspect the retrieval paths, and build the pages that close the gaps.
