AI Visibility

ChatGPT stopped browsing the web. Now it asks your website directly.

On 8 August, ChatGPT changed how it searches. It reads roughly twice as many pages as before and credits far fewer of them — and a large share of what it reads, it now requests from named domains, by name. For the first time since AI answers started mattering, the lever that moves the outcome is one you own outright.


Something in ChatGPT's search behaviour changed in the second week of August, and the first people to notice were the ones watching Reddit.

The symptom everyone noticed

Reddit had spent two years as one of the most-cited domains in AI answers. Then, over about a week, it very nearly stopped appearing. Promptwatch, monitoring ChatGPT's real interface, measured Reddit's share of all citations falling from a 3.83% average across 18 July to 7 August, to a 0.52% average across 14 to 17 August — a relative drop of 86.4%. The decline started on 8 August and crossed below one percent on the 14th.

The figure doing the rounds secondhand — a fall from 4.5% — is a peak day, not a baseline. The measured before-and-after is 3.83% to 0.52%.

The interesting part is not that Reddit fell. It is why it fell, because Reddit did not get worse and OpenAI did not ban it. Reddit fell because of a change in how the question gets asked — and that change has a much bigger consequence for anyone selling something.

What actually changed

When you ask ChatGPT a question, it does not run your question as a search. It decomposes it into a set of its own internal queries — the industry calls these fan-out queries — runs those, reads what comes back, and writes an answer. You never see them.

What changed on 8 August is the shape of those queries. ChatGPT started scoping a large fraction of them to a single named site, using the site: operator that has existed in search engines for twenty years. Instead of asking the open web what the best espresso grinder is, it asks site:yourstore.com what your grinder costs.

A site:-scoped query cannot return Reddit unless the model asks for Reddit. Reddit's collapse is not a penalty. It is arithmetic. Mechanism per Promptwatch, 20 August 2026

Four independent measurements of how far this has gone, and they do not agree — which is itself the most useful thing to know about them:

  • 0.37% → 16.8% of all fan-out queries used site: (Promptwatch).
  • 0.3% → about 23% over roughly a year (David Konitzny, Peec AI).
  • 64% of queries contained site: on ChatGPT 5.6 Sol, across about 4,000 prompts — with site:, "official" and "gov" as the three most common words in the query set (Chris Long, Nectiv).

The spread is wide because those samples cover different prompts and different model tiers. The floor is what matters: even the most conservative measurement puts site-scoped retrieval at 16.8% of ChatGPT's internal searches, up from 0.37%. At the top end it is approaching two-thirds. Whichever number you take, this went from a rounding error to a primary retrieval mode inside a year, and it stepped up sharply on 8 August.

Reading more, crediting less

The second half of the change is that ChatGPT is now doing considerably more work per answer. Searches per response roughly doubled. On the reasoning tiers the increase is much steeper — one measurement puts fan-out queries per prompt at 2.17 rising to 7.61, with the longest observed chain going from 4 searches to 29. Prompts that a single round of fan-out could satisfy fell from 94.0% to 43.5%.

And yet the number of sources it credits went down. Unique domains cited per response fell from 19 to 15. In one detailed analysis of 57 conversations covering 3,554 retrieved pages, only 110 pages — 3.1% — made it into a final answer.

Retrieved

The page was fetched and read. It shaped what the model said, including whether your brand got named at all. Nobody linked to you and no analytics tool recorded a thing.

Cited

The page appears as a source the reader can click. This is the only one of the two that can send you a visitor — and it is the one getting rarer.

Keep those two words apart when you read anything about AI visibility, including your own dashboard — we wrote about why that distinction reorganises the whole problem earlier this month. A tool reporting "visibility" without telling you which of the two it counted has not told you anything. Most of the contradictions between published studies this year dissolve the moment you check which one each was measuring.

The model decides who to ask about before it reads anything

This is the finding that should change what you do on Monday.

In that same analysis of 3,554 retrieved pages, 21 of 27 opening fan-out queries contained a brand name the user had never mentioned. The model brought those names with it. And the difference in outcome was enormous:

Your brand appears in the model's own fan-out query

cited 68.9%

The model already knew you belonged in this conversation, went looking for you specifically, and credited what it found roughly seven times out of ten.

Your page merely turns up and gets fetched

cited 2.1%

You were in the pile. The pile is now 24 pages deep and three of them survive.

The contest is not for the citation. It is for being one of the names the model types into its own search box before it has read anything at all. Suganthan Mohanadasan, 57 conversations / 3,554 retrieved pages

That reframes the work. Optimising a page for a query the model was never going to point at your domain is effort spent on the 2.1% path. Being known as one of the handful of names in your category is what puts you on the 68.9% path — and once you are there, what the model finds on your own site decides what it says about you.

What the shift rewards, and what it punishes

The mix of pages ChatGPT retrieves moved at the same time. Product pages now account for 16.39% of retrieved pages. Listicles, how-to guides and comparison round-ups all lost retrieval share.

Those three formats are precisely what the last eighteen months of "generative engine optimisation" produced at scale — the "best 10 tools for X" page whose ranking exists to be harvested by a model. Reddit now flags around 25,000 spam posts a day using its own LLM detection. Whether OpenAI's change is aimed at that, or is simply a cheaper and faster way to get a reliable answer, the effect on that playbook is the same.

Four things you actually control

Here is why this shift matters more than the others. Every previous version of AI-visibility advice ended somewhere you could not reach: earn press, get reviews, be discussed on Reddit. A model that types site:yourstore.com is asking a question you get to answer.

01

Put the facts in crawlable HTML

Prices, specifications, model numbers, compatibility, shipping and returns, warranty terms — as readable text on the page. Not baked into an image, not assembled by JavaScript after load. The crawlers that fetch pages at question time do not reliably run JavaScript. If your spec table is rendered client-side, the model asked and got nothing.

02

Make the official domain unambiguous

Use "Official Site" naturally in your title tags and meta descriptions, and keep one canonical domain reinforced consistently everywhere you appear. This matters most if you share a name with another company or are not yet the obvious answer in your category.

03

Invest in product pages over round-ups

The retrieval mix moved toward the page that holds the actual answer about the actual product. Your own comparison content is worth less to a model than a product page that states its specifications plainly.

04

Be a name, not a result

Getting into the fan-out query itself is a notoriety problem, not a markup problem — reviews, comparisons, coverage, being the brand people name unprompted. Slower than the other three, and worth more than all of them.

A test you can run this week

Pick five prompts a customer would plausibly ask before buying from you. Run each one five times — not once; a single run of a probabilistic system tells you almost nothing. Then check something most people skip: whether your brand appears in the searches the model ran, not just in the answer it wrote.

If you are absent from the searches, no amount of on-page work will fix it, and you have a notoriety problem. If you are present in the searches but absent from the answer, the model asked about you and did not like what it found on your site. Those are different problems with different budgets.

The failure mode nobody is telling you about

If the model is going to ask a domain by name, it has to know which domain. Sometimes it does not. Malte Landwehr documented ChatGPT running site:census.com when researching a company that actually owns getcensus.com — every result belonging to an unrelated business.

This is not a rare edge case. A 2025 Netcraft study found that roughly 33% of brand login links produced by AI models pointed at domains the brand did not own, and 29% pointed at domains that were unregistered, inactive or parked.

A brand in that position is invisible for a reason no content strategy will ever fix, and nothing in a standard analytics stack will tell them. If you have a generic name, a name you share, a recent rebrand, or a legacy domain still floating around, this is worth checking before you spend anything on content.

One engine, not "AI"

Scope matters here, because it decides what you do next. This is ChatGPT. Google's AI surfaces moved 11–30% over the same window — a drift, not a cliff.

The five engines that matter are diverging, and they reward different things. Perplexity and Google's AI surfaces still lean heavily on Reddit and forum discussion. ChatGPT increasingly asks the brand's own domain. A single blended "AI visibility score" averages those into a number that cannot tell you which lever to pull — which is why we report each engine separately, and keep retrieval and citation apart inside each one.

This is also unlikely to reverse. A site-scoped fetch is cheaper and faster than evaluating the open web, so whether OpenAI made the change to fight spam or to cut inference cost, the economics point the same way.

The durable part is small enough to state in one sentence, and big enough to plan around: the assistant is now asking your website about your products, and what it finds there is what your customers will be told. That has never been true before. Unlike every other lever in this field, it is one you own.

Sources

  1. Lily Ray, What we can learn from evolving ChatGPT fan-out queries, 17 August 2026 — synthesis, and the source for the Konitzny, Long, Resoneo, Mohanadasan and Landwehr figures cited above.
  2. Klaas Foppen, Why did ChatGPT stop citing Reddit? Inside the August 2026 citation collapse, Promptwatch, 20 August 2026 — real-interface citation monitoring, 18 July to 17 August 2026.
  3. David Konitzny, Peec AI — fan-out volume, site: share and retrieved page-type analysis.
  4. Chris Long, Nectiv — approximately 4,000 prompts on ChatGPT 5.6 Sol against a 2025 baseline.
  5. Suganthan Mohanadasan — network analysis of 57 conversations and 3,554 retrieved pages.
  6. Olivier de Segonzac and the Resoneo team — unique cited domains per response.
  7. Netcraft, 2025 — AI-generated brand login link study.
  8. Ahrefs, ChatGPT's most cited pages — 1.4M prompts, February 2026, for the retrieval-versus-citation baseline.

Every figure above names who measured it and on what sample. Where measurements differ, the range is given.

Find out what the assistants find

On Belay probes five AI engines every week, keeps retrieval and citation apart instead of blending them into one score, and turns the result into a ranked plan for your own site. It is included in every plan — not a per-seat add-on.

Get started