Most AI visibility tools give you one number. Your score went from 14 to 17. Nobody can tell you what to do about that, because a single score collapses three different failures into one digit — and the three need completely different work.
The useful frame, the one the research converged on through 2026, is that there are two separate things a brand wants from an AI answer, and they are not the same event.
Being used is not the same as being cited
When an assistant answers a question, it may pull your page into its context and let that page shape what it says — without linking to you, and sometimes without naming you. That is being used. Separately, it may attribute part of the answer to a specific URL the reader can click. That is being cited.
These come apart more than you would expect. Ahrefs ran 1.4 million prompts through ChatGPT and found that only 49.98% of the URLs it retrieved went on to be cited. Half of everything the model read was discarded before attribution. Semrush, working from a corpus of 126 million prompts, found that the overlap between brands mentioned in an answer and domains cited as its evidence can be as low as 30% on Gemini.
That second number is the one worth sitting with. The single largest body of content shaping ChatGPT's answers produces almost no visible citations. If your dashboard counts citations, all of that influence is invisible to it — and so is any work you do to earn it.
Three stages, and only one is about your writing
Every answer passes through the same sequence. A brand can fail at any of the three, and the failures look identical in a blended score.
The engine looked it up
Engine behaviourIt ran a live web search instead of answering from memory. You do not control this — it is a property of the engine and the question being asked.
It cited one of your pages
CitingOf the answers where it searched, this many pulled in a page from your domain. Failing here is a citing problem — the engine never opened your door.
It named you in the answer
NamingOf the answers that cited your page, this many actually said your name. Failing here is a naming problem — your page was open and your name still wasn't in the text.
The third stage is the one that surprises merchants. An engine can open your product guide, read it, and then recommend three competitors — because your page never put a plain, liftable sentence anywhere the model could find one.
Retrieval engines and memory engines
Assistants answer in one of two modes, and the mode decides what kind of work can move them. This is the single most useful distinction for deciding where to spend a quarter.
Searches the live web, then answers
It opens pages at question time. What is on your site today can change the answer today. Publishing, restructuring and fixing extraction all move these engines, often within days.
Moves on: on-site content, page structure, titles that match real questions, crawler access.
Answers from what it already learned
No live search. The answer comes from training. Nothing you publish this week can reach it. These engines move on third-party authority — being discussed where the training data comes from.
Moves on: coverage on high-authority third-party sites, community discussion, video, editorial roundups.
Mode is not a fixed property of an assistant. The same product can switch when its underlying model changes. We have measured one major assistant go from running a live web search on roughly 9% of questions to nearly half of them — with no announcement, and no change on our side beyond the model it was pointed at.
A visibility programme built on last quarter's assumption about one engine is a programme aimed at the wrong lever. This is a thing to re-measure, not a thing to look up once.
What the research says actually earns a citation
Two findings are worth more than the rest of the literature combined, and neither is technical.
The first is from a Princeton-led study presented at KDD 2024, which measured content changes lifting visibility in AI answers by up to roughly 40%. The three highest-impact techniques were, in order: citing credible sources in the body text, adding specific statistics, and quoting named authorities. Page speed was not a significant factor.
The second is about shape. Retrieval systems work at passage level, not page level. Your opening paragraph, each section, each FAQ answer and each table row is scored separately and competes separately for citation. A March 2026 framework from the University of Tokyo and University of Tsukuba found that document architecture predicts citation probability largely independently of what the content actually says.
Ahrefs adds a third, narrower finding: the best predictor of a retrieved page surviving to become a citation was whether its title and URL matched the specific sub-question the assistant had broken the prompt into. Not domain authority. Not freshness.
49.98%
of URLs ChatGPT retrieves go on to be cited
Ahrefs · 1.4M prompts
~40%
visibility lift measured from content changes alone
Princeton GEO · KDD 2024
11.4%
conversion rate on AI search traffic, against 5.3% organic
Similarweb
30%
how little brand mentions and cited domains can overlap on one engine
Semrush · 126M prompts
That conversion figure is the reason none of this is optional. AI search sends less traffic — Pew found click-through on the classic blue links drops from 15% to 8% when an AI summary is present — but the traffic it does send converts at roughly twice the rate. Smaller, and much better qualified.
Two axes, four positions
Once you measure citing and naming separately, a brand lands in one of four positions, and each one has a different correct answer. This is the diagnosis On Belay produces.
The reason to keep the axes apart is that they genuinely move independently. In our own measured portfolio one retailer sits at high citing with low naming, and another is close to the exact inverse — near-perfect pages that engines almost never open. A single composite score puts those two brands beside each other and prescribes the same work to both. The content sprint that fixes the first would cost the second a quarter to move a number that is already at its ceiling.
From diagnosis to a plan
A verdict is only useful if it names the next piece of work. The two positions point in opposite directions.
The pages losing you the most answers
We can name the individual URLs where an engine opened the page, read it, and still answered without you — ranked by answers lost, not by traffic. That is a specific, fixable extraction failure on a specific page.
The authority levers, ranked by your own data
When engines aren't opening your site, more pages won't help. We rank the external places engines cite instead of you by how many of your answers each appears in while your brand is absent — video, community forums, editorial roundups, review sites, in whatever order your data puts them.
The verdict, the numbers and the list of pages are all computed. No language model decides any of them, because identical data has to produce an identical verdict — in a measurement product, a sentence that rewords itself week to week is indistinguishable from a finding that changed.
Only the brief explaining how to fix a page is written, and every claim it makes about what an engine did has to be traceable to the answer text we stored.
Where to start
If you do nothing else this month, check two things. First, whether your site is readable by the
query-time crawlers at all — OAI-SearchBot, PerplexityBot,
ClaudeBot and Google-Extended are separate from the training crawlers,
and it is entirely possible to allow one while silently blocking the other. Allowing
GPTBot while blocking OAI-SearchBot keeps you in the training data and
removes you from every citation.
Second, find out whether you have a citing problem or a naming problem, because the two most common mistakes in this field are running a content sprint for a brand whose pages already convert, and buying third-party coverage for a brand whose problem is a paragraph.
Sources
- Ahrefs, ChatGPT's most cited pages — 1.4M prompts, February 2026.
- Semrush, 2026 AI Visibility Index — 126M prompts across 22 industries.
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024.
- University of Tokyo & University of Tsukuba, GEO-SFE structural framework, March 2026.
- Similarweb, AI search referral conversion analysis; Pew Research Center, AI summaries and click-through.
- Search Engine Land, Used or cited: the two ways brands appear in AI search.
This post cites its sources in the body text, gives specific figures, and is structured so each section stands on its own. That is not a stylistic preference — it is the three techniques the Princeton study measured as most likely to earn a citation, applied to itself.
See which problem you have
On Belay measures citing and naming separately across five AI assistants, every week, and turns the result into a ranked plan. It is included in every plan — not a per-seat add-on.
Get started