How AI Assistants Decide Who to Cite

Not ranking. Retrieval, corroboration and whether a passage is safe to quote. The practical consequences are more specific than most advice admits.

· 7 min read

Worth saying up front: nobody outside these companies knows the exact mechanics, and anyone claiming otherwise is guessing with confidence. What follows is what can be observed from behaviour and from how retrieval systems generally work. Treat it as a working model, not gospel.

With that caveat, the patterns are consistent enough to act on.

It is retrieval, not ranking

A search engine ranks documents against a query. An assistant answering a question does something different: it retrieves candidate passages, then synthesises an answer from them and attributes what it used.

The unit is the passage, not the page. This has a consequence people consistently miss: a page can rank first and never be cited, because nothing in it is a self-contained quotable chunk. Conversely a modest page with one crisp, specific paragraph answering exactly the question can get lifted repeatedly.

Three things appear to decide it

Can it read you at all. The unglamorous prerequisite. GPTBot, ClaudeBot and PerplexityBot largely do not execute JavaScript. If your content is painted client side, there is nothing to retrieve. This is the most common cause of zero citations and it has nothing to do with content quality.

Can it identify you. Models are noticeably more willing to name an entity they can resolve and corroborate. Consistent naming, structured data that says who you are, and third-party sources that agree. Ambiguity is a reason not to attach your name to a claim. For professionals this cuts both ways: individual people are entities too, and a named expert with a consistent footprint is easier to cite than a faceless brand.

Is the passage safe to quote. Specific, self-contained, and falsifiable. A sentence that survives being lifted out of its page gets lifted. A sentence that only makes sense after three paragraphs of setup does not.

What that means you should actually do

Answer the question in the first paragraph under the heading. Then elaborate. The inverted pyramid, which journalism worked out a century ago, happens to be exactly what a retrieval system wants.

Make headings match real questions. Not keyword phrases — the way someone phrases it in a chat box, which is longer and more conversational than what they would type into a search bar.

Be specific enough to be wrong. “Results vary by industry” is unquotable because it says nothing. “In legal, clicks run $50 to $300, so a modest conversion-rate problem is ruinous” is quotable because it commits to something. The willingness to be pinned down is what makes a source useful.

Attribute your own claims. Where a number comes from, when it was measured, what the sample was. Models appear to prefer sources that show their working, which is also just good practice.

Keep your entity consistent. Same organisation name, same person names, same descriptions across your site, your profiles and anywhere else you appear. This is tedious and it is a large share of the durable gain.

What does not appear to work

Keyword stuffing for models. The retrieval layer is semantic. Repeating a phrase does not increase the odds of being retrieved and makes the passage worse to quote.

llms.txt as a ranking play. It is worth having — this site publishes one — because it costs nothing and makes the site easier to understand. It is not a ranking factor and I would be wary of anyone selling it as one.

Volume. Publishing more mediocre pages does not help. The retrieval unit is a passage; a hundred pages with no quotable passage in any of them is a hundred pages of nothing.

Trying to game it. These systems change frequently and are actively adversarial to manipulation. The things that work are the things that would work anyway if you were writing to be genuinely useful, which is a boring conclusion but appears to be the true one.

How to find out where you stand

Ask the models the questions your buyers ask, and record what comes back. Not keyword phrases: actual questions, phrased the way a person would.

Do it on a schedule, because model behaviour shifts and a single snapshot tells you nothing about direction. Record who gets cited instead of you — that competitor set is usually the most actionable output of the whole exercise.

The free AI visibility scan does a version of this. AI visibility consulting is the full version, with monitoring and a competitor benchmark. If it turns out the problem is that crawlers cannot read your site at all, that is a technical SEO job first and everything else is premature.

Where this leads

Related services

AI Visibility

Get cited by ChatGPT and AI Overviews

SEO

Dominate search rankings

Want this looked at properly?

Book a free discovery call. I'll come having already looked at your site, and you keep the findings either way.

Book a Free Discovery Call