How AI engines decide
what to cite.
Most AEO advice is guesswork. The engines published the documentation. Six things that are actually written down, and one measurement that changes what you fix.
Most advice about getting cited by AI is guesswork. The engines publish how their crawlers work and what governs their answers. Almost nobody selling this reads those pages.
What follows is the documented part, with a link to the primary source on every claim. Where something isn't documented, this guide says so rather than filling the gap with confidence.
One thing to get straight first, because it changes what you measure: being named in an answer and being cited as a source are different events. A model can describe your business accurately and link to somebody else. That happens more than people expect, and it's the gap most reporting hides.
Six things that are actually written down.
What is actually documented
Every claim below links to the company that published it.
There is no single "AI crawler," and blocking one is not blocking another
OpenAI runs four separate crawlers and they do different jobs. GPTBot gathers content for model training. OAI-SearchBot is the one that decides whether you can surface in ChatGPT's search results. ChatGPT-User fetches a page because a person asked, which means robots.txt may not govern it at all.
Google publishes its own list, and Bing publishes theirs.
The common mistake: blocking GPTBot to opt out of training, then wondering why you never appear in ChatGPT. Those are different bots. Read your own robots.txt and check which one you actually disallowed.
Named and cited are separate results, and the split is the finding
We baselined a real estate development firm in August. Ten buyer questions, one engine, every answer stored verbatim, the whole run repeated four times across two endpoints so a single lucky reply couldn't move the number.
Named in 2 of 10 answers. Cited as a source in 1 of 10. The split matters more than either number. Both branded questions hit. All five category questions missed. All three pricing questions missed.
Findable if you already know the name. Invisible if you don't. That is a different problem from "we need more content," and it needs a different fix.
What this means for you: a report that gives you one visibility score has averaged away the only useful distinction. Ask for the branded and unbranded numbers separately.
Structured data helps a machine parse you, and Google says what it does not do
Schema markup is worth adding. It tells a parser which text is a question, which is an answer, which is a price. Google's own introduction is clear about what it is for and what it is not.
It is a parsing aid. It is not a lever that makes a thin page authoritative. If the underlying page doesn't answer the question, marking it up as an answer changes nothing.
The tell: anyone who sells schema markup as the AEO fix, without touching what the page says, is selling the cheap half.
Google publishes what governs its AI features, and it is mostly ordinary
Google's page on AI features is the closest thing to an official answer that exists. Read it before you buy anything described as a proprietary AEO methodology.
Most of what it asks for is unglamorous and already familiar. Be crawlable. Be useful. Be clear about who wrote it.
Worth noticing: there is no published ranking formula for AI answers. Anyone quoting one is quoting themselves.
Part of citation is decided inside the model, not on your website
Anthropic documents a citations feature where the model returns the specific source passages its answer rests on. That decision happens at answer time, over whatever documents were supplied.
You cannot optimise your way into a mechanism running on someone else's servers. What you can do is be the source that is easy to quote: a clear claim, in one place, attributable to a named author.
The useful consequence: write passages that survive being lifted out of context, because that is exactly what happens to them.
One ask proves nothing, because an answer is a distribution
Ask the same question twice and you can get two different answers. Any measurement built on a single query is noise reported as a finding.
Ask each question several times. Record every answer verbatim. Treat a change as real only when it survives repetition. Our baseline above was run four times across two endpoints before a single number went in a document.
Ask your vendor: how many times did you run each query, and can I see the raw answers? If the answer is once, and no, the number is decoration.
What we would check first
Open your robots.txt and find out which crawlers you are actually allowing. Most people have never looked, and a surprising number have blocked the bot that decides whether they can appear in ChatGPT at all.
Then run ten real buyer questions, not your brand name, several times each. Write down how often you are named and how often you are cited. Those two numbers, kept separately, are the honest starting point.
Everything after that is content work, and content work is slow. Knowing which of the two numbers is broken tells you which kind of slow work to do.
Keep reading
The 6-Point Check For AI Work Before It Leaves Your Company
Most AI failures don't look like errors. Here's how to grade output you didn't build, and turn it into an error rate you can actually act on.
Read the guide → Operations · 6 min readThe 30-Minute AI Inventory: What Is Running In Your Business And Who Owns It
Four questions, four groups, one specific order. Surfaces the undocumented setup carrying real volume before the person running it resigns.
Read the guide →