Ranks fine on Google.
Invisible to AI.
What an answer engine actually says about a business with healthy rankings. One real baseline, written out in full, including the parts that do not generalise.
A real estate development firm came to us ranking perfectly well on Google. We asked an answer engine ten questions their buyers actually ask. It named them twice and cited them once.
This is that run, in full. The numbers, the method, and the parts that do not generalise. It is the same baseline referenced in how AI engines decide what to cite, written out properly.
One framing to get straight before the numbers, because it decides what you would fix: ranking and being named are not the same event, and neither is being cited. This business was doing well at the first and badly at the other two, and nothing in their analytics would have told them.
The split is the whole finding
Ten questions. Two named. One cited. Where those hits landed is the part that matters.
| Question type | Asked | Result |
|---|---|---|
| Branded — "who is [company]?" | 2 | Named both times |
| Category — "best [service] in [city]" | 5 | Absent, all five |
| Long-tail — "what does [service] cost?" | 3 | Absent, all three |
Findable when you already know the name. Invisible when you don't.
That is a business with no search discovery at all. Every customer an answer engine sends them is one they already had. A single averaged "visibility score" would have reported something around 20% and hidden the entire finding, because the only queries that hit were the ones that prove nothing.
What to ask for: branded and unbranded numbers, reported separately. Any vendor who gives you one blended score has averaged away the useful half.
Three things the run surfaced
None of these were visible in analytics, rankings, or traffic.
The one page an engine cited was the wrong page
The single citation in the entire run pointed at a blog post. The post is good. It attracts a completely different audience than the one that hires them, and the service pages written to convert a buyer were never reached at all.
Worth checking on your own site: which page gets cited, not just whether one does. A citation that brings the wrong reader is a vanity metric.
When an engine cannot identify you, it guesses
Asked who the company was, the engine pulled sources from four unrelated businesses with similar names, including the Wikipedia page for a venture capital firm, and hedged its answer with "appears to be."
That is not a ranking problem and more content would not have fixed it. Nothing on the open web corroborated that the entity existed, so the model reached for the nearest thing that looked like it.
The tell: if an engine hedges when describing you, the gap is identity, not authority.
Three buying questions had no winner at all
A pricing question, a materials-comparison question, and a how-to-choose-a-contractor question. Zero local businesses named on any of them. The engine fell back to generic aggregators and out-of-market blogs.
Nobody owns those questions. They need no domain authority to win, because there is no incumbent to displace. They are the cheapest citation available and almost nobody builds for them, because everyone is busy competing for the question that already has a winner.
The move: find the questions in your category that no one answers, before you fight for the ones everybody does.
Why the engine could not quote them
A separate audit the same morning scored the site 26 out of 100. Grade F.
| Dimension | Score | What it means |
|---|---|---|
| Reachable | 13 / 15 | Crawlers get in fine. Not the problem. |
| Identity | 4 / 25 | Bare Organization schema. No address, phone, geo or service area. |
| Questions | 5 / 25 | No FAQPage schema anywhere on the site. |
| Answerable | 0 / 20 | Median page 210 words. 12 of 17 pages thin. Eight pages are a headline and nothing else. |
| Entity | 4 / 15 | No sameAs links. Nothing on the web corroborates the company exists. |
The most useful number there is Reachable at 13 out of 15. The crawlers were never blocked. The site was available and legible the whole time. There was simply nothing on it long enough or specific enough to lift into an answer.
That matters because the crawler-access layer is what most AEO tooling checks first, and it is the layer this business had already passed. One competitor kept getting named for work outside its specialty. It is not the better operator. It published a page that answered the question.
The method, and what it does not measure
Stated first rather than buried, because it is the part that decides whether any of the above is worth anything.
A vendor API, not the consumer app a buyer opens. Those run different retrieval stacks, and some surfaces expose no API at all. What this measures is a close proxy, measured identically every month, which is what makes a trend line honest. It is not a claim about what any individual buyer saw.
An answer is a distribution, not a fact. The same question returns different answers on different days. Every number above is the result of repeated sampling: four runs across two endpoints, with every response stored verbatim. That is the reason we trust it, and it is why a single query tells you almost nothing.
The tracker records which model actually answered and warns if that changes between runs, because a vendor can re-point a preset at any time. If next month's number moves, the first question is whether the site changed or the instrument did.
A score is not a measurement
While building the monthly tracker, a commercial tool in this category rendered us a clean report headlined "0 of 100 on the first run. This is the baseline."
Nothing had been measured. Every query had errored out on a rate limit. Zero calls completed, zero dollars spent, no responses stored. The tool reported a score anyway.
The dangerous part was that the fake zero agreed with the real baseline, which made it more believable rather than less. A number that looks like a finding and happens to match what you expected is the hardest kind of wrong to catch.
Three questions for anyone selling you a visibility score: how many times did you ask, which engine answered, and what happens to the number when a query fails?
Keep reading
The 6-Point Check For AI Work Before It Leaves Your Company
Most AI failures don't look like errors. Here's how to grade output you didn't build, and turn it into an error rate you can actually act on.
Read the guide → Operations · 6 min readThe 30-Minute AI Inventory: What Is Running In Your Business And Who Owns It
Four questions, four groups, one specific order. Surfaces the undocumented setup carrying real volume before the person running it resigns.
Read the guide →