The finding
We took every measured run in our archive — 1,984 runs, 139 frozen question sets, six engines, real local markets — and for each question compared the set of businesses one engine named against the set another engine named.
Mean overlap across all engine pairs: about 12%.
| engine pair | overlap |
|---|---|
| Google AI Mode vs ChatGPT | 6.7% |
| ChatGPT vs Perplexity | 7.8% |
| ChatGPT vs Gemini | 8.5% |
| Claude vs Perplexity | 9.5% |
| Google AI Overviews vs Gemini | 10.7% |
| Gemini vs Perplexity | 11.3% |
Same question. Same city. Same week. Roughly seven out of eight names differ.
Why this matters more than it sounds
Most AI-visibility tools report a single number, or report one engine. Both are undermined by this result.
A single-engine report describes one surface. If a vendor tells you where you stand in ChatGPT, they have told you about a system whose answer set overlaps Gemini’s by roughly 8%. Your position on the other five surfaces is not implied, estimated, or inferable from it.
A blended score is worse. Averaging six barely-overlapping populations produces a number that describes none of them. It can move because one engine changed while five did not, and it cannot tell you which. This is the specific reason we report every engine separately and refuse to publish a composite “visibility score.”
It also means partial wins are normal. A business is routinely a top pick on one engine and absent on another, in the same market, in the same week. That is not an error in the measurement or a failure of the business. It is what six independent retrieval systems produce.
The companion finding: none of them hold still
The same archive lets us ask how much each engine changes between repeat runs of the same question.
| engine | repeats its own answer | changes it |
|---|---|---|
| Perplexity | 48.6% | 51.4% |
| Google AI Mode | 44.9% | 55.1% |
| Google AI Overviews | 37.7% | 62.3% |
| ChatGPT | 33.5% | 66.5% |
| Claude | 25.5% | 74.5% |
| Gemini | 23.2% | 76.8% |
Every engine we measured changes more of its answer than it keeps. Perplexity is roughly twice as stable as Gemini or Claude, and even Perplexity is under half.
Two consequences. A single run is not a measurement — it is one draw from a distribution, which is why our panels repeat every high-intent question. And “durable” means something different per surface: a position won on Gemini is close to a coin flip on re-measure, while the same position on Perplexity is meaningfully more likely to persist.
How this was measured
Six engines: ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews and Google AI Mode. Questions are frozen before the panel starts and never changed. Every session is run from a clean, logged-out vantage so nothing is flattered by account history.
For each question we extract the businesses named in the answer text, then compare sets between engines using overlap as a share of the combined set. Business-name extraction was validated before use against a hand-curated roster of nine firms in a market we had scored by hand, and recovered 100% of them.
One honest caveat, because it is the kind of thing we would want disclosed to us. The overlap figure moved from 9.6% to 12.4% when we corrected how business-name variants are collapsed — treating “Dr. Jane Smith” and “Smith Plastic Surgery” as one practice rather than two. The direction and the order of magnitude are stable; the exact decimal is not. Read this as “roughly one in eight,” not as 12.4%.
We are also measuring named businesses in answer text, not clicks, traffic or revenue. It tells you who the engines put in front of a customer asking a buying question. It does not tell you what that was worth.
What we would do with this if it were our business
Measure all six, separately, more than once. Expect to be strong somewhere and absent somewhere else. Treat any vendor’s single “AI visibility score” as a number that has averaged away the only thing you needed to know.