Why isn't one check enough?
Because the same question, asked twice, gives two different answers. AI engines are non-deterministic by design: they generate a fresh set of search queries almost every time, and the businesses they name shift from run to run. Research on repeated prompting has found that only a small fraction of citations stay consistent across three runs of the identical prompt. One check tells you what happened once. It does not tell you what your customers see.
So we don't ask once. Try it — here is one real buying-intent question, run three times on the same engine, minutes apart:
Same engine, same question, same day. Three answers. Now imagine deciding your strategy off the first one. Anonymized from a real panel.
What does Nameworthy measure — the API or the app?
The surface your customers actually use. An API call and the consumer app can return different model versions, different search behavior, and different citations; an API request without live search returns the model's memory, not what a person sees on screen. So we collect buying-intent questions from the real consumer interface. Manual browser runs are ground truth by definition — the same standard behind the published local-search studies this field is built on.
Why measure from a clean, logged-out vantage?
Because measurement shouldn't depend on whose account is asking. Google's own AI Overviews documentation says plainly that “people's interactions with Search and these AI experiences” — what they search, what feedback they give — are used to improve its generative models, with human review in the loop. Signed-in surfaces carry history and personalization; a clean vantage answers the only question that matters: who does the engine name for a stranger? Every run in our logs is collected under that standard, and where an engine offers a memory-off mode, we use it — a discipline we've cross-checked against our own logged-in instruments to confirm the archive measures the market, not the account.
How do you score an answer?
Scoring is pre-registered — defined before the baseline runs, and never reinterpreted mid-engagement. Every run is graded against four fixed events, for your business and for each competitor:
- Mention. Your business is named anywhere in the answer text.
- Recommendation. You appear in a recommended list — and we record your ordinal position.
- Citation. Your URL is linked as a source the engine drew from.
- Sentiment. How the answer treats you: positive, neutral, or negative.
Why report every engine separately?
Because each engine runs on a different source stack, and winning one says nothing about the others. In our panels, any two engines name overlapping businesses only about 12% of the time — 1,984 runs across 139 question sets, with per-pair overlap running from 6.7% to 11.3%. Blend them into one score and you hide the exact thing you're paying to see. So we never blend. Every fix we prescribe is aimed at a named engine.
Why these six engines, and not ten?
Some tools advertise coverage of ten or more engines. We measure six on purpose: the surfaces where local, commercial-intent questions actually get asked and answered — ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, and Google AI Mode — each measured in the consumer app, three runs per question. A wider engine list measured shallowly, through APIs that don't match what customers see, is a logo wall, not a measurement. When another surface earns real commercial-intent share, we add it the same way we measure the rest: app, repeated runs, per-engine report. Until then we'd rather be right about six than decorative about twelve.
What's the statistics behind it?
A citation rate is a proportion, so we treat it like one. We report the rate across runs with a confidence interval, not a bare percentage. Repeated-generation measurement with panel-level means is now the published standard for evaluating language models; the variance research is clear that single-run estimates swing by several points even at temperature zero, and that prompt wording is a dominant source of error. That last finding is why the panel is frozen: we never edit a prompt mid-engagement, because changing the words changes the number.
Read next
The research — how AI engines actually source local-business recommendations.
Ready to see yours
Get your free report — your customers' real questions, the major engines, 48 hours.