Industry report

The AI Citation Durability Report (2026): how long an AI recommendation actually lasts

Everyone in AI visibility talks about winning the answer. Almost nobody measures what happens next. We re-ran the same frozen questions on consecutive days across six engines, hundreds of times, and counted how often yesterday's recommendation survived to today. The expected lifetime of an unmanaged #1 is roughly one to two days. This is the first edition of a benchmark we'll keep updating as the archive grows.

Updated August 9, 2026Reading time 7 min

The question nobody was measuring

The AI-visibility category has settled on a comfortable story: get recommended by ChatGPT, and the customers follow. What the category hasn’t published is the follow-up question — once an engine recommends a business, how long does that last?

Nameworthy measured 383 first-position transition pairs to produce this benchmark.

We could answer it because our measurement protocol accidentally builds the dataset: frozen questions, re-run on consecutive days, across six engines, logged every time. Pull every pair of runs where the same frozen question hit the same engine on two different days, and you can count survival directly: was yesterday’s first-named business still first today? Still present at all?

This first edition draws on 383 first-position transition pairs and over 1,800 any-presence pairs from our panel archive — multiple verticals, multiple metros, every run logged. The businesses measured were unmanaged: nobody was defending these positions. That’s the point. This is the natural decay rate — the “before” against which any managed program should be judged, including ours.

The headline number

An unmanaged #1 AI recommendation has an expected lifetime of roughly one to two days.

Per engine, the probability that the first-named business on a question is still first when the same question is re-run a day or two later:

Engine First-position hold rate (cross-day) Expected lifetime at #1
Perplexity 52% (45/87) ~2.1 days
ChatGPT 50% (21/42) ~2.0 days
Gemini 37% (45/122) ~1.6 days
Google AI Mode 37% (32/86) ~1.6 days
Google AI Overviews 16% (4/25) ~1.2 days
Claude 14% (3/21)* see note

* Claude’s sample is the smallest and spans a mode split (local-places answers vs. web-research answers) this edition can’t fully separate from churn — treat it as a floor, not a point estimate. Expected lifetimes are geometric extrapolations from day-scale gaps: directional, not decimal-precise. We publish the raw fractions so you can check the arithmetic.

Even presence — being named anywhere in the answer, not just first — is unstable: any-presence hold rates ran from 61% (Perplexity, the stickiest surface we measured) down to the thirties and forties on Google’s AI surfaces.

Read the table honestly: on the majority of surfaces, yesterday’s top recommendation is more likely than not to be someone else today. If a vendor shows you a single-day screenshot as proof of a “ranking,” you’re looking at a coin flip, not an asset.

Three independent checks that say the same thing

A number this stark needs corroboration, and it has it — one external, two internal:

  • SISTRIX (external), analyzing 82,619 qualified prompts and 1,548,213 snapshots over 17 weeks (17 December 2025 to 8 April 2026), found ChatGPT replaces up to 74% of its cited sources every week. Week-scale source rotation is exactly what day-scale position churn compounds into.
  • Our paired re-run study (internal): across paired runs of identical frozen questions, the first-named business changed 52.2% of the time.
  • Our AI Overviews competitor tracking (internal): the set of competitors named on a question changed 51% cross-day — the churn isn’t just position one; the whole slate reshuffles.

Three instruments, three datasets, one conclusion: AI answers are re-drawn, not maintained.

What decays slower: the seven-day anchor

One longitudinal case in the archive is worth the whole table. A panel subject — a retailer with a published, corroborated, exactly-quotable claim on its site — held #1 on the exact-match question at day seven, while in the same seven-day window losing a generic question’s #1 down to #4. Zero competitor action in either case. Same business, same site, same week — the position anchored to a concrete published fact held; the position held by generic relevance didn’t.

That’s a single case, so we hold it lightly. But it points where the durability program goes: positions are not equally perishable, and the published, verifiable, liftable fact is the most durable substrate we’ve measured. It’s also consistent with everything else our field studies keep finding — engines hold on to what they can quote and verify.

What this means if you run a business

Monitoring is not maintenance. A dashboard that tells you weekly that your citation vanished is a subscription to bad news. The number that matters isn’t “were you recommended once” — it’s weeks held: how long a won position survives, per engine, against the one-to-two-day unmanaged baseline above. Holding requires somebody to notice the decay event and re-harden the source that slipped — that month, not next quarter.

And no one can promise you a position. This data is also why we refuse to guarantee rankings and put that refusal in our pricing: any specific position is a probabilistic outcome that churns daily. What can be honestly sold is measurement, the fix work, the re-hardening loop — and the durability record that proves whether it’s working.

Method, limitations, and what the next edition adds

Method. Every pair consists of the same frozen question, on the same engine, run on two distinct dates (median gap 1–2 days), from clean measurement vantages, drawn from our panel archive. States are scored as first-named / named / absent; a hold means the state didn’t degrade. Ongoing streaks are right-censored. Engines are reported separately, never blended.

Limitations we know about. Business-name matching is automated and imperfect, which inflates apparent churn modestly. Day-scale gaps extrapolated to lifetimes assume the decay is memoryless. Claude’s mode split needs conditioning. All of this is fixable and disclosed — the numbers here are honest fractions of logged runs, not projections.

Next edition adds true week-scale holds from monthly re-measurement cycles, hand-checked entity resolution, and — as managed data accumulates — the comparison this report exists to enable: weeks-held under management vs. the unmanaged baseline. This page will be updated in place, with the revision dated.

Common questions

How long does an AI recommendation last?
In Nameworthy's cross-panel measurement (383 first-position transition pairs, July 2026), an unmanaged #1 recommendation had an expected lifetime of roughly one to two days. Cross-day hold rates ranged from 16% on Google AI Overviews to 50% on ChatGPT and 52% on Perplexity, with Claude lower still on a small, caveated sample — meaning on most engines, yesterday's top pick is more likely than not to be different today.
Why do AI answers change so much run to run?
The engines re-retrieve and re-select sources continuously rather than maintaining a fixed ranking. [SISTRIX](https://www.sistrix.com/blog/ai-citation-drift-how-stable-are-sources-in-ai-search-results/) found ChatGPT replaces up to 74% of its cited sources every week across 82,619 prompts; our own paired re-runs show the first-named business changing 52% of the time on identical questions. Selection is probabilistic — which is why single-pass measurement misleads and repeated runs are the only honest instrument.
Does anything hold its position longer?
Yes — in our longitudinal measurement, positions anchored to a published, corroborated, exactly-quotable fact decayed measurably slower than generic-question positions held by the same business in the same window. Nothing held unassisted indefinitely, but concrete published facts are the most durable substrate we've measured.

Your baseline

Find out how long your positions are lasting.

The free Snapshot shows where you stand today. The retainer is how a position measured in days becomes a streak measured in months.

Get your AI visibility report