There is no portable benchmark for a good AI citation rate, and the published figures prove it rather than contradict it. MaxAEO's 2026 AI visibility benchmarks reported a median brand mentioned in roughly 31% of relevant non-branded answers across eight engines, with the top quartile clearing 58% and the top decile at 74% or higher, while AirOps' 2026 retrieval study found only 15% of the pages ChatGPT retrieved were cited at all. Those numbers are not measuring the same thing: the denominator is whatever prompt set the publisher chose, and changing the prompts changes the rate more than any site change will. Two comparisons do hold. Your own rate against your own earlier runs on an unchanged question set, and your rate against the named competitors inside that same set. Everything else is somebody else's prompt list.
A good AI citation rate is one that is higher than your own last measurement on the same question set, and competitive with the companies named alongside you in that set. No absolute figure survives contact with a different prompt list, which is why the published benchmarks are not comparable with one another, let alone with your own set.
The reason is structural rather than a flaw in anyone's research. Citation rate is the share of prompts in a set where you are named or cited, so the denominator is a choice. A set built from narrow questions your product genuinely answers will produce a high rate. A set built from broad category questions dominated by review sites will produce a low one. Both are honest measurements of different things.
That makes the number useful in exactly two directions. Compared against itself over time, with the questions held fixed, it reads as a trend. Compared against the other companies appearing in the same answers, it reads as share of voice. Compared against a figure from an industry report, it reads as nothing at all.
What people are usually asking when they ask this question is whether they should be worried. A more answerable version is: of the questions a buyer in our category would ask, how many name us, and who is named instead? That has a specific answer for any company and it takes an afternoon to produce.
Our methodology page keeps citation outcomes in a separate series from the page-level readiness score for the same reason. Folding an outcome rate into a page score makes a week when the engines rotated their sources look identical to a week when a page got worse.
Published AI citation benchmarks disagree because they measure different units over different windows on different engines, and almost none of them share a prompt set. Reading four of them side by side makes the incompatibility obvious, which is more useful than picking one.
| Published figure | What it measured | Source |
|---|---|---|
| Median brand mentioned in roughly 31% of answers, top quartile 58%, top decile 74% | Non-branded prompts across eight engines, rolling 90 days to mid-2026 | MaxAEO, 2026 |
| Median half-life of 11 days from peak | Citation share per page per engine, 883,000 pages | Profound, 2026 |
| Median brand at half its peak in 31 days | Brand-level citation counts, 10,991 brands over 10 months | Trakkr, 2026 |
| 15% of retrieved pages cited | 548,534 pages retrieved across 15,000 ChatGPT prompts | AirOps, 2026 |
The two decay figures in that table are the clearest illustration. Profound's 2026 citation decay study measured citation share per page per engine across 883,000 pages and seven engines over twelve months, producing 1.19 million page-and-engine lifecycles, and found the median page lost half its citation share within 11 days of peaking. Trakkr's "The Half-Life of AI Citations", which tracked 108,650 citation URLs at the URL level and 10,991 brands across 857,000 daily reports over ten months, found the median brand fell to half its peak citation count in 31 days.
Neither is wrong and the gap is not a contradiction. One measures a page losing its share of a category's citations; the other measures a brand losing its total count across all its pages. A brand can hold its count while every individual URL behind it churns, which is exactly what the Trakkr study found: 73.5% of citation URLs appeared once and never returned.
So when a benchmark is quoted at you, the question to ask is what the denominator was. If the answer is a prompt set you have not seen, the figure cannot tell you whether your own rate is good, only that somebody measured something.
A citation rate measures the share of prompts in a defined set where your domain is cited, or where your brand is named, and those two are different metrics that often get the same label. Being named in the text of an answer and having a URL of yours cited as a source are separate events, and a brand can be mentioned in answers that cite nobody's pages at all.
The denominator deserves most of the attention because it is the part under your control. A prompt set is a commercial document: it encodes which buyers you care about, which use cases, which stage of the purchase, and whether competitors are named in the question. Two suppliers measuring the same company with different sets will report different rates and both will be defensible.
The numerator has a subtlety too. AirOps' 2026 retrieval study examined 548,534 pages retrieved across 15,000 ChatGPT prompts and found only 15% of retrieved pages were cited, which means a page can be fetched and read on every run and still never appear. Retrieval and citation are separate stages, and a rate that counts only citations will show a zero for a page that is one sentence away from winning.
There is also the question of what counts as an appearance. Some tools count a mention anywhere in the answer, some only count a linked citation, some count position within the answer, and some weight the first-named company more heavily. None of these is wrong; they are simply not comparable, which is why switching tools mid-programme usually produces an apparent jump or drop that reflects nothing.
The practical upshot is that a citation rate is only meaningful alongside its definition. A report that gives a percentage without stating the prompt set, the engines, the window and whether mentions count is not reporting a measurement.
Movement of tens of percent between runs is normal on a small prompt set, which is why a single month-to-month change is rarely evidence of anything. Trakkr's "The Half-Life of AI Citations" reported that a typical brand's week-over-week citation count swings by 51.8%, across 857,000 daily reports covering 10,991 brands.
Volume is what separates signal from sampling. A brand cited many times across many prompts has a rate that moves slowly, because each individual answer is a small share of the total. A brand cited twice has a rate that halves when one answer changes, and no rewrite caused that.
The practical consequence is that the first thing a measurement programme should establish is its own noise floor, which is what running the baseline twice a week apart is for. If two identical runs a week apart differ by eight points before anything was changed, then an eight-point gain in month two is not a result.
Rotation makes this worse in a specific and predictable way. Trakkr's "The Half-Life of AI Citations" found 73.5% of citation URLs appeared exactly once and did not return, so the set of pages behind a stable brand-level number is churning underneath it. A page dropping out is therefore the normal case rather than a failure, and only a sustained pattern across several runs is worth acting on.
What does count as a real move is direction held across three consecutive runs on an unchanged question set, or a change in which companies are named rather than how often. Our guide to measuring AI citation decay works through the arithmetic, including how to size your own variance before reading a trend.
Compare your citation rate against your own previous runs and against the competitors who appear in the same answers, and stop there. Both comparisons use the same prompt set, which is the only way to hold the denominator still, and both are available from the log you are already keeping.
The time comparison is the one that tells you whether the work is working. Same questions, same engines, same counting rule, read across at least three runs. Anything else introduces a change in measurement at the same time as a change in the site, which makes the result uninterpretable in either direction.
The competitive comparison is the one that tells you whether the rate is good. If you appear in 4 of 12 shortlist answers and one competitor appears in 11, the gap is specific and diagnosable: look at what the answers cited when they named them. If every company in the category appears in roughly a third of answers, a third is what good looks like in that category, and no industry benchmark was needed to learn it.
A third comparison is occasionally worth running and often misread: your rate per engine. The engines retrieve differently, and a company can be strong in ChatGPT and absent from Google AI Overviews because one leans on vendor documentation and the other on pages that already rank. Our guide to which AI engine to optimise for first covers how to decide where that gap matters.
What is not worth doing is building a composite visibility score out of several of these. A single number that blends mention rate, position and sentiment moves for reasons nobody can trace, which defeats the purpose of measuring at all.
Report four numbers to leadership and keep them stable: how many questions in the set name you, how that compares with the previous run, who else is named, and which of your pages the answers cited. Those four fit on a slide and each one is traceable back to a logged run.
Deliberately leave the composite score out. Leadership will ask what good looks like, and the honest answer is that the comparison is against the last run and against named competitors rather than against an industry figure. Giving a benchmark you cannot defend buys one quiet quarter and an awkward conversation later, because the first person to look up a different benchmark will find one that disagrees.
Where an external figure genuinely helps is in explaining variance rather than in setting a target. Telling a board that the median brand's week-over-week citation count swings by 51.8%, as Trakkr's "The Half-Life of AI Citations" measured, is what buys permission to report trends over quarters instead of months. The ConvertMate GEO Benchmark 2026 finding that 83% of AI citations come from pages outside Google's organic top 10 serves the same purpose for the question of why rankings do not predict this.
Pair the four numbers with one sentence on what changed on the site before them. A report that shows movement without naming the change behind it cannot be used to decide what to do next, which is the only reason to produce it.
If you do not yet have the first run to report, a free visibility assessment produces exactly those four numbers across ChatGPT, Claude, Perplexity and Google AI Overviews, with the prompt log, which is the baseline every later comparison needs.
It depends entirely on which prompts are in the denominator, which is why MaxAEO's 2026 benchmark study found a median of 31% and a top decile of 74% on one fixed prompt panel alone. On a set of narrow questions your product directly answers, 20% is low. On a set of broad category questions where review sites dominate the retrieval, 20% is strong. The figure becomes meaningful once you know your own previous rate on the same set and the rate of the competitors appearing beside you.
Citation rate is the share of prompts in your set where you appear. Share of voice is your share of all the companies named across those same answers, so it moves when a competitor gains even if your own appearances are unchanged. Both come off the same run log, and reporting them together is what distinguishes losing ground from a category getting more crowded.
Almost certainly because the new tool counts differently or uses a different prompt set, not because anything happened to your visibility. Tools vary in whether they count unlinked brand mentions, whether they weight position in the answer, which engines they query and how often. A tool change resets the series, so the sensible move is to run both for a period or to treat the switch as a new baseline.
Three consecutive runs moving in the same direction on an unchanged question set is the usual threshold, and the reason is the size of the underlying noise. The Trakkr citation decay study measured a typical brand's week-over-week citation count swinging by 51.8% across 857,000 daily reports, so two points of data cannot distinguish a trend from the normal churn. Sizing your own variance with two baseline runs a week apart is what makes the threshold defensible for your set.
Both, because they fail differently and the gap between them is diagnostic. Profound's 2026 study measured per-page citation share and found a median half-life of 11 days from peak, while Trakkr measured brand-level counts and found a 31-day median. A brand holding its count while its URLs churn is working as intended; a brand losing its count while individual pages hold is usually a coverage problem rather than a decay problem.
A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews and reports how many name you, who is named instead, and which pages the answers cited, with the prompt log included.
Request a free visibility assessment →