AI visibility is not a single score — it is share of voice across four engines (ChatGPT, Google AI Overview, Perplexity, Claude), measured per buyer query, and the four numbers rarely agree. Tracking it means running the questions your buyers actually ask, on each engine's default settings, multiple times, and recording which firms get named and how often. A single run is unreliable because the same query returns a partly different answer next time — which is exactly why share of voice is measured across repeated runs, not read from one response.
AI visibility is how often, and how prominently, an AI engine names your company when a buyer asks a question your company could answer. It is the AI-era equivalent of being on the shortlist — except the shortlist is now assembled by the engine before the buyer contacts anyone, and there is no second page to climb to.
The unit that makes it measurable is share of voice: across repeated runs of a given buyer query, the percentage of those runs in which an engine names your firm. A firm named in four of five runs has 80% share of voice on that query and engine; a firm named in zero has 0%, a position no competitor owns and one the firm could claim.
The four engines read from different indexes and weight sources differently, so a firm dominant on one is frequently absent from another. This is not noise — it is structural and stable across repeated runs. In a SIGNALS study of 27 buyer queries, no firm held a strong position across all four engines at once in 20 of them (SIGNALS, 2026). A firm can lead Google, Perplexity, and Claude and still be missing from ChatGPT entirely.
The consequence for measurement: a single blended "AI visibility score" hides the thing that matters. You need the number per engine, because the gap is almost always engine-specific — and the fix usually is too.
Start from the questions a buyer types when sourcing what you sell — not your product names, but the buyer's phrasing: "best [category] for [application]," "who builds [thing]," specialist and standard-specific variants. Specialist phrasings and generic phrasings return very different rosters, so include both to see where you can compete and where the category consolidates around large incumbents. This query set is the measurement surface; everything else is run against it.
Four axes are usually enough to build the set, and the difference between them is not cosmetic. Take one category from the SIGNALS pharmaceutical process skid study as a worked example. The buyer-phrased query "best CIP skid fabricators for pharmaceutical manufacturing" left no firm above 55% in that study, which is an open field. The same requirement worded the way a generalist would say it, "leading automated CIP system vendors", consolidated on GEA at 100% on every engine. Add a region and the roster changes again: "pharmaceutical process skid fabricators in Texas" surfaced Mako Industries at 60% and A&G Piping Services at 50%, neither of which appears in the national answers at all (SIGNALS, 2026).
That is one requirement, three phrasings, three different markets. A query set built from only one of those axes will look stable and will be measuring a slice of the demand. So take the buyer phrasing, the generalist phrasing, at least one standard or specification-led variant, and the regional form of each query that matters, and keep them as separate rows rather than averaging them into a category score. The averaging is what hides the finding.
Size the set by what it has to decide rather than by coverage. Ten queries that sit in front of real deals, measured repeatedly and held steady for a year, are worth more than a hundred measured once, because the value of this work is in the series and a wide single pass is not part of a series. Queries get added when the business enters a new category, not when someone thinks of one.
Run every query on ChatGPT, Google AI Overview, Perplexity, and Claude, using each engine's default consumer configuration — because that is what a real buyer experiences. Capture the full response, including the sources each engine cites. Note that public APIs do not reliably reproduce the consumer surfaces; the answer a buyer sees comes from the default interface, so that is what has to be measured.
A single run is unreliable: the same query, run again, returns a partly different roster. Run each query several times per engine and measure share of voice as the percentage of runs in which a firm is named. This is what turns a noisy one-off answer into a stable, comparable number — and it is why credible AI visibility measurement is repeated, not snapshot.
For each query and engine, record which firms were named and how often (their share of voice), and capture which sources the engine cited. The citation source matters. In the SIGNALS pharmaceutical process skid study, between 73% and 89% of citations pointed to firms' own pages rather than to directories or listicles, and the band held across six separate categories (SIGNALS, 2026). That tells you the lever for improving the number is the firm's own content rather than third-party placement.
A number is only useful against a baseline and a competitor set. Establish where you stand today across the four engines, identify the queries where a competitor is named and you are not, then re-measure after you change anything. AI citation takes two to six weeks to shift after structural changes, so monthly re-measurement during active work shows what is actually moving — and which engine recovered.
You measure AI citation decay by re-running the same query set on the same engines at a fixed interval, then comparing how often a given page is still cited from one interval to the next. Decay is a slope across several measurements, never a reading taken once. A page cited in half the runs of a query in March and in a fifth of them in June has decayed, and the only way to know that is to have measured in March. This is the one part of AI visibility measurement that cannot be reconstructed after the fact: without a baseline, a low number today is indistinguishable from a number that was always low.
Citation decay is the gradual loss of citations by a page that engines were previously citing. Two other patterns get mistaken for it. Run to run variance is the difference between two runs of the same query minutes apart, which is why share of voice is measured across repeated runs rather than read from a single response. A citation drop is a step change you can date to one interval, usually with a technical cause behind it. Decay is neither of those. It is a page quietly losing ground while nothing visibly breaks, which is what makes it easy to miss and expensive to leave alone.
The three patterns look similar in a spreadsheet and ask for different responses. Separating them is most of the work.
| What you observe | What it usually is | How you confirm it | What it asks you to do |
|---|---|---|---|
| Cited in one run, absent in the next, cited again in the third | Run to run variance | It shows up inside a single measurement window, across repeated runs of the same query | Nothing to the page. Raise the number of runs per engine until the number settles |
| Citation share falls by a little at every measurement | Citation decay | A downward slope across three or more consecutive measurements on unchanged queries | Refresh the page itself, then re-measure after the lag rather than the next day |
| Citation share falls to near zero at one measurement and stays there | A citation drop | A step change you can date to a single interval | Look for a cause at that date: a blocked crawler, a changed URL, a removed page |
| Your share falls on one engine while a competitor's rises on the same query | Displacement | The roster changed and the sources the engine cites changed with it | Read the page the engine now cites, then close that gap on your own page |
Citation half-life is the interval over which a page's citation share falls by half, and you calculate it from your own measurement history rather than from a published benchmark. Take the page's share of voice at its highest measured point, find the later measurement where it has fallen to half of that, and the gap between those two dates is the half-life. Two conditions have to hold for the number to mean anything. The query wording and the engine settings must be identical between measurements, and every measurement must use the same number of runs per engine. Change either one and you are measuring your own method rather than the page.
The series below is illustrative arithmetic rather than measured data. It exists to show the shape you are looking for and how the half-life falls out of it.
| Measurement | Runs citing the page | Citation share | Reading |
|---|---|---|---|
| March | 16 of 20 | 80% | Peak. This is the baseline everything after is measured against |
| April | 13 of 20 | 65% | A single fall. Not yet a slope |
| May | 11 of 20 | 55% | Second consecutive fall. Now it is a trend, not variance |
| June | 8 of 20 | 40% | Half of the March peak. Half-life is roughly three months |
| July | 6 of 20 | 30% | Still falling. The page is refreshed this month |
| August | 9 of 20 | 45% | Recovery shows up one measurement after the refresh, not immediately |
Read that series as a curve rather than as six separate readings. The page did not break in June. It had been losing ground since March, and the refresh in July surfaces in the August measurement rather than the same week, which is the two to six week lag described in step five. A monthly cadence catches a decline like this on the second reading. A quarterly cadence would have shown the March number, then the June number, and made a slow slide look like a sudden collapse.
The practical rule that follows: measure at an interval shorter than the half-life you are trying to detect, and refresh the page before the curve reaches its floor. A page that has already fallen to zero has to earn its position back from nothing, which is the expensive version of the same job. For the underlying scoring, see how SIGNALS measures citation readiness and the seven SIGNALS dimensions.
Detecting decay reliably takes about twenty runs per engine per measurement, repeated at roughly monthly intervals. The two numbers are linked, and most measurement that fails fails on one of them. Runs per measurement decide how small a change the number is physically able to show. The interval decides how much real movement has happened between one reading and the next.
Start with the runs. Share of voice measured across five runs can only come out as 0, 20, 40, 60, 80 or 100 percent, because those are the only values five runs can produce. A page whose citation share has genuinely slid from 80 to 65 percent has nowhere to report that. It will read 80 or it will read 60, and which one you get is close to a coin toss. Resolution is only half the problem. Every measurement is a sample, so it carries sampling error even when nothing about the page or the engine has changed at all.
The table below is arithmetic rather than a SIGNALS measurement. It takes a page whose true citation share is 60 percent and shows how far a single reading can sit from that figure by chance alone, at different numbers of runs.
| Runs per engine | Smallest change the number can show | How far one reading can sit from the truth | What a single measurement is good for |
|---|---|---|---|
| 5 | 20 points | About 43 points either way | Establishing that the page is cited at all. Nothing finer than that |
| 10 | 10 points | About 30 points either way | Presence or absence on an engine, not movement |
| 20 | 5 points | About 21 points either way | A workable baseline, read as a slope across several measurements |
| 50 | 2 points | About 14 points either way | Movement between two measurements starts to be readable on its own |
| 100 | 1 point | About 10 points either way | A single measurement carries a number you can quote |
Two things follow from that table, and the second one is the useful one. The margin on any single reading is wider than the change you are usually hunting for, so no one measurement can tell you a page is decaying. But if each reading were pure noise around a flat line, a fall would be as likely as a rise, and three consecutive falls would happen about one time in eight. Four in a row happens about one time in sixteen. That is why the slope is trustworthy while the individual readings are not, and why the decay table above asks for three or more consecutive measurements before you call it.
Published research agrees on the direction, and it reached the conclusion from data rather than from arithmetic. The argument above is a sampling argument made from first principles, so it is worth saying where it sits against work that measured the engines directly. Schulte and colleagues at the University of St. Gallen set out the case in the April 2026 arXiv preprint Don't Measure Once: Measuring Visibility in AI Search (GEO), whose finding is that visibility has to be treated as a distribution rather than a reading, because a single observation does not describe it.
The measurement that makes this concrete is source churn. The same study reported the sources cited on consecutive days overlapping by only 34% to 42% (Schulte et al., arXiv:2604.07585). A roster that turns over by more than half between one day and the next is the reason a single pass tells you almost nothing, and it is the same effect the table above describes as sampling error, observed from the outside instead of derived. It also puts a floor under the interval: re-run a query set a week apart and most of what moved is this churn, not the page.
Our own reading of that paper's run-count analysis, and what it implies for how often a report should be produced, is set out on how to automate AI visibility reporting. The short version is that the published minimum is lower than the twenty runs above, because it is answering a narrower question. Seven to ten runs is enough to say with confidence whether a brand appears at all. Twenty is what you need before a change in share of voice between two measurements is readable as a change rather than as noise, which is the harder problem and the one decay work actually poses.
Two caveats, because this is a young literature. The work is a preprint rather than peer-reviewed, and it measured a set of engines at one point in time, on defaults that the vendors change without announcing. So treat the numbers as the best published evidence rather than as constants. What the caveats do not touch is the finding itself. One run is not a measurement.
Re-measure monthly for most pages. The interval has a floor and a ceiling. The floor comes from how long the engines take to reflect a change, which this page puts at two to six weeks: a query set re-run seven days after the last round is mostly re-measuring run to run variance, and a query set re-run days after you edited a page is reading the lag rather than the result. The ceiling comes from the half-life you are trying to catch. Measure quarterly and a page with a three month half-life hands you its peak and then its floor with nothing in between, which reads as a collapse rather than the slide it actually was. Monthly sits between the floor and the ceiling for most pages, and it is the cadence the worked example above assumes.
The cost is the reason most firms have no such series. Twenty runs, four engines, every query, every month is real work, and there is no shortcut that leaves the answer intact. What you can do is narrow the query set. A decay series on ten queries that decide deals is worth more than one pass over a hundred that do not, because the series is the thing you are building and a single pass is not part of it.
When a page starts decaying, read the sources the engine cites now, before changing anything on your own page. A page that is losing citations while its own content sits unchanged is usually losing them to something else that changed: a newer page answering the query more directly, a competitor's page that carries the specific figure the engine wants, or a shift in what the engine retrieves for that wording. The citation list captured in every run is the evidence for which of those it was, which is the reason step four records citations and not just names.
One caveat worth holding on to. Not every decaying page is worth defending. If the query itself has lost volume, or the engines have started answering it without citing anyone, the page can be doing nothing wrong while its number falls. That is why the citation list matters more than the share of voice figure: a roster that changed tells you that you were displaced, and a roster that shrank for everybody tells you the query changed underneath all of you.
SIGNALS tracks this with PULSE. It runs your buyer queries across ChatGPT, Google AI Overview, Perplexity, and Claude on default settings, captures the full responses, and reports your share of voice query by query and engine by engine — against the competitors winning your category. It is built for exactly the behavior above: four numbers, measured across repeated runs, not one blended score from a single pass.
The percentage of runs of a given buyer query in which an AI engine names your company. Measured per query and per engine across repeated runs, it quantifies how often you appear in the answers buyers actually receive.
You can report it in one place, but it cannot honestly be one number. Visibility differs by engine because each reads from a different index — so a single blended score hides the engine-specific gaps that are the whole point of measuring. Track four numbers, one per engine.
AI engines introduce variation between runs, so a single query returns a partly different roster each time. This is why share of voice is measured across multiple runs rather than read from one response — repetition is what makes the number reliable.
APIs do not reliably reproduce the default consumer surfaces buyers actually see, so they are a poor proxy for real visibility. Measurement has to reflect the default interface a buyer uses.
Monthly during active optimization, since AI citation takes two to six weeks to shift after changes; quarterly once stable. Always re-measure immediately after structural changes to confirm they registered.
AI citation decay is the gradual loss of citations by a page that answer engines were previously citing. It differs from run to run variance, which is the difference between two runs of the same query minutes apart, and from a citation drop, which is a step change you can date to a single interval. Decay is a downward slope across consecutive measurements while nothing visibly breaks.
Re-run the same query set on the same engines at a fixed interval, keeping the query wording, the engine settings, and the number of runs per engine identical each time, then compare how often each page is still cited from one interval to the next. Decay is the slope across those measurements rather than a single reading, so it can only be detected against a baseline recorded earlier. The worked method is set out above.
Citation half-life is the interval over which a page's citation share falls by half. You calculate it from your own measurement history: take the page's share of voice at its highest measured point, find the later measurement where it has fallen to half of that, and the gap between those two dates is the half-life. Measure at an interval shorter than the half-life you are trying to detect, or the curve will look like a collapse rather than a slide.
About twenty runs per engine. Five runs can only produce the values 0, 20, 40, 60, 80 or 100 percent, so a slide from 80 to 65 percent has nowhere to register. Twenty runs give five-point resolution, and even then one reading can sit around twenty points either side of the truth by chance, which is why decay is read as a slope. The arithmetic is set out above.
Monthly for most pages. Engines take two to six weeks to reflect a change, so re-measuring seven days later mostly captures run to run variance. Measure quarterly and a page with a three month half-life shows you only its peak and its floor.
Read the sources the engine cites now, before you change your own page. Rule out a citation drop first, since that is a step change with a technical cause such as a blocked crawler. Then close the specific gap the comparison identifies and re-measure after the lag.
Request a free PULSE visibility assessment. It establishes your share of voice across all four engines, query by query, against your named competitors — the baseline everything else is measured against.
A free, PULSE-powered visibility assessment maps exactly where you're cited and where you're invisible across ChatGPT, Google AI Overview, Perplexity, and Claude — query by query, against your competitors.
Request a free visibility assessment →