SIGNALS
Measurement

How do you know if your AEO is actually working?

The short version

AEO is working when citation share on a fixed prompt set moves in the same direction across three consecutive runs, against the same competitors, on prompts that have a buyer behind them. A single better week is not evidence. BrightEdge found that domains cited frequently moved about 0.7% week to week while sporadically cited domains swung by more than 50%, and put the stability threshold near 50 citations, where weekly volatility drops from roughly 50% to 8%. Below that threshold your own measurement noise is larger than any effect your agency could produce in a month, which is why the baseline has to be run twice before anybody changes a page.

How do you know if your AEO is actually working?

You know AEO is working when citation share on an unchanged prompt set moves in the same direction across three consecutive runs, measured against the same named competitors, on prompts a buyer would actually type. Every clause in that sentence is doing work. The set has to stay unchanged because swapping prompts resets the series, and three runs are the minimum because two points cannot be told apart from noise. Competitors have to be named because absolute citation share depends mostly on how easy your own prompt list is.

What will not tell you is any single reading. Ask an engine the same question twice in an hour and the sources differ, which is a property of how these systems sample rather than a change in your standing. A vendor who sends a screenshot of a good answer is sending you one sample from a distribution, and the same prompt an hour later may not contain you at all.

The second test is whether the prompts that moved are the ones that matter. Citation share averaged across 60 prompts can improve while every prompt with commercial intent behind it stays exactly as it was, because informational questions are easier to win. Split the report by intent and look at the comparison and recommendation prompts on their own, since those are the ones with a buyer at the other end.

The third test is whether anything on your pages actually changed. Programmes that report movement without a list of specific URLs edited in that period are reporting the weather. Ask for the change log next to the citation numbers, and the two together tell you whether you are watching cause and effect or watching variance with a narrative attached.

Why does your citation count change so much week to week?

Your citation count changes week to week because AI engines sample rather than rank once and hold, and because most brands sit below the volume at which the average stabilises. The size of the swing is now measured rather than anecdotal, and it is large enough that it dominates everything else in the first months of a programme.

BrightEdge's citation volatility analysis, run across ChatGPT, Perplexity, Google AI Overviews and AI Mode, found that domains cited frequently moved about 0.7% week to week while sporadically cited domains swung by more than 50%, a 70-fold gap that held on every platform. It put the stability inflection near 50 citations, where weekly volatility falls from roughly 50% to 8%, and recorded different baselines by engine, with Perplexity at 72.8% and Google AI Overviews at 41.8%.

Individual URLs behave worse than brands do. Trakkr Research logged 108,650 citation URLs across 10,991 brands and eight tracked models over a ten-month window and found that 73.5% of cited URLs appeared exactly once and never returned, with the median brand falling to half its peak citation count in 31 days. A page that won a citation last month is not holding it by default.

Repetition within a single day varies too, which catches people who assume a same-day check is a controlled comparison. Schulte, Bleeker and Kaufmann of the University of St. Gallen ran identical prompts several times on the same day and found cited sources overlapped in only 32% to 43% of cases, with around 65% of cited sources turning over from one day to the next. Our page on why AI answers change every time you ask covers the mechanism behind it.

What should a monthly AEO report show you?

A monthly AEO report should show four lines and the threshold that makes each one meaningful, because a number without a threshold is an invitation to read noise as progress. The lines below are what a report has to carry for somebody outside the work to judge it. Anything else in the document is context.

Line Question it answers What counts as a real move
Citation share, fixed prompt set Are we named when buyers ask Same direction across three consecutive runs
Share against named competitors Are we gaining or losing relative position Our share up while a named rival's is flat or down
Prompts won, split by intent Did we gain on questions with a buyer behind them Movement on comparison and recommendation prompts
Pages changed this period What was actually done Specific URLs, with dates, next to the numbers

Variance belongs in the report as a stated figure rather than as a caveat in the appendix. Two identical baseline runs a week apart give you your own noise floor, and a report that shows this month's change next to that floor is honest in a way that a bare percentage is not. Our guide to measuring AI citation decay sets out the arithmetic for half-life and the one-and-done rate, which belong in a quarterly view rather than a monthly one.

Engine breakdown matters more than a blended score. Averaging four engines into one number hides the common case where a company is named consistently by ChatGPT and absent from Google AI Overviews, which is a different problem with a different fix, and our page on why ChatGPT cites you but Google does not goes through why the two diverge.

How soon should you expect the first real movement?

Expect the first real movement two to six weeks after the page changes are live, and expect the first reading you can defend a quarter in. The gap between those two dates is where most programmes get cancelled, so it is worth writing both into the plan before anybody starts editing pages.

What the early re-measurement can tell you is narrower than people hope. A run two weeks after a fix answers whether the page is being retrieved and considered again, which is a real signal and not the same as winning. Whether the page is chosen over the alternatives takes several more runs, because the effect has to be larger than the week-to-week swing described above before it is visible at all.

Sequence changes the timeline as well. Winning a question nobody currently answers well is faster than displacing an incumbent who is cited on it every time, because the second case needs the engine to prefer you over a source it already trusts. Our page on how long AEO takes to work separates those two cases and what each realistically takes.

One honest caveat belongs in every plan. Citation frequency can move for reasons that have nothing to do with your work, including interface and ranking changes at the engines themselves, so a month that looks flat is not proof the work failed and a month that looks excellent is not proof it succeeded. Direction across runs, measured against competitors, is the only reading that survives that.

What are the signs the work is not working?

The signs that AEO work is not working are specific, and none of them is a flat month. The first is direction: three consecutive runs with citation share flat or falling while competitors in the same answers hold or gain. That combination rules out the usual excuse, because an engine-side change moves everybody in the retrieved set and a real failure moves only you.

The second sign is a page-level one. When the pages that were rewritten are still absent from the retrieved set two months later, the problem is upstream of the writing: crawler access, rendering, or a page that answers a question nobody asks. A Vercel and MERJ analysis of more than 500 million crawler fetches found that none of the major dedicated AI crawlers execute JavaScript, so a rewrite shipped into a client-rendered template can be invisible no matter how good the copy is, which our page on whether AI crawlers can read JavaScript covers in detail.

The third sign is a vocabulary mismatch nobody has addressed. If the prompts you are losing use words that do not appear on the pages you are losing them with, the work has been aimed at the wrong target, and more of the same will not fix it. Discovered Labs' 2026 citation analysis found vocabulary alignment between page and query was the only page-level signal with a causal effect that survived controlling for domain authority, at an effect size of β=+0.37.

The fourth sign is organisational rather than technical. Programmes stall when the diagnosis is produced by one party and the pages are edited by another with no shared queue, so recommendations accumulate and nothing ships. Count the URLs changed per month. When that count is near zero, the measurement is not the problem.

What should make you suspicious of an AEO report?

A report should make you suspicious when it cannot be checked. The strongest tell is a summary with no underlying log: no prompt list, no engine, no dates, no run numbers, no cited URLs. Without those, every figure in the document is unverifiable, and the next supplier cannot pick up where this one left off because the baseline leaves with them.

A second tell is a prompt set that quietly grows. Adding prompts you are likely to win raises average citation share without anything improving, and the change is invisible unless the report states the set size every month. Ask for the count and the diff, and treat any change to the set as an event that needs explaining rather than a housekeeping detail.

A third tell is a single blended score with no engine breakdown and no competitors named. Both omissions remove the comparison that would make the number falsifiable. A fourth is a guarantee: a promised citation count or a promised position inside an answer, which the mechanism does not support, since the engines sample and their interfaces change without notice.

The last tell is a refusal to report a bad month. Given the volatility figures above, a programme that has never had a down month is either measuring something insensitive or presenting selectively. If you are reading a report you cannot audit and you want a second opinion on the same prompts, a free visibility assessment runs your buyer questions across the four engines and hands back the log, which is the artefact a report should have been built on.

What else do people ask about judging AEO results?

How long before AEO work shows measurable results?

Two to six weeks for citation frequency to respond to a structural change, and roughly a quarter before the direction is reliable. The first re-measurement after a fix answers a narrower question than most people expect: whether the page is being retrieved again at all. Whether it is winning the answer takes several more runs, because a single run cannot be told apart from variance.

Is referral traffic from ChatGPT a good way to judge AEO?

Referral traffic is a lagging and badly undercounted signal, so it is a poor primary metric and a reasonable secondary one. A large share of AI-driven visits arrive with no referrer at all and land in direct, and the engines do not tag links consistently. Judging a programme on AI referral sessions in its first quarter will usually get the programme cancelled while it is working.

What is a reasonable citation share to aim for?

No single figure applies, because the denominator is your own prompt set and a set full of easy prompts produces a flattering number. The useful target is relative: named more often than a specific competitor on the prompts where a buyer is choosing. Absolute share becomes meaningful only once the prompt set is fixed and large enough that a few swapped citations do not move the percentage.

Should you change the prompt set if results look flat?

Not while you are still measuring the effect of a change. Swapping prompts resets the series, and a series you have reset cannot show you a trend. Add prompts to a separate, clearly labelled set if a new question becomes important, keep the original running, and only retire the original when a page has genuinely stopped mattering commercially.

How do you tell a real improvement from an engine-side change?

Check whether competitors moved at the same time. When your citation share and three rivals' all shift in the same week, the likely cause is a change at the engine rather than anything on your site, and OpenAI's May 2026 interface change that made links more prominent is the kind of event that does it. When you move and the rest of the retrieved set does not, the change is more plausibly yours.

Related guides

The Assessment

Find out who gets named when buyers ask about your category.

A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.

Request a free visibility assessment →
SIGNALS · A BlackSig Systems company