AEO is sold as a project and as a monthly retainer, and the split is not arbitrary: structural work stays done, and everything downstream of it does not. Crawler access, rendering, heading structure, vocabulary alignment and a first baseline are project work, finished once and holding. Measuring who gets named, publishing answers you do not yet cover, and refreshing pages that have gone quiet are recurring, because the answer is regenerated on every ask. A 2026 Profound analysis of 240 million citations put the median half-life of a cited source at roughly 4.5 weeks, and according to AirOps' 2026 State of AI Search study only about 30% of brands stayed visible from one answer to the next. The sensible order is to measure first, do the structural work as a scoped project, and decide on a retainer with the baseline in hand rather than before it.
AEO is sold both ways, and the useful answer is that it starts as a project and only becomes a retainer if the recurring half is worth paying for at your company. The work splits into two kinds. One kind stays done: making pages reachable and parseable by the crawlers, fixing a broken heading order, rewriting a page so its vocabulary matches what buyers actually type, publishing the answer to a question nobody at your company had answered. The other kind does not stay done: which companies get named when a buyer asks, which of your pages the engines pull from, and what a competitor published last month that is now being quoted instead of you.
Agencies that only sell retainers tend to describe the first kind as trivial. Agencies that only sell projects tend to describe the second kind as somebody else's problem. Both descriptions are shaped by what the seller has to sell, which is why the shape of the engagement is worth deciding before you start reading proposals.
The question that settles it for a specific company is simple: after the project ends, who keeps publishing and who keeps measuring? If the answer is that a content team will keep shipping pages and somebody will re-run the question set every month, a project plus internal discipline genuinely works. If the answer is nobody, a project buys you a document, and the document ages at the same speed as the answers it describes.
Sequence matters more than the label. Buying a twelve month retainer before anybody has measured who currently gets named in your category means agreeing to a programme without knowing the size of the gap. Running the measurement first costs almost nothing and changes what you are negotiating about.
AI visibility work does not stay done because an answer is generated fresh each time somebody asks, not stored and served like a ranking. Two pages can be equally well built and get different treatment on consecutive days, which means a single good result is not a position you now hold.
Researchers at the University of St. Gallen measured this directly. In "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv:2604.07585, April 2026), Julius Schulte, Malte Bleeker and Philipp Kaufmann tracked ChatGPT, Perplexity, Gemini and Google AI Mode over 45 days and found cited sources turning over roughly 65% from one day to the next. Identical prompts re-run at the same moment overlapped only 32% to 43% of the time.
Decay shows up over longer windows too. A 2026 Profound analysis of 240 million citations put the median half-life of a cited source at roughly 4.5 weeks, with 40% to 60% of the domains cited for a given query changing month to month and 70% to 90% rotating over six months. Our guide to measuring AI citation decay covers how to separate that drift from normal variance.
Staying present looks harder still when you measure brands instead of URLs. According to AirOps' 2026 State of AI Search report, which tracked citation patterns across ChatGPT, Perplexity, Google AI Overviews and Gemini, only about 30% of brands stayed visible from one answer to the next and only about 20% were present across five consecutive runs. The same report found pages not updated quarterly were about 3 times more likely to lose citations.
None of that makes a retainer automatically correct. What it rules out is the version of the sales pitch where a one-off rebuild makes you permanently citable, and the version where month-to-month movement proves the work is failing. Both readings misunderstand a system that regenerates its answer every time.
A one-time AEO project delivers the structural half, and that half is real. Crawler access is either fixed or not: a page that renders its content only after client-side JavaScript either gets converted to server-rendered HTML or stays invisible to the engines, and once it is fixed it stays fixed through the next deploy. Heading order, self-contained sections and a comparison table where options exist are the same kind of change.
Structure is worth doing properly because the cited population looks structurally different from the web at large. The ConvertMate GEO Benchmark 2026, an observational study of 12,500 queries across 8,000 domains, found 68.7% of cited pages using a strict H1 to H2 to H3 hierarchy, and 83% of AI citations coming from pages outside Google's organic top 10.
The second thing a project delivers is a baseline, and a baseline has a longer useful life than the fixes do. A recorded set of buyer questions, run across the engines, with the companies and URLs that came back, is the only document that lets you tell in six months whether anything changed. Without it every later claim about improvement is an assertion.
Vocabulary work also belongs in a project, not in a monthly cycle. Rewriting a page so it uses the nouns buyers use is a one-time edit per page, revisited only when the market renames the thing you sell. That edit carries more weight than its effort suggests: Discovered Labs' 2026 citation analysis found vocabulary alignment was the only page-level signal with a causal effect that survived controlling for domain authority, at an effect size of β=+0.37. Our methodology page explains why that finding carries the heaviest weight in the score.
A retainer buys three things a project cannot: repeated measurement, a publishing cadence, and somebody whose job it is to react to what the measurement says. Each one is a recurring cost because the thing it responds to recurs.
Repeated measurement is the least glamorous and the hardest to fake. A fixed question set re-run on a schedule is what turns "we think we are doing better" into a series you can read. One round of prompts is a screenshot; the St. Gallen variance figures above are the reason a screenshot cannot be compared to anything.
Publishing cadence is where most of the actual movement comes from, and it is the part companies most often fail to resource. Winning a question you currently lose usually means a page that answers that question exists, is reachable, and says the thing a buyer asked in the words they used. If your category has forty such questions and you have covered nine, no amount of re-measuring the nine changes the answer.
Reacting to decay is the third recurring job. Pages that earned citations stop earning them, and the AirOps 2026 figures quoted above put pages left unrefreshed for a quarter at roughly 3 times the risk of losing citations. Noticing which specific page went quiet, and why, is work that only exists if somebody is looking every month.
What a retainer should not buy is activity for its own sake. A monthly report listing tasks completed, with no movement in who gets named, is the failure mode to watch for, and our guide on telling whether your AEO is actually working sets out what the readout should contain instead.
The parts of AEO that are one-off fixes are the structural ones, and the parts that recur are the ones answering to somebody else's behaviour. Splitting the work item by item is the fastest way to size a project against a retainer, because the split is not a matter of opinion: some of it is a change to a file that stays changed, and some of it is a response to something outside your control that keeps happening. The table below is how we scope it.
| Work item | One-off or recurring | Why it falls there |
|---|---|---|
| Crawler access and server-side rendering | One-off, per site | A rendering change holds until somebody rebuilds the front end |
| Heading order and self-contained sections | One-off, per page | Structure does not degrade on its own |
| Vocabulary alignment on existing pages | One-off, revisited when the market renames things | The words buyers use change slowly, not monthly |
| Baseline measurement of who gets named | One-off to establish, then recurring to be useful | A baseline with nothing to compare against is a screenshot |
| Publishing answers to questions you do not cover | Recurring | Coverage is a backlog, and competitors keep adding to theirs |
| Refreshing pages that have stopped being cited | Recurring | Cited sources rotate, so yesterday's winner goes quiet |
| Third-party presence and review profiles | Recurring | Other people's pages are maintained on their schedule |
Read down the first column and the honest scope of a project becomes visible: the top four rows, done once, properly. Read the bottom three and you have the case for a retainer, or for hiring somebody internally to do exactly those three things.
A pilot is the right way to test AEO, and its minimum length is set by measurement variance rather than by anyone's sales preference. The St. Gallen work cited above recommends reading per-brand trends over rolling windows of two to four weeks, with repeated runs of each prompt instead of one. That arithmetic means a 30 day pilot gives you roughly one readable window, and a 90 day pilot gives you three: enough to tell a trend from a wobble.
Four terms make a pilot readable. Fix the question set in writing before anything starts, and do not change it mid-pilot, because changing the questions resets the series. Fix the engines and the run cadence the same way. Agree what counts as movement, in terms of how often you are named and from which page, rather than in terms of tasks completed. And agree the exit: what happens, concretely, if the number has not moved.
Scope narrow, not broad. One product line, one buyer type, one language and a few dozen prompts will tell you more than a sprawling set across three regions, because a narrow set can be run often enough to be read. Our guide to how many prompts you need works through the sizing.
Insist on keeping the raw prompt log. Without it you cannot verify a single line of the summary, and you cannot carry the baseline to a different supplier if the pilot goes badly. Some vendors treat the log as proprietary, which is a position they are entitled to hold and a bad thing to find out in month three.
If you have never measured at all, the cheapest version of a pilot is the measurement half on its own. A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews and reports who is named and from which page, which is the baseline you would otherwise be paying a project to establish.
A retainer turns into an unreviewed subscription when the monthly report measures the supplier's activity instead of your visibility. Preventing that is a contract problem and a reporting problem, and both are easier to fix before signing than after.
Put the question set in the contract. A named, fixed list of buyer questions, with the engines and the run frequency, is what makes every later report comparable, and it stops the quiet substitution of easier questions when the hard ones are not moving. Agree in the same clause that the raw runs belong to you.
Define what the monthly readout shows. At minimum: how often you were named across the question set, which pages of yours were cited, who else was named, and what changed since the previous window. A readout built on that can be wrong but cannot be vague. A readout listing pages optimised and blog posts shipped tells you what was spent, not what happened.
Set a review point with a real decision attached, far enough out to be fair. Two to three months is too soon for a structural change to show, and twelve months is long enough for nobody to remember what was promised. A review at the end of a quarter, against the baseline, with cancellation genuinely on the table, keeps both sides honest.
Expect the honest answer to include things that cannot be guaranteed. Nobody can promise a number of citations or a position inside an answer, because the engines change their interfaces and their retrieval behaviour without notice. A provider willing to put a guaranteed citation count in a proposal is selling something the mechanism does not support, and our guide on choosing an AEO agency lists the other claims worth pushing back on.
Three months is the shortest length that can produce a readable result, because of how much the answers move on their own. The University of St. Gallen study on measuring AI search visibility recommends reading per-brand trends over rolling windows of two to four weeks, which means a single month gives you roughly one window and no trend. Anything shorter than a quarter is a pilot, and it is more honest to call it one.
The structural work holds and the coverage stops growing, so you decay slowly rather than falling off a cliff. According to a 2026 Profound analysis of 240 million citations, the median half-life of a cited source is roughly 4.5 weeks and 70% to 90% of the domains cited for a given query rotate over six months, so pages that were earning citations will give some of them back. Companies that stop usually lose ground on new questions first, because those are won by publishing rather than by holding position.
Yes, and for a company with a disciplined content team it is often the better split. The hard part is not running the prompts, it is running the same prompts the same way every month and keeping a log you can compare, which is where most in-house attempts quietly stop. If you go this route, write the question set down before the project ends and put the monthly run in somebody's calendar with their name on it.
Usually not at first. A small surface area means the structural work is genuinely finishable, and a narrow category means the question set is small enough to run yourself. The case for a retainer gets stronger when the category has more questions than you can cover, when competitors are publishing faster than you, or when several product lines or regions each behave like a separate retrieval problem.
Cost is driven by the size of the question set, how many engines and competitors are tracked, how many of your pages get read in detail, and how much of the writing and page editing the supplier does instead of your team. A single product line in one language is a different engagement from five lines across three regions. Current SIGNALS figures live on the pricing page rather than inside articles, because a number copied into an article goes stale the day it changes.
A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews, records who is named and from which page, and gives you the baseline every later decision gets compared against.
Request a free visibility assessment →