An AEO engagement is three recurring jobs wrapped around one piece of setup. The setup is a baseline: a fixed set of buyer questions run across the engines, with every company and source the answers named, recorded with dates. The three recurring jobs are re-running that question set, fixing pages that are retrievable but not quoted, and publishing answers to questions nothing on your site covers. Month one is mostly measurement and technical checks; from month two the work is a queue. The reason the measurement half never finishes is that citations rotate: Trakkr's "The Half-Life of AI Citations" found 73.5% of citation URLs appeared exactly once and did not return, with the median brand falling to half its peak citation count in 31 days. Anything billed monthly that does not include a re-run of the same questions is a content retainer with an AEO label on it.
An AEO engagement month to month includes a measurement run, a set of page-level fixes and a small number of new pages, reported against a question set that does not change. Those are the three lines that should appear on every invoice, and the proportions shift over the engagement rather than the items.
The measurement run is the spine. A fixed list of buyer questions is put to ChatGPT, Claude, Perplexity and Google's AI surfaces on a schedule, and every answer is logged with the companies named and the URLs cited. Without that log, nothing else in the engagement can be assessed, because there is no series to compare a change against.
Page-level fixes are the second line and they are where most of the early movement comes from. The work is unglamorous: rewriting a page so its vocabulary matches how buyers ask, restoring a broken heading order, making a client-side rendered page readable to a crawler, adding the source to a statistic that was asserted. Each fix is small and each one is checkable.
Publishing is the third line, and it is a backlog rather than a cadence target. The questions your category gets asked that nothing on your site answers are a finite list, and working down that list is what wins questions you currently lose. A provider who quotes a fixed number of articles a month before seeing the gap is selling volume rather than coverage.
What should also appear monthly is a short written reading of the numbers, which is different from a dashboard. A reading says which questions moved, what changed on the site before they moved, and what will be done next. A dashboard says the number is 14.
The first month of an AEO engagement is spent establishing what is true now, because almost every later decision is a comparison against it. Three things get built: the question set, the baseline runs, and a list of page-level faults ordered by what they cost.
The question set is the part worth arguing about while it is cheap to change. It should be written in the words buyers use rather than the words the company uses, cover each distinct use case or buyer type separately, and include the shortlist questions that name competitors. Once it is fixed it should not be edited mid-engagement, because changing the questions resets the series.
Baselines get run at least twice, a week apart, before anything is fixed. Running them twice is not duplication: it measures your own noise, which is substantial. AirOps' 2026 retrieval study examined 548,534 pages retrieved across 15,000 ChatGPT prompts and found only 15% of retrieved pages were cited, and which pages win rotates between runs, so a single run cannot tell a real change from sampling.
The technical pass runs in parallel and usually finds the cheapest wins on the whole engagement. Crawler access, server-side rendering, heading structure and gated content are yes or no questions with known fixes, and our guide to what an AI visibility audit includes sets out the full checklist.
What should not happen in month one is publishing. Writing pages before the baseline exists means the first thing anybody asks, whether the new pages changed anything, has no answer.
After the first month the work settles into four workstreams, and a useful proposal names who does each one. The table below is how we scope an engagement, with the cadence that matches what the measurement can actually resolve.
| Workstream | What gets done | How often |
|---|---|---|
| Measurement | The same question set re-run across the engines, logged with every source cited | Monthly, weekly for a small set of buying questions |
| Page fixes | Rewrites for vocabulary, structure, sourcing and crawlability on existing pages | Continuous, worked as a queue |
| New coverage | Pages answering questions in the set that nothing on the site covers | Continuous, sized by the gap not by a quota |
| Reading and decisions | A written reading of what moved, what changed before it moved, and what is next | Monthly |
| Decay checks | Finding pages that were cited and stopped, and why | Monthly once anything is winning |
Monthly is the right default for the measurement run, and the reason is arithmetic rather than habit. Structural changes take weeks to show up in citation frequency, so a run a fortnight after a fix tells you whether a page is being retrieved again and not whether it is winning. Profound's 2026 citation decay study, which measured 883,000 pages across seven engines over twelve months and produced 1.19 million page-and-engine lifecycles, found the median page lost half its citation share within 11 days of peaking, which is why quarterly checks meet a sliding page too late.
The ratio between fixing and publishing should change over time. Early on, most of the gain is in pages that already exist and nearly work. Once those are done, the remaining gap is coverage, and the engagement starts to look more like a publishing operation with a measurement loop attached. A provider whose mix never changes is probably not reading their own numbers.
Most of an AEO engagement is content work, and the technical share front-loads into the first few weeks and then nearly disappears. The technical faults that matter are a short list and they are either fixed or not: whether crawlers can reach the page, whether the content survives without JavaScript, whether the heading order is intact, whether anything important sits behind a form.
Content is the larger share because it is where the dominant signal lives. Discovered Labs' 2026 citation analysis, built on more than two million citations using domain fixed effects and double machine learning, found content alignment was the only page-level signal with a causal effect that survived controlling for domain authority, at β=+0.37. Rewriting a page so it uses the nouns buyers type is a content task, not an engineering one.
Structure and sourcing sit in between and both are content tasks in practice. The ConvertMate GEO Benchmark 2026, an observational study of 12,500 queries across 8,000 domains, found 68.7% of cited pages used a strict H1 to H2 to H3 hierarchy, and Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024 measured a 41% lift in AI visibility from adding statistics to a page and 28% from adding quotations, across a benchmark of 10,000 queries.
The implication for staffing is that an engagement weighted heavily towards developer time is usually either early or mis-scoped. A proposal that reads as a three-month engineering project with a schema deliverable at the end has the ratio backwards, unless the site is a single-page application, in which case the rendering work genuinely comes first.
Schema markup deserves a specific note because it is sold hard. Structured data helps engines resolve what a page is about and it is worth having, but it does not substitute for the page saying the thing; our guide on whether schema markup helps AI citations covers what the evidence supports.
An AEO engagement needs four things from your team, and an engagement that stalls has usually lost one of them rather than hit a technical wall. The first is the ability to publish: somebody who can put a page live, or approve a page going live, within days rather than weeks.
The second is subject-matter time, and it is the one most often underestimated. The pages that win are the ones containing specifics only your company knows: what the integration actually syncs, what the limits are, what happens in the edge case a buyer is worried about. An hour a month with somebody technical is usually enough, and nothing substitutes for it.
The third is access. Server logs, analytics, Search Console and the staging environment are what let a supplier verify that a change landed rather than assert it. Engagements run without access produce reports that cannot be checked, which is a slow way to discover a problem.
The fourth is a named decision-maker for the question set. The list of buyer questions is a commercial document, not a technical one, and somebody in marketing or sales has to own whether it reflects how buyers actually ask. Our guide on who should own AI search visibility goes through where that usually sits.
What your team should not have to supply is the measurement discipline. Running the same prompts the same way every month and keeping a comparable log is precisely the part that in-house attempts tend to drop, and it is a reasonable thing to buy.
Several things buyers expect to be included usually are not, and finding out in month three is worse than asking in week one. Getting you onto third-party lists and review sites is the big one: appearing in somebody else's round-up is outreach work with a different skill set and a different success rate, and it is frequently scoped separately or not at all.
Review platform management is similarly excluded. Soliciting and responding to G2 or Capterra reviews touches customers and legal review, and most AEO providers will advise on it rather than run it. That matters because review sites appear in a lot of shortlist answers, as our guide on whether G2 and Capterra reviews affect AI recommendations works through.
Development work is the third common gap. A provider will usually specify the rendering fix and hand it to your engineers rather than implement it, which is fine if your engineers have capacity and a problem if the whole programme is waiting on a ticket nobody has prioritised.
Guarantees should be refused rather than merely left out of scope. Nobody can promise a citation count or a position inside an answer, because the engines change retrieval behaviour and interfaces without notice and the same prompt returns different sources on consecutive runs. A proposal containing a guaranteed number of citations is describing something the mechanism does not support, and our guide to choosing an AEO agency lists the other claims worth pushing back on.
Correcting outdated facts the engines state about your company is a related but separate piece of work, and in our own testing in September 2026 verified outdated answers turned out to be rarer than the marketing around them suggests. Scope it only if you have found one.
The difference between a real engagement and a report subscription is whether anything on your site changes as a result of the numbers. A subscription produces a dashboard and leaves the work to you, which is a legitimate product and should be priced like software rather than like a service.
Four questions separate them quickly. Does the provider rewrite pages or only recommend rewrites? Does the monthly report name specific pages and specific questions, or show aggregate scores? Does the question set belong to you in writing? And is the raw run log, every prompt with every source cited, something you can export?
The last one is worth insisting on even when the answer is awkward. Without the raw log you cannot verify a single line of the summary, and you cannot carry the baseline to a different supplier if the engagement ends. Some providers treat the log as proprietary, which they are entitled to do and which is better discovered before signing.
A tell in the other direction is a provider who reports tasks completed. Pages optimised and articles shipped are inputs, and an engagement that reports only inputs is asking you to accept effort as evidence. Our guide to telling whether your AEO is working sets out what the readout should contain instead.
If you are weighing a specific proposal and want a baseline to read it against, a free visibility assessment runs your buyer questions across the four main engines and reports who is named today, which makes the first month of any proposal considerably easier to judge.
There is no defensible standard number, and a proposal that names one before seeing your site is quoting capacity rather than scope. What determines it is how many of your existing pages are close to working, how many questions in your category have no page at all, and how fast your team can approve changes. A useful proposal gives a range for the first quarter and revises it once the baseline and the technical pass are done.
Weekly runs earn their cost on a small set of buying questions and waste it on everything else. The reason is that structural changes take weeks to register, so most weekly movement is sampling noise rather than progress, and Profound's 2026 study of 883,000 pages found the median page lost half its citation share within 11 days of peaking, so the fast-moving signal is real but hard to act on. The common compromise is monthly for the full set and weekly for the handful of questions with a buyer behind them.
Yes, and for a company with a disciplined content team it is often the better split. The part that is genuinely hard to sustain in-house is running the same prompts the same way every month and keeping a log you can compare, which is where most internal attempts quietly stop. If you take this route, write the question set down, name the person who runs it, and put the run in their calendar rather than in a plan.
The question set, the engines, the run frequency, who owns the raw run log, what the monthly readout contains, and a review point with cancellation genuinely on the table. Fixing the question set in the contract is what makes every later report comparable and stops easier questions being quietly substituted when the hard ones are not moving. Our guide to choosing an AEO agency covers the rest of the contract terms worth naming.
Cost is driven by the size of the question set, how many engines and competitors are tracked, how many of your pages get read in detail, and how much of the writing and page editing the supplier does rather than your team. A single product line in one language is a different engagement from five lines across three regions. Current SIGNALS figures live on the pricing page rather than inside articles, because a number copied into an article goes stale the day it changes.
A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews and reports who is named and from which page, so the first month of any engagement has something to be measured against.
Request a free visibility assessment →