An AI visibility audit answers two questions, and a good one answers both. The first is who gets named when buyers ask your category questions, measured by running a fixed prompt set across ChatGPT, Claude, Perplexity and Google AI Overviews, repeatedly rather than once. The second is what on your own pages is stopping you from being one of the names, which is a page-level diagnosis of crawler access, parsing, vocabulary and quotable structure. An audit that only runs prompts tells you that you are losing without telling you why, and an audit that only reads your pages tells you nothing about who you are losing to. Repetition is not optional: researchers at the University of St. Gallen found that identical prompts run several times in one day returned overlapping source lists only 32% to 43% of the time.
An AI visibility audit includes a measurement half and a diagnosis half, and a scope document that only describes one of them is describing half a job. The measurement half builds a set of prompts a real buyer would type, runs them across ChatGPT, Claude, Perplexity and Google AI Overviews, and records for every run which companies were named and which URLs were cited. The diagnosis half takes the prompts you lost and reads your own pages against them, looking for the specific reason the engine passed you over.
Both halves produce a list, and the two lists answer different questions. The measurement list says who owns your category inside the answer, which is usually a mix of competitors, review sites, trade publications and Reddit threads. The diagnosis list says what to change, in the order worth changing it. A buyer who reads only the first list knows they have a problem and has no idea what to do on Monday.
Competitors belong in the scope from the start rather than as an add-on. Being absent is not the finding; being absent while three named rivals are present in the same answer is the finding, and it is the version that survives a conversation with a CEO. The same run produces both at no extra cost, so an audit that reports only your own citation count has thrown away the more useful half of what it already collected.
One thing an audit is not is a crawl report with an AI label on it. Plenty of what is sold as an AI visibility audit is a standard technical SEO crawl plus a paragraph about schema markup. The test is simple: ask whether the deliverable contains the actual text of AI answers to your buyer questions, with the sources listed. If it does not, nobody asked the engines anything.
An AI visibility audit differs from an SEO audit in what it treats as the outcome. An SEO audit works backwards from ranking position in a list of blue links. An AI visibility audit works backwards from inclusion in a generated answer, which is a different selection process with different inputs, and the two can disagree completely about the same page.
How completely they disagree is measurable. The ConvertMate GEO Benchmark 2026, an observational study across 8,000 domains, found that 83% of AI citations come from pages outside Google's top 10 results, which means the page that wins the answer is usually not the page that wins the ranking. An audit built on rank tracking is looking at the wrong column.
The inputs diverge too. Discovered Labs' 2026 citation analysis, the largest published to date, tested page-level signals against citation frequency with domain fixed effects and found that vocabulary alignment between the page and the way buyers actually search was the only signal with a causal effect that survived controlling for domain authority, at an effect size of β=+0.37. Most of the classic SEO signals collapsed to near zero once domain was held constant. Our methodology page sets out how that finding shapes the weighting.
Practically, this changes what gets read. An SEO audit reads titles, internal links and Core Web Vitals. An AI visibility audit reads whether a section can be lifted out of the page and still make sense to somebody who never saw the section above it, because that lifted passage is the unit an engine quotes. Both are worth doing. Only one of them predicts whether you are named in the answer, and our comparison of AEO and SEO goes through where the two practices overlap.
The prompt set should be large enough to cover the decisions your buyers actually make, and every prompt in it has to run more than once. Running a prompt once and writing down the answer is the most common defect in audits sold today, and it produces a document that looks precise and describes nothing repeatable.
Repetition is not a nicety. In "Don't Measure Once: Measuring Visibility in AI Search (GEO)", Schulte, Bleeker and Kaufmann of the University of St. Gallen ran identical prompts multiple times and found the cited sources overlapped in only 32% to 43% of cases within a single day, with roughly 65% of cited sources turning over from one day to the next. Their convergence analysis recommends at least seven runs per prompt per day, and reading per-brand trends only over rolling windows of two to four weeks.
Set size follows from coverage rather than from a round number. A single product sold into one market can be represented by a few dozen prompts spread across the questions buyers ask at different stages. A company with several product lines, buyer types or regions needs a set several times that size, because each of those is a separate retrieval problem. Our guide to how many prompts you need to track AI visibility works through the arithmetic.
Ask any prospective auditor for the prompt list before they run it. Reading it tells you more about the quality of the engagement than the methodology section will: prompts written by somebody who understands your category look like things buyers say out loud, and prompts written by somebody who does not look like keywords with question marks added.
On your own pages an audit checks four things in sequence, because a failure at an early stage makes everything after it irrelevant. The sequence comes from the AgentGEO pipeline analysis published as arXiv:2603.09296 in March 2026: retrieval, whether crawlers can reach the page at all; parsing, whether the content can be extracted from the HTML; ranking, whether the page beats the others in the retrieved set; and generation, whether anything on it is quotable as a standalone unit.
Retrieval and parsing are where the cheap, embarrassing failures live. A Vercel and MERJ analysis of more than 500 million crawler fetches found that none of the major dedicated AI crawlers execute JavaScript, including OpenAI's GPTBot, Anthropic's ClaudeBot and PerplexityBot, so content that only exists after client-side hydration is invisible to every one of them. A site can pass a Google audit cleanly and still serve an empty shell to the engines that matter here, which is the subject of our page on whether AI crawlers read JavaScript.
Ranking is where most pages actually lose, and it is a vocabulary problem more often than a quality problem. A page written in the language a company uses internally competes for a query written in the language a buyer uses, and loses to a worse page that happens to say it the buyer's way. Reading the failed prompts next to the page copy is the only way to catch this, and no crawler can do it for you.
Generation is the last check and the easiest to fix. A section that opens with "This means that" is unusable to an engine assembling an answer from fragments, because the fragment no longer has a subject. Restating the subject at the top of every section costs nothing and changes whether a passage can be quoted.
At the end of an audit you should get four artefacts, and the fourth is the one people forget to ask for. The raw runs matter as much as the summary: without them you cannot check the work, and you cannot compare the next round against this one. The table below is the deliverable set we hand over, and what each part is for.
| Deliverable | What it contains | What you do with it |
|---|---|---|
| The prompt log | Every prompt, engine, date, run number, answer text and cited URL | Check the work, and use it as the baseline for the next round |
| The competitive picture | Who is named, how often, and which page of theirs is cited | Show leadership the gap in terms they already argue about |
| The page diagnosis | Per page, which pipeline stage it fails at and the evidence | Decide what gets fixed and in what order |
| The fix list, prioritised | Specific changes to specific URLs, with effort and expected effect | Hand to whoever edits the pages |
Priority order matters more than completeness. A list of 60 findings sorted by severity is a way of not making a decision, because everything above "minor" looks urgent. A list that says which six changes to make first, and why those six, is a plan. The ordering should follow the pipeline: no point rewriting copy on a page the crawlers cannot reach.
Ask for the raw prompt log in writing before the engagement starts. Some vendors treat it as proprietary, which is a reasonable position to hold and a bad one to discover afterwards, because without the log you cannot verify a single claim in the summary and you cannot take the baseline with you if you change supplier.
You can run a useful version of an AI visibility audit yourself in an afternoon, and it will be worth doing even if you later pay somebody. Write down the ten questions a buyer would ask before shortlisting a company like yours. Ask each one in ChatGPT, Claude, Perplexity and Google AI Overviews. Record who gets named and which URLs are cited. That single exercise usually settles the internal argument about whether this matters, because most people have never seen their category answered without their company in it.
Where the do-it-yourself version breaks down is repetition and diagnosis. Ten prompts run once across four engines is 40 observations, and the St. Gallen variance figures say a large share of what you write down will be different tomorrow. Doing it properly means running the same set several times, on a schedule, and keeping a log in a form you can compare, which is where enthusiasm usually runs out.
The second gap is harder. Knowing you lost a prompt does not tell you why you lost it, and the honest reasons are often unflattering: the page answers a question nobody asked, or the useful content sits behind a form, or the one section a buyer needed is four scrolls down under a heading nobody would search. Reading your own pages for this is difficult because you already know what they say. Our 27-point AI citation checklist is the closest thing to a self-serve version of the page-level half.
If the afternoon version shows competitors named and you absent on questions that matter commercially, the next step is a proper baseline rather than a rewrite. A free visibility assessment runs your buyer questions across the four engines and reports who is cited and from which page, which is the measurement half done properly and at no cost.
Two to four weeks for the version worth acting on. The prompt runs themselves take days rather than weeks, but a single round of runs cannot separate a real absence from engine variance, so the schedule has to allow a second round a week later. Page-level diagnosis runs in parallel. An audit delivered in 48 hours is a single round of prompts, which is a screenshot rather than a measurement.
The size of the prompt set, the number of engines, the number of competitors tracked alongside you, and how many of your own pages get read in detail. A single product line in one language with 30 prompts is a different job from five product lines across three regions. Current SIGNALS figures are published on the pricing page rather than quoted inside articles, because a figure copied into an article goes stale the day it changes.
A tool subscription replaces the measurement half and not the diagnosis half. Trackers are good at telling you your citation share moved, and they do it continuously, which an audit does not. What they do not do is read your pages against the prompts that failed and tell you which stage of retrieval you are losing at, because that is a judgement rather than a metric.
All four, even when ChatGPT dominates your referral numbers. Previsible's 2026 AI traffic study found ChatGPT carrying 92.4% of trackable standalone referral traffic, but referral share and citation share are different questions, and the engines disagree about who to name. A company cited by ChatGPT and invisible in Google AI Overviews has a specific, fixable problem that a single-engine audit never surfaces.
A guaranteed number of citations, a guaranteed position inside an answer, or a fixed timeline to either. Citation frequency shifts two to six weeks after structural changes and moves for reasons outside your site, including interface changes at the engines themselves. An audit can tell you what is blocking you and what the realistic ceiling looks like. Any provider putting a guaranteed citation count in a proposal is selling something the mechanism does not support.
A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.
Request a free visibility assessment →