An automated AI visibility report is four moving parts: a fixed prompt set of the questions buyers really ask, scheduled runs of every prompt on every engine, a parser that records the brands named, the pages cited and the competitors that appeared, and a dated store you can query. Everything else is presentation. The two mistakes that ruin these reports are sampling too little, since a single run is a draw rather than a measurement, and changing the prompt wording mid-quarter, which resets the trend line without anybody noticing. Google Search Console now shows AI impressions on Google's own surfaces, with no clicks, no position and no API, so it complements this report rather than replacing it.
Automating AI visibility reporting means turning an occasional manual check into a scheduled job with four parts. First, a prompt set: the questions your buyers ask, written the way they would say them out loud, held fixed so that this month compares with last month. Second, a runner that sends every prompt to every engine you care about, several times each, in clean sessions. Third, a parser that reads each response for your brand, your competitors and the URLs cited. Fourth, a store that keeps every observation with a timestamp, the engine, the prompt and the raw text.
Building the prompt set is the part that decides whether the report is useful, and it is the part most teams rush. Twenty to fifty questions covering the buying journey beats several hundred generated from a keyword tool, because you will be reading these every week and defending movements in them. Include the category question, the comparison question, the problem question a buyer asks before they know your category exists, and the direct question about your own company.
Keep the runner boring and repeatable. Fresh sessions with memory and custom instructions disabled, one fixed geography unless location is part of your market, the same wording every run, and a full copy of the raw response saved rather than a parsed summary. The raw text is what lets you answer next month's question about why a number moved, and no parser you write today will extract everything you will want later.
Then decide what "appeared" means before the first run, since that definition moves the number more than any engine does. A brand named in passing, a brand recommended, and a page cited without the brand being named are three different outcomes. Write the rule down, apply it consistently, and revisit it only at a version boundary you record.
An AI visibility report should contain five measures, reported per engine rather than blended into one score. Mention rate answers how often you were named across runs. Citation rate answers how often a page of yours was linked, and which page. Share of voice puts your mention rate next to the competitors named on the same prompts. Coverage gaps list the prompts where rivals appear and you do not. Movement shows the change since the last period, expressed as a range rather than a point.
Blending engines into a single score is the most common way these reports mislead. Strong ChatGPT visibility with weak Google AI Overview visibility is a specific, fixable, index-level problem, and an average hides it behind a comfortable middling figure. Keep the engines in separate columns for the same reason a financial report keeps currencies separate.
Add the source list, which is the diagnostic half of the report. Recording which URLs each engine cited tells you whether a competitor is winning on their own pages, on a listicle neither of you controls, or on a forum thread. Those three situations call for different work, and none of them is visible in a mention count. Where the buying query returns comparison pages rather than vendor sites, your own category comparison page becomes the asset that competes.
| Metric | What it answers | What a bad number means |
|---|---|---|
| Mention rate, per engine | How often you are named across repeated runs | The engine does not consider you part of the category |
| Citation rate and cited page | How often your pages are linked, and which ones | You are known but your pages are not the evidence |
| Share of voice | Your mention rate against named competitors | Someone else owns the answer you are paying for |
| Coverage gaps | Prompts where rivals appear and you never do | A missing page, or a page in the wrong vocabulary |
| Movement, with a range | Change since the previous period | A single-run comparison reporting sampling noise |
Google Search Console can now report AI Overviews and AI Mode, but only impressions and only on Google's own surfaces. Google introduced generative AI performance reports on 3 June 2026 and, according to Search Engine Land's coverage of the global rollout, finished releasing them to all properties worldwide on 31 August 2026. The data breaks down by page, country, device and date.
What is missing shapes how you can use it. There are no clicks, no click-through rate, no queries and no average position in the report, so it tells you that Google's AI surfaces showed your page and nothing about what that appearance was worth. Search Engine Journal's report on the worldwide rollout makes the same point about the missing commercial metrics.
The second limitation is programmatic. The Search Analytics API does not expose an AI Overviews or AI Mode type, and the data is absent from the BigQuery bulk export, so an automated pipeline cannot pull it. A scheduled manual CSV export, filed with a date, is the honest workaround until that changes, and it is worth doing because it is the only first-party impression data anyone publishes.
Treat it as one panel in the report rather than the report. Search Console covers Google, your prompt runs cover the engines that Google does not, and your analytics covers the small number of people who click through. None of the three alone describes AI visibility, and a team that adopts the Search Console panel as its whole AI measurement has quietly narrowed its view to one vendor's surfaces. Our guide to what AI search actually sends in traffic terms covers the analytics half.
Run the full prompt set weekly, a small high-value subset daily, and present the trend monthly. That cadence matches the rate at which the underlying reality changes: pages get published and indexed over weeks, engines change retrieval behaviour over months, and anything faster is mostly sampling noise being plotted as a movement.
Repetition per run matters more than frequency. Schulte and colleagues at the University of St. Gallen, in the April 2026 arXiv preprint Don't Measure Once: Measuring Visibility in AI Search (GEO), reported the standard error of a brand's estimated detection rate falling below 0.10 at seven runs per prompt per day, with source coverage needing eight, and found the sources cited on consecutive days overlapping by only 34% to 42%. Seven to ten runs is the working minimum for any prompt you intend to make a decision on.
| Cadence | What to run | What it is for |
|---|---|---|
| Daily | Five to ten buying prompts, all engines, several runs each | Catching a competitor taking an answer you rely on |
| Weekly | The full prompt set, all engines, seven or more runs each | The working number the team acts on |
| Monthly | Aggregation, Search Console AI impressions, assistant referrals | The trend that goes in front of a stakeholder |
Budget the volume before choosing tools, because prompts multiplied by engines multiplied by runs multiplied by frequency is the number that decides cost and runtime. Thirty prompts on four engines at seven runs each is already 840 responses a week, which is fine for a machine and impossible by hand.
Stopping a report from measuring its own noise comes down to three disciplines, and none of them is technical. Sample enough per prompt, because a movement built from single runs is a coin toss. Freeze the prompt wording, because rewording a prompt changes the question and silently resets the baseline. Keep a change log of every alteration to the prompt set, the parser or the definitions, dated, next to the data.
The change log earns its keep the first time a number jumps. With it, you can say that mention rate rose the week three prompts were reworded and treat the jump as an artefact. Without it, somebody will attribute the jump to last month's content work, and the programme will spend a quarter repeating something that did nothing.
Corroborate against a second source where you can. Assistant referrals in your analytics, Search Console AI impressions, and inbound enquiries that mention an assistant all move slowly and independently. When the prompt-run numbers rise and none of the three move at all over a quarter, the measurement deserves the scrutiny before the market does.
Report the uncertainty rather than hiding it. A visibility figure presented as a range across a stated number of runs survives being re-run in the room by a sceptical executive, and a single percentage does not. The credibility of the whole programme rests on that one presentational choice, and it costs nothing. Our note on why AI answers change between runs explains the mechanism to anyone who asks why the range exists.
Fix a prompt set of the questions your buyers actually ask, run every prompt several times on each engine on a schedule, parse each response for your brand, your competitors and the sources cited, and write every observation to a dated store you can query later. The report is then a view over that store rather than a document somebody assembles by hand. Everything that varies between runs, wording, geography, account state, has to be held constant or the trend line measures the setup rather than the market.
Five things, per engine rather than blended: mention rate across runs, citation rate with the page that was cited, share of voice against the named competitors on the same prompts, the coverage gaps where competitors appear and you do not, and the movement since the previous period with a range rather than a single figure. Blended scores hide the engine-level failure that is actually costing you the answer.
Partly. Google introduced generative AI performance reports on 3 June 2026 and completed the worldwide rollout on 31 August 2026, covering impressions in AI Overviews, AI Mode and generative features in Discover, broken out by page, country, device and date. Clicks, click-through rate, queries and average position are not in that report, and it covers Google's surfaces only, so it cannot tell you anything about ChatGPT, Claude or Perplexity.
Weekly for the full prompt set, daily for a small group of high-value buying prompts, and monthly for the trend you actually present. Reporting more often than your content changes mostly reports sampling noise, and reporting less often than monthly means a competitor can take an answer and hold it for a quarter before anyone notices.
Sample enough per prompt, hold the prompt set fixed, and keep a change log. The University of St. Gallen preprint on measuring AI search visibility reported the standard error of a brand's detection rate falling below 0.10 at seven runs, so anything under roughly seven runs per engine is a draw rather than a measurement. A movement that coincides with a prompt rewrite is a measurement artefact, and the change log is what lets you say so.
A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.
Request a free visibility assessment →