A set in the low dozens is enough for most companies, and the number of prompts matters less than how many times each one runs. Researchers at the University of St. Gallen found that identical prompts run repeatedly on the same day produced cited-source lists overlapping only 32% to 43% of the time, with roughly 65% of sources turning over from one day to the next, and their convergence analysis recommends at least seven runs per prompt per day with trends read over two to four week windows. A set of 30 prompts run seven times across four engines is a better measurement than 200 prompts run once, and it costs less to maintain.
Most companies need a set in the low dozens to track AI visibility usefully, and the count matters far less than the number of runs behind each prompt. A focused product sold into one market is well represented by 30 to 50 prompts covering the questions buyers ask at different stages. A portfolio with several product lines, buyer types or regions needs a multiple of that, because each of those is a separate retrieval problem with its own competitors.
Set size is a coverage decision rather than a statistical one. The question a prompt set has to answer is whether it represents the decisions your buyers actually make, so the sensible way to size it is to list those decisions first and then write the questions that lead to each. A set assembled the other way round, by generating variations until the number looks impressive, is mostly synonyms competing with each other.
| Company shape | Workable set size | What drives it |
|---|---|---|
| One product, one market | 30 to 50 prompts | Stages of one buying decision, plus a brand group |
| Several products, one market | 50 to 150 prompts | One decision set per product line, competitors differ per line |
| Several products, several regions or languages | 150 or more | Retrieval sets diverge by language and locale |
| Single high-value question you must own | 5 to 10 prompts, run daily | Depth of runs rather than breadth of coverage |
Published practitioner sets sit in the same range. Kevin Indig's prompt-tracking structure uses 40 seed prompts split into 12 brand, 12 category and 16 problem-focused questions, with three personas applied to the 28 category and problem prompts, and around five consecutive runs once a week. The shape of that set is more instructive than its size: brand, category and problem are three different measurements, and mixing them produces a number that moves for reasons nobody can name.
Running a prompt once tells you nothing because the engines sample. Two identical questions asked an hour apart return different source lists, so a single observation records one draw from a distribution and presents it as a position. Treating that draw as a measurement is the most common defect in AI visibility reporting, and it produces documents that look precise and cannot be reproduced.
The variance is now quantified. In "Don't Measure Once: Measuring Visibility in AI Search (GEO)", Schulte, Bleeker and Kaufmann of the University of St. Gallen ran identical prompts multiple times and found cited sources overlapped in only 32% to 43% of cases within a single day, with roughly 65% of cited sources turning over from one day to the next. Brand sets were steadier and still unstable, showing 45% to 59% consistency across runs.
Individual URLs are the least stable unit of all. Trakkr Research logged 108,650 citation URLs across 10,991 brands and eight tracked models and found 73.5% of them appeared exactly once and never returned. Designing a report around which URL was cited, rather than around whether your company was named, means most of what you record is noise by construction.
The practical consequence is a change in what you write down. Instead of "we rank third for this prompt", the honest record is "we were named in five of seven runs of this prompt on this engine this week", which is a probability and behaves like one. Our page on why AI engines give a different answer every time explains the sampling behind it.
Run each prompt at least seven times per day when you need a reading you can defend, which is the figure the St. Gallen convergence analysis recommends, and read per-brand trends only over rolling windows of two to four weeks rather than single days. Below that, the measurement moves more than the thing you are measuring, and the report describes the engine's sampling rather than your standing.
Seven runs a day on every prompt is more than most programmes will pay for, so the trade is made on breadth. Cutting the set to 30 prompts and running them properly gives you a series you can act on; keeping 200 prompts and running each once a month gives you a spreadsheet. When the budget forces a choice, cut prompts and keep runs, because the prompts you drop cost you coverage while the runs you drop cost you the ability to say anything at all.
Frequency can differ by prompt group, and should. A handful of questions with a buyer visibly at the end of them justify daily runs with several repeats. Informational questions that support the category can run weekly. Brand prompts, which also serve as an accuracy check on what engines state about you, are fine monthly unless something has changed on your site.
One rule protects the whole series: never change the number of runs without recording it. Citation share is a rate, and a rate compared between a week of seven runs and a week of two is not a comparison. The run count belongs in the log next to every observation, which our guide to automating AI visibility reporting treats as a schema requirement rather than a nicety.
Choose prompts by working backwards from buying decisions, then writing each question the way a person would say it out loud. Four groups cover most categories: problem questions asked before a solution has a name, category questions comparing approaches, recommendation questions asking who to use, and brand questions naming you directly. Each group answers a different commercial question, and each deserves its own line in the report.
Wording matters more than coverage, because vocabulary is the dominant signal in whether a page is retrieved at all. Discovered Labs' 2026 citation analysis found alignment between a page's language and the way buyers actually search was the only page-level signal with a causal effect that survived controlling for domain authority, at an effect size of β=+0.37. A prompt set written in internal product language measures a market that does not exist.
Real query data beats intuition for the wording. According to our own Google Search Console export for signalscite.com, covering 93 queries in the 90 days to 27 September 2026, 38% of them run to six words or longer, and the Bing queries that reach the site are often whole sentences with clauses. Mining your own search console, sales call notes and support tickets for phrasing produces prompts that look like the ones arriving.
Expect each prompt to fan out behind the scenes, which is why adjacent questions belong in the set. A single user question is decomposed into several internal sub-queries before retrieval, so a page that answers only the headline question can lose to one that also covers the neighbouring ones. Our page on query fan-out covers how to write for that decomposition.
Run the same prompt set on every engine, unchanged, so that a difference between engines is a finding rather than an artefact. ChatGPT, Claude, Perplexity and Google AI Overviews disagree about who to name on identical questions, and that disagreement is one of the more actionable things a programme produces, because it points at a specific gap instead of a general one.
Engine baselines differ enough that a blended score hides the signal. BrightEdge's citation volatility analysis recorded materially different week-to-week churn by platform, with Perplexity at 72.8% and Google AI Overviews at 41.8%, and found that domains cited frequently moved about 0.7% week to week against more than 50% for sporadically cited ones. Averaging a volatile engine with a stable one produces a number that belongs to neither.
Weighting by where your buyers are is reasonable; changing the prompts is not. If referral evidence says one engine dominates your traffic, running that engine more often is a sensible allocation of runs. Previsible's 2026 AI traffic study, covering 6.77 million AI-driven sessions across 166 websites, found ChatGPT carrying 92.4% of trackable standalone referral traffic, which is an argument about run frequency rather than about prompt wording.
Keep the engine list stable too. Adding an engine mid-quarter changes the denominator of any blended figure, so add it as a separate series and let it accumulate its own history. Our guide to which AI engine to optimise for first covers how to sequence the work when you cannot do all four at once.
Change the prompt set when a question has stopped mattering to the business, not when the numbers are disappointing. Swapping prompts resets the series, and a reset series cannot show a trend, which means every change costs you the history you were about to use. Treat the set as an instrument rather than as content.
Additions are safer than replacements. When a new product or a new buyer question becomes important, add a clearly labelled second set and keep the original running in parallel, so the old series stays comparable while the new one accumulates. Reporting the two together, with set sizes stated, is the only way a reader can tell growth from redefinition.
Volume gives you one legitimate reason to grow the set. BrightEdge's volatility analysis found that the stability inflection sits near 50 citations, where weekly volatility falls from roughly 50% to 8%, so a company sitting well below that threshold is reading mostly noise and more observations genuinely help. Growing the set for that reason is defensible; growing it because the current prompts are hard is not.
Write the change rule down before the first run, because the pressure to edit the set arrives with the first bad month. A single line stating what would justify a change, who approves it, and that the old set keeps running for a quarter afterwards is enough. If you want the set built once and run properly rather than argued about quarterly, a free visibility assessment starts from your buyer questions and hands back the log with the runs recorded.
Ten prompts is enough to settle an internal argument and not enough to manage a programme. Run ten questions across four engines and you will usually discover whether competitors are named while you are absent, which is the finding that gets attention. What ten prompts cannot do is support a trend, because a single swapped citation moves the percentage by ten points.
Yes, kept as a separate group and never mixed into the headline number. Prompts containing your company name measure something different from prompts describing the problem you solve, and they are far easier to win, so blending them inflates citation share without telling you anything about acquisition. Keeping them apart also makes them useful as a check on answer accuracy.
As long as the question a person would actually ask, which in practice is longer than a keyword. According to our own Google Search Console export for signalscite.com, covering 93 queries in the 90 days to 27 September 2026, 38% of them run to six words or more, and the Bing queries that reach the site are frequently full sentences. Writing prompts as tidy keyword phrases measures a question nobody asks.
No, and using different sets makes the engines incomparable. Run one set everywhere so that a gap between ChatGPT and Google AI Overviews is a finding about the engines rather than an artefact of your own design. Engine-specific prompts belong in a separate side experiment if you want them at all.
Three to five named rivals, chosen because a buyer would consider them rather than because they are the biggest firms in the market. Relative position is the only reading that survives engine-side changes, since a change at the engine moves everybody in the retrieved set while a change on your pages moves you. Tracking twenty competitors produces a table nobody reads.
A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.
Request a free visibility assessment →