Small sites do get cited, and the published evidence on how much authority helps genuinely disagrees. Surfer analysed roughly 5 million citation source URLs from 20,000 prompts and found that domain authority correlated with citation at close to zero. SE Ranking's dataset of 129,000 domains found the opposite shape, reporting that sites above 350,000 referring domains averaged 8.4 ChatGPT citations against 1.6 to 1.8 for sites under 2,500. Both can be true, because authority buys you candidacy and the page decides the rest. The lever a small site actually controls is specificity: answering a narrow question completely, in the words buyers use, on a page a crawler can read.
Small sites are cited by AI engines routinely, and the citation data makes that easier to show than most people expect. Ahrefs found that only 12% of AI-cited URLs ranked in Google's top 10 for the original prompt, which means the great majority of cited pages are not the established page-one incumbents for the question being asked.
A second Ahrefs study points the same way from the other end. Looking at 863,000 keywords and 4 million AI Overview URLs, it found that 38% of cited pages also ranked in the top 10 for the same query, down from 76% in an earlier run, with the rest split almost evenly between positions 11 to 100 and beyond position 100. Engines are reaching deep into the index for source material.
What a small site cannot do is win on presence alone. Being cited requires being retrieved, and being retrieved requires the crawler to reach the page, the page to answer the question in extractable form, and the wording to match how the question was asked. None of those three has a minimum domain size attached, and all three are commonly failed by large sites.
The realistic expectation is narrow wins rather than broad ones. A small site rarely becomes the default source for a category term, and regularly becomes the cited source for a specific question nobody large has answered properly. That is a smaller prize with a much higher hit rate, and for most businesses it is also the prize that produces enquiries.
Domain authority predicts AI citations weakly in some datasets and clearly in others, and the disagreement is worth understanding before you plan around either. Surfer's analysis of roughly 5 million unique citation source URLs from 20,000 tracked prompts across AI Mode, AI Overviews, ChatGPT and Perplexity found domain authority correlating with citation frequency at close to zero, testing three separate measures including PageRank and harmonic centrality from Common Crawl.
| Study | Scale | What it measured | Finding |
|---|---|---|---|
| Surfer, 2026 | 20,000 prompts, about 5 million citation URLs | Correlation between domain authority metrics and citation | Close to zero, slightly negative |
| SE Ranking, 2026 | 129,000 domains | Average ChatGPT citations by referring-domain count | 8.4 citations above 350,000 referring domains, 1.6 to 1.8 below 2,500 |
| Discovered Labs, 2026 | More than 2 million citations, domain fixed effects | Which page-level signals survive controlling for domain | Vocabulary alignment survives, effect size 0.37 |
| Ahrefs, 2026 | 863,000 keywords, 4 million AI Overview URLs | Overlap between cited pages and Google's top 10 | 38% of cited pages ranked top 10, down from 76% |
The SE Ranking study of 129,000 domains found that the opposite held, with sites above 350,000 referring domains averaging 8.4 ChatGPT citations while sites under 2,500 referring domains averaged 1.6 to 1.8. Link diversity showed the clearest correlation of the factors it tested.
Both results can hold at once, because they measure different things. A correlation across all cited URLs asks whether the cited pages skew authoritative, and finds they do not much. A comparison of averages per domain asks how often a domain gets cited at all, and large domains publish vastly more pages on vastly more topics, so they accumulate more citations without any single page being favoured. The Discovered Labs study is the one that separates the two, because controlling for domain leaves vocabulary alignment as the page-level signal that survives.
Ranking on page one is neither required nor sufficient for citation, and the gap has widened. The Ahrefs study of 863,000 keywords and 4 million AI Overview URLs found that 38% of cited pages also ranked in Google's top 10, down from 76% in its earlier run, with roughly 31% coming from positions 11 to 100 and another 31% from beyond position 100.
Other measurements put the overlap lower still. BrightEdge reported around 17% of AI Overview citations overlapping with organic top 10 results in an analysis published earlier in the same year, and the ConvertMate GEO Benchmark 2026 found 83% of AI citations coming from pages outside Google's top 10, a figure summarised on our methodology page. The methodologies differ and the direction does not.
The reason is structural rather than mysterious. A generated answer is assembled from several sub-queries, not from one ranking, so a page that answers a narrow part of the question well can be pulled in even when it would never rank for the headline term. Our explainer on query fan-out covers how one prompt becomes many hidden searches.
For a small site the consequence is encouraging and slightly uncomfortable. The encouraging half is that you do not have to outrank an incumbent to be quoted alongside them. The uncomfortable half is that classic rank tracking stops telling you whether the work is succeeding, so the measurement has to change at the same time as the content does.
Big domains keep showing up because AI citation is concentrated, and pretending otherwise would be dishonest. 5W Research's Q1 2026 citation audit, built on a Similarweb dataset of roughly 600,000 United States citation events, put Wikipedia at 13.15% and Reddit at 11.97% of ChatGPT citations, so two domains account for more than a quarter of what one engine cites, as our page on why AI engines cite Reddit sets out in more detail.
Concentration of that kind is not evidence that authority is being rewarded directly. Wikipedia and Reddit are cited heavily because they hold broad, plainly worded, question-shaped content on almost every topic, which is exactly what retrieval favours, and because they are easy for a crawler to read. The lesson available to a small site is in the shape of that content, not in the size of the domain.
The concentration also varies sharply by question type. Broad definitional prompts pull encyclopedic sources. Specific commercial prompts pull comparison pages, vendor documentation and trade sources, which is where a specialist site has a genuine chance, and where the citation is worth more commercially anyway.
So the picture to hold is two-layered. Aggregate citation counts are dominated by a handful of enormous domains, and the citations that decide a purchase are spread much more widely. Optimising against the first number is hopeless for a small site. Optimising against the second is ordinary work.
A small site can be completely specific, and a large one usually cannot afford to be. Answering one narrow question exhaustively, with the numbers, the caveats, the failure cases and the vocabulary a practitioner would use, is cheap for a specialist and expensive for a publisher covering ten thousand topics at scale.
Specificity is also the thing the strongest evidence rewards. The Discovered Labs 2026 analysis of more than 2 million citations found that vocabulary alignment between page and query was the only page-level signal to survive domain fixed-effects controls, at an effect size of 0.37, which is why it carries 35% of the SIGNALS weighting on our methodology page. Matching how your buyers actually phrase the question is not a soft factor, it is the measured one.
Controlled testing points the same way about evidence. The GEO experiment published by Aggarwal et al. at ACM SIGKDD 2024 found that adding verifiable statistics, credible quotations and cited sources raised visibility in generated answers by up to roughly 40% across its 10,000-query benchmark, while keyword stuffing did nothing. A small site can source every claim on a page in an afternoon. Where those gains land in the pipeline, and why the methods that failed were the ones aimed at ranking rather than generation, is set out in how much structure changes the probability of being cited.
The last advantage is speed. A specialist can publish an answer to a question that appeared this month while a large publisher is still scheduling it, and being the first credible source on a question is how small domains enter citation sets they could not otherwise reach.
A small site should fix eligibility first, because nothing downstream matters if the crawler leaves with an empty page. Confirm that the AI crawlers are allowed in robots.txt, that your pages return their content in the initial HTML rather than after JavaScript runs, and that the pages you care about are not orphaned. Our guides on whether AI crawlers read JavaScript and how to check if AI can read your website cover the tests.
Then pick questions instead of keywords. Write down the ten questions your buyers ask before they buy, in their words, and check which of them your site answers in a form an engine could lift: a heading worded as the question, the answer in the first sentence under it, and a section that stands alone for a reader who arrived there directly.
Third, source every claim on the page in the same paragraph as the claim. A number with a named source beside it survives a model's shallow verification check, and a bare number reads as invented. Where your page is too thin to support a claim, leave the claim out rather than padding, since a page that overstates is a page that gets checked and dropped.
Finally, get named somewhere other than your own domain. The ConvertMate GEO Benchmark 2026 found brands mentioned on third-party domains receiving 6.5 times more citations than brands present only on their own site, and our page on getting into the lists AI engines quote covers the practical version of that work for a small company.
The evidence disagrees, which is itself the answer. Surfer analysed about 5 million citation source URLs from 20,000 prompts and found that domain authority correlated with citation at close to zero. SE Ranking's 129,000-domain study found that sites above 350,000 referring domains averaged 8.4 ChatGPT citations against 1.6 to 1.8 for sites under 2,500. Large domains accumulate more citations overall, while individual pages are not chosen for their domain.
A new website can be cited once its pages are crawlable and it answers a question better than what already exists, though it starts with no third-party corroboration, which slows things down. Expect narrow questions to come first and category-level prompts to take much longer, and measure with a repeated prompt set rather than a single check.
No. Ahrefs found that 38% of AI Overview citations came from pages that also ranked in Google's top 10, down from 76% in its earlier run, with the remainder split between positions 11 to 100 and beyond position 100. A separate Ahrefs analysis found only 12% of AI-cited URLs ranking in the top 10 for the original prompt.
Wikipedia and Reddit are cited heavily because they carry broad, plainly worded, question-shaped content on nearly every topic and are easy to crawl. 5W Research's Q1 2026 citation audit, drawn from roughly 600,000 United States citation events, found that Wikipedia accounted for 13.15% and Reddit for 11.97% of ChatGPT citations. Concentration is highest on broad definitional questions and much lower on specific commercial ones, which is where a specialist site competes.
Crawler access and rendering, because a page the crawler cannot read is not eligible for any citation at all. Check robots.txt for the AI user agents, then fetch a page without JavaScript and confirm your answer is in the HTML that comes back. After that, reword headings as the questions buyers ask and put the answer in the first sentence under each one.
A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.
Request a free visibility assessment →