Answer engines cite Reddit heavily because it is a giant archive of real questions answered in plain language, and because Google and OpenAI have paid for clean access to it. The published numbers disagree wildly, from roughly 12% of citations to roughly 40% of answers, because each study counts something different. For a business the useful conclusion is narrow: engines reach for question-shaped, first-person, unpromotional text, and you can supply that on your own pages without buying an account and pretending to be a customer.
AI engines cite Reddit constantly because its archive matches what a retrieval system needs almost perfectly. A prompt arrives as a question in conversational language. Reddit is millions of questions in conversational language, each with answers written by people describing what actually happened to them. The vocabulary match that decides retrieval is effortless there, in a way that a polished product page never manages.
The format helps as much as the volume. Reddit's answers are first-person and specific, which reads as experience rather than marketing, and the threads are dated and continuously refreshed, so a 2026 question has 2026 replies. A single comment can also be lifted into an answer without the surrounding page, which is exactly the unit an engine wants to quote.
Commercial access matters as well, and it is the part most explanations omit. Reddit signed a content licensing agreement with Google reported at 60 million dollars a year in February 2024, and a comparable arrangement with OpenAI followed the same year, according to CNBC's reporting on the deal in July 2026. Licensed, structured, real-time access is a different proposition from crawling a site politely.
None of this reflects a judgement by the engines that Reddit is authoritative. Retrieval rewards documents that resemble the query and can be quoted cleanly, and Reddit produces those at a scale nobody else does. Our explainer on how AI engines choose which sources to cite covers the selection mechanism in more detail.
The percentage of AI citations coming from Reddit depends entirely on what the study counted, and the honest answer is a range rather than a number. Two different measurements circulate as though they were the same thing. One is Reddit's share of all citations issued. The other is the share of answers that contain at least one Reddit link. Because a single answer commonly cites five or ten sources, the second number is several times the first, by construction.
| Study | What it measured | Scale | Reddit result |
|---|---|---|---|
| 5W Research, 2026 | Share of all United States ChatGPT citations | Not a single-domain sample | Reddit 11.97%, Wikipedia 13.15% |
| Semrush Reddit study, 2026 | Which Reddit content gets cited, and why | 217,000 prompts, 248,000 Reddit URLs | Q&A threads dominate the cited set |
| Profound citation dataset | Domain share across ChatGPT, AI Overviews, Perplexity | 680 million citations, Aug 2024 to Jun 2025 | Reddit near the top, Wikipedia 7.8% on ChatGPT |
The lower end is the one to quote in a board meeting. 5W Research reported in 2026 that Wikipedia at 13.15% and Reddit at 11.97% together account for more than a quarter of United States ChatGPT citations, with major national newspapers absent from the top twenty entirely.
Volatility is the second thing the numbers hide. Reddit's share has swung hard between measurement windows as engines have changed their retrieval and their licensing, which means a strategy built on last quarter's citation mix is already out of date. Treat any single figure as a snapshot of one engine in one window, and rerun the measurement yourself rather than inheriting someone else's.
The Reddit content that gets cited is question-shaped, specific and unpopular, which surprises most marketers. Semrush analysed 248,000 Reddit URLs cited across Google AI Mode, Perplexity and ChatGPT Search, drawn from 217,000 unique prompts, and found that more than half of the cited content came from question and answer threads rather than news, images or discussion.
Engagement turned out to be close to irrelevant. In the same study, most cited posts had fewer than 20 upvotes and fewer than 20 comments, which tells you the retrieval layer is scoring textual relevance to the prompt rather than community approval. A thread with three replies that answers the exact question outranks a viral thread that answers a neighbouring one.
The engines also paraphrase rather than quote. Semrush measured the similarity between AI answers and the cited Reddit posts at roughly 0.53 to 0.54, against 0.04 to 0.05 similarity with the user's prompt, which is the signature of a model rewriting a source in its own words while keeping the substance. Being cited does not require being quotable word for word, but it does require being the clearest available statement of the answer.
What those findings describe is a specification, and it applies to any page rather than only to a forum thread: ask the question in the buyer's words, answer it directly and specifically, and keep the answer self-contained. Our AI citation checklist encodes the same specification, and it works the same way on a site you own.
Post on Reddit only if you can do it honestly, under your own name, in a community where your expertise is genuinely useful. Reddit's communities are moderated by people who have seen every promotional tactic, and accounts that arrive to recommend their own product get removed. A removed thread cites nobody, and a company caught staging recommendations acquires a permanent, searchable reputation problem.
What does work is unglamorous. Find the threads where people ask the question your product answers, reply with the specific detail you know because of your job, disclose your affiliation plainly, and accept that half your value will be delivered to people who never buy anything. Over time you become the person who answers that question, which is the thing an engine retrieves.
Watch two structural risks while you do it. Reddit's relationship with the engines is a commercial arrangement rather than a law of nature, and CNBC reported in July 2026 that the Google agreement was approaching expiry with renewal openly in question. And a citation of a Reddit thread that mentions you is not a citation of you: the link goes to Reddit, the traffic goes to Reddit, and your name appears only if the model chooses to repeat it.
Instead of chasing Reddit citations, make your own pages the cleanest available answer to the questions your buyers ask, and build third-party presence where it is legitimate to have one. Reddit is a symptom of what retrieval rewards, not a channel you have to buy into. Everything the engines like about a Reddit thread can be reproduced on a page you control.
Copy the format deliberately. Use the buyer's question as the heading. Answer it in the first sentence. Keep each section able to stand alone, since a reader arriving from a citation never saw the section above. Name the source of every number in the same paragraph as the number. Include the things a forum thread includes and a brochure omits: what it costs to change your mind, what the thing is bad at, and what happens when it goes wrong.
Then widen the surface honestly. Answer questions where your customers already ask them, keep your entries on review and comparison sites accurate and current, and publish the data only you have. Third-party corroboration is a real signal, weighted at 15% in the SIGNALS framework, and it is earned by being genuinely present in your category rather than by manufacturing mentions. Where your category's buying queries return listicles rather than vendor sites, the honest response is to publish a comparison of your own, as we did for AEO agencies.
Because Reddit holds an enormous archive of questions answered in ordinary language by people with no incentive to sell, which is close to the ideal shape for a retrieval system answering a conversational prompt. Licensing helps too: Reddit signed a content agreement with Google reported at 60 million dollars a year in February 2024 and a similar arrangement with OpenAI, which gave both companies clean, current access to that archive.
The published figures range from about 12% to about 40%, and the spread is a measurement artefact rather than a disagreement about reality. A study that counts the share of all citations puts Reddit near 12%, as 5W Research did for United States ChatGPT citations. A study that counts the share of answers containing at least one Reddit link lands far higher, because one answer can cite many sources. Always ask which of the two a number is before you quote it.
Question and answer threads, mostly, and not the popular ones. Semrush analysed 248,000 Reddit URLs cited across Google AI Mode, Perplexity and ChatGPT Search and found that more than half of cited Reddit content came from Q&A threads, and that most cited posts had fewer than 20 upvotes and fewer than 20 comments. Relevance to the question decides citation, not engagement.
Only if you can participate honestly under your own name in communities where your expertise is genuinely useful. Reddit communities remove promotional accounts quickly, and a removed thread cites nobody. The tactic that works is answering questions in your subject area with real detail and disclosing who you work for. The tactic that fails, publicly and expensively, is seeding fake recommendations of your own product.
Nobody should plan on it. Reddit's citation share has already swung sharply between measurement windows, and its commercial arrangements are not permanent: CNBC reported in July 2026 that the Google agreement was approaching expiry with renewal in question. A visibility strategy that depends on one third-party domain inherits that domain's negotiations, which is a good reason to own the answer on your own pages as well.
A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.
Request a free visibility assessment →