Schema markup does not appear to increase AI citations. The best controlled study available, run on 1,885 pages against matched controls, found no gain in ChatGPT or AI Mode and a small drop in AI Overviews. Cited pages do carry schema more often, but that is a property of the sites that bother with schema rather than an effect of the markup. Keep your structured data for search features and machine-readable facts, and spend the next hour on vocabulary and sourced claims instead.
Schema markup does not measurably help you get cited by AI, on the strongest evidence published so far. That conclusion is narrower than it sounds, so it is worth stating precisely: adding JSON-LD to a page that lacked it does not raise how often ChatGPT, Google AI Mode or AI Overviews quote that page. Structured data still does the job it was designed for in classic search, and it still makes facts about your organisation legible to machines.
The confusion comes from a genuine correlation. Pages that appear in AI answers do carry structured data more often than pages that do not, and that observation has been recycled into a thousand posts recommending schema as an AEO tactic. Correlation of that kind is exactly what a controlled test exists to break, and when someone finally ran the test, the effect went away.
This matters commercially because structured data is the easiest thing on an AEO checklist to sell and to implement. A plugin emits it, an audit confirms it, everyone reports progress, and the citation rate does not move. Meanwhile the two changes with measured effects, matching buyer vocabulary and grounding claims in named sources, need a writer and an editor rather than a plugin, so they get postponed.
The Ahrefs schema study tracked pages before and after they added JSON-LD and compared them against pages that did not. Louise Linehan and Xibeijia Guan followed 1,885 pages that added schema between August 2025 and March 2026, matched them against roughly 4,000 control pages, and measured citation changes across Google AI Overviews, Google AI Mode and ChatGPT.
Nothing moved in the direction the industry expected. Pages that added schema performed slightly better than controls in AI Mode and ChatGPT by margins small enough to be noise across thousands of URLs, while Search Engine Journal's write-up of the study notes a 4.6% decline in AI Overview citations that was small but statistically significant against the matched controls.
A study is not a law, and this one has limits worth naming. It measures adding schema to pages that already existed, not building a page with schema from the start. It covers three surfaces, not Perplexity or Claude. And a 4.6% decline is more plausibly a quirk of what those particular pages were doing than evidence that markup repels engines. What the study rules out is the large positive effect that AEO checklists have been promising, and it rules it out at a sample size no single agency can match.
Not every experiment agrees on every surface. Otterly.ai's own schema experiment reported movement in AI Overviews and no isolated connection to ChatGPT visibility, which is consistent with schema affecting eligibility for Google surfaces rather than the choice of who gets quoted. Where two studies disagree, the one with matched controls and a four-figure sample carries more weight.
Cited pages carry schema more often because structured data is a marker of a well-run site, and well-run sites do the things that actually earn citations. In the Ahrefs data, pages appearing in AI answers were almost three times more likely to have JSON-LD than pages that did not appear, which reads as a strong signal until you ask what kind of site emits JSON-LD in the first place.
The answer is: sites with a maintained CMS, a technical owner, an editorial process and a habit of publishing sourced material. Every one of those traits independently raises the chance of being retrieved and quoted. Markup rides along with them without doing the work, which is the textbook shape of a confounded variable, and it is the same trap that made domain authority look like a content signal for years.
This is why the SIGNALS methodology weights signals by measured effect after controls rather than by how often they appear on cited pages. Discovered Labs analysed 2 million AI citations across 10,000 pages with domain fixed effects and found content alignment was the only page-level signal to survive, at an effect size of beta = +0.37. Almost everything else that correlates with citation collapsed once the domain was controlled for.
The practical test for any AEO recommendation is therefore one question: has anyone measured this against a control, or is it a description of what cited pages happen to have? Most of the checklist circulating in 2026 is the second kind. Our methodology page sets out which signals we weight and what evidence each weight comes from.
Google says no special structured data is required for AI Overviews or AI Mode, and has said for years that structured data is not a ranking factor. Google's guidance on AI features in Search states plainly that there is no special schema.org markup you need to add to be eligible, and points site owners at the same fundamentals that govern ordinary indexing.
The role structured data does play is eligibility for search features. John Mueller has described it as getting you into the party rather than getting you a good seat: markup makes a page qualify for a rich result, and other factors decide whether it is shown. That framing is a useful one to carry into AEO, because it separates what a format unlocks from what a selection algorithm rewards.
Where Google is quieter is on how much its systems parse markup when choosing sources for a generated answer, and no amount of confident blogging changes the fact that this is undisclosed. When the vendor says the thing is not required and the controlled test finds no effect, the honest conclusion is that schema is a search-features tool that has been reclassified as an AI tactic by people with markup to sell.
Instead of adding more schema, spend the same hour on the two things with measured effects: the words the page uses, and the evidence it carries. Both are visible to a reader, both are checkable, and both survive engine updates because they change what the page actually says rather than how it is annotated.
| Change | Evidence it moves AI citation | Effort | Verdict |
|---|---|---|---|
| Add JSON-LD schema | Controlled test found no gain, small AI Overview decline | Low, usually automated | Keep it, do not count on it |
| Match buyer vocabulary in titles and headings | Only signal surviving domain controls, beta = +0.37 | Medium, needs research | Do this first |
| Add statistics with named sources | Among the strongest tactics in the Princeton GEO trials | Medium, needs real sourcing | Do this second |
| Answer the heading in the first sentence | Improves extractability of the quoted unit | Low, editing only | Do this everywhere |
| Let retrieval crawlers in | Prerequisite: a page not fetched cannot be cited | Low, one file | Check before anything else |
Evidence is the half most sites skip. In Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, content changes tested across roughly 10,000 queries lifted visibility in generated answers by up to 41%, and the strongest performers were citing sources, adding quotations and adding statistics. A statistic without a source in the same paragraph does not count, because a reader cannot check it and an engine has nothing to attribute.
Keep the schema types that either qualify a page for a search feature or state a fact about you that no prose can express unambiguously. Article and its subtypes carry author, publisher and dates. Organization and Person tie your entity to profiles elsewhere. Product, Review, Event, Recipe and JobPosting all gate rich results in classic search. FAQPage and HowTo have narrower rich-result eligibility than they once did, but both express structure cleanly and cost nothing to emit.
One rule matters more than the choice of type: the markup must match the visible copy exactly. Schema that describes questions and answers not present on the page is a quality problem in classic search and a credibility problem everywhere else, and it is the most common defect in auto-generated FAQ blocks. If the visible text changes, regenerate the markup in the same commit.
This site emits Article, FAQPage and BreadcrumbList on every resource page, and every FAQ answer in the markup is the same sentence a reader sees. The reason is not that it wins citations. It is that the cost is zero once the template emits it, the facts are true, and search features are worth qualifying for. That is the right size of claim to make about structured data in 2026.
No, not on the current controlled evidence. The Ahrefs study tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against roughly 4,000 matched control pages and found that citations did not rise meaningfully in ChatGPT or Google AI Mode, and a small decline of 4.6% in AI Overviews. Schema remains worth having for search features and for machine-readable facts, but it is not the lever that decides whether an engine quotes you.
Because schema tends to sit on well-maintained sites, and those sites do everything else well too. In the Ahrefs data, cited pages were almost three times more likely to carry JSON-LD than uncited pages, which is exactly the pattern you would expect from a confounded variable: the same teams that add structured data also write clearer pages, publish more sources and earn more links. Adding markup to a thin page does not import the rest of that behaviour.
No. Google's own AI features documentation states that there is no special schema.org structured data needed to appear in AI Overviews or AI Mode. Google has also said for years that structured data is not a ranking factor and instead makes a page eligible for particular search features. Eligibility is a real benefit, but it is a different thing from being chosen as a source.
No. Keep it. Structured data still drives rich results in classic search, still expresses author, organisation and product facts in a machine-readable form, and costs nothing once your templates emit it. The argument here is about where marginal effort goes, not about deleting work already done. Removing correct markup buys you nothing and loses you search features you currently qualify for.
Vocabulary that matches how buyers ask, direct answers in the first sentence of each section, and checkable evidence. Discovered Labs' 2026 analysis of 2 million citations found content alignment was the only page-level signal that survived domain controls, at an effect size of beta = +0.37. The Princeton GEO study measured visibility gains of up to 41% from citing sources, adding quotations and adding statistics. Both are changes to the words on the page rather than to the markup around them.
A free visibility assessment scores your pages on the seven dimensions, weighted by measured effect rather than by checklist folklore, and shows what to fix first.
Request a free visibility assessment →