SIGNALS
Failure modes

Does gated content get cited by AI, or does the form make it invisible?

The short version

Gated content does not get cited, because an AI crawler cannot fill in a form, log in or buy a subscription. It requests a URL and reads whatever comes back, which on a gated asset is the landing page in front of the gate. The cost is measurable: the Paywall Penalty study published by Everything-PR and 5W Public Relations in July 2026 tested 40 queries and found that seven major paywalled publications took zero citations while open-web publishers took 91.3% of the citation inventory. The fix is not to ungate everything. Put the argument, the method and the numbers on an open page, and keep the form on the thing people actually want to download.

Does gated content get cited by AI?

Gated content is not cited by AI engines, because the crawlers that feed them never see it. A crawler makes a request and reads the response. It does not type an email address into a field, click submit, wait for a confirmation, or open the link in an email. Whatever sits behind that sequence is not part of the web as far as the engine is concerned.

What the crawler does read is the page in front of the gate, and that page is almost always thin by design. A title, a subtitle, three bullets about what you will learn, a cover image and a form. There is nothing there to quote. The 60 pages of original analysis your team spent a quarter producing contribute nothing, and the landing page that represents it says less about your expertise than a competitor's blog post.

The same logic covers every kind of wall, not just the classic lead form. A login wall on documentation, a customer portal holding your best technical answers, an email course, a webinar recording behind a registration page, a community that requires an account: each is content that exists for humans and does not exist for retrieval. Conductor's guide to gated content and AI discoverability, published in 2026, makes the same point about form fills, logins, passwords and subscriptions, which are all one obstacle from a crawler's point of view.

Worth separating from a decision you may also have made on purpose. Blocking GPTBot in robots.txt keeps crawlers away from pages that are otherwise public, and that is a policy choice with its own tradeoffs, covered in our piece on whether you should block AI crawlers. A gate is different: nobody chose to be invisible to AI, it is a side effect of a lead capture decision made before answer engines mattered.

How much visibility does a paywall actually cost?

The cost has been measured on the hardest version of the problem, which is news. The Paywall Penalty study, published by Everything-PR with 5W Public Relations in July 2026, ran 40 queries across AI answer engines and recorded which sources were cited. Seven major paywalled publications, including The Wall Street Journal, the Financial Times, Bloomberg, The New York Times, The Washington Post, The Economist and The Atlantic, captured zero citations. Open-web publishers took 91.3% of the citation inventory on those same queries.

Those are outlets with enormous authority, decades of reporting and the kind of brand recognition no B2B company will ever have. If a paywall can zero out The Economist on a query it is qualified to answer, a lead form in front of a 12 page report is not going to fare better. The gate outranks the reputation.

There is a second-order effect that hurts more than the missing citation. AI answers describe categories using whatever material is retrievable, so a market where your competitors publish openly and you publish behind a form ends up with your competitors' framing as the default description of the category. Buyers then arrive at the question with your rival's vocabulary already in their heads, which is the thing the scoring model treats as decisive: the Discovered Labs 2026 analysis of more than 2 million citations found vocabulary alignment between page and query to be the only page-level signal that survived domain-level controls, at an effect size of 0.37, as set out on our methodology page.

One caveat that gets misread. Engines can still describe a gated report if other people wrote about it, because a press release, a summary in a newsletter or a LinkedIn post carries fragments into the index. The citation then goes to whoever wrote the summary. Your research is being used and someone else is getting named for it.

What does an AI crawler see when it hits your gate?

An AI crawler sees your gate as a short HTML document that promises something and delivers nothing. Different gates fail in slightly different ways, and the differences matter when you are deciding which ones to keep, so the table below separates them by what the crawler actually gets back.

Gate type What the crawler receives Can it be cited Usual verdict
Lead form before a PDF Landing page copy only, typically under 200 words No Publish the findings as HTML, keep the PDF as the download
Hard paywall Headline and standfirst, sometimes a first paragraph Rarely, and only the free fragment Own an open version of the argument somewhere you control
Login wall on docs or support Sign-in page No Open the documentation, it is the most quotable asset you have
Soft gate, full text then a prompt The full text, if it is in the initial HTML Yes Keep, this is the pattern that works
Webinar or video behind registration Registration page No Publish a transcript or a written summary openly

Two rows in that table are worth reading together. A soft gate that renders the full article and then asks for an email after a scroll is readable, because the text was in the response the crawler received. A gate that loads the article by JavaScript only after the email is accepted is not, for the same reason a client-rendered page is not: the text was never in the HTML. Our piece on whether AI crawlers read JavaScript covers that mechanism, and it applies to gates as much as to frameworks.

The PDF question comes up constantly and has a boring answer. An open PDF can be fetched and parsed by most crawlers, so it is not invisible, but it is a weak candidate: headings are unreliable, text may sit in columns or inside images, and the file almost never uses the wording a buyer typed. The same content as an HTML page with question-shaped headings is a far stronger citation target, with the PDF offered beside it for people who want something to print.

Should you ungate everything, then?

Ungating everything is the wrong conclusion, and it is the one most articles on this subject jump to. A form is not a mistake. It is a trade: you give up retrievability in exchange for contact details, and the trade is good whenever the thing behind the form has value that a summary cannot carry.

The rule that survives contact with a real marketing calendar is to ungate the argument and gate the artefact. Your position on how a problem should be solved, the method you used, the numbers you found and the conclusion you drew are the argument. They are also exactly what an engine needs in order to name you when someone asks about that problem. The artefact is the thing a person wants to hold: the template, the spreadsheet model, the calculator, the benchmark dataset, the audit. Nobody quotes a spreadsheet, so gating one costs you nothing in citations.

Applied to a typical B2B research report, that means the findings page is open and carries the headline numbers, the method, the sample size and the charts, and the form sits on the full dataset or the slide deck. You lose the download from people who would have skimmed the PDF and never replied. You gain a page that can be cited on every query the research answers, and a set of buyers who arrive already knowing what you found.

There is a demand-side reason to do this beyond retrieval. Skyword's 2026 guide to ungating for AI search describes the same hybrid: open the top of the funnel for discoverability, keep gates where the asset is genuinely worth an email. Buyers who self-educate first tend to show up later in the process and further along, which changes what the form is for rather than removing the need for one.

How do you keep the leads and the citations at the same time?

Keeping both means restructuring one asset rather than rewriting your whole content programme. Take your best gated report, the one sales actually uses, and build an open findings page in front of it. The page carries the question the research answers as its H1, the headline numbers with their method, a section per finding, and a clear statement of who ran the research and when. The form stays, one click away, attached to the full document.

Write the open page for extraction rather than for persuasion. Sections of 200 to 400 words that stand alone, headings phrased the way a buyer would ask the question, and a direct answer in the first sentence under each heading. The Princeton GEO study, published by Aggarwal and colleagues at ACM KDD 2024, tested 22 content strategies across 10,000 queries and found that adding statistics with citations raised visibility in generative engines by 41% and adding expert quotes by 28%, which is a fair description of what a research findings page already contains.

Then check what the gate is doing to the rest of the site. Documentation, pricing detail, integration lists and technical specifications are the pages that get retrieved for commercial questions, and each one behind a login is a question your competitors get to answer instead. Most teams find at least one of these by accident when they run the audit in our guide to checking whether AI can read your website.

Measure the swap rather than assuming it. Record which prompts name you before the change, publish the open page, and re-run the same prompts on a schedule so the change shows up as a series rather than a feeling. Citation counts move week to week for reasons that have nothing to do with you, which is why our page on measuring AI citation decay treats a single check as worthless.

What about AI browsers that read paywalled pages for the user?

AI browsers are a real exception and a misleading one. An agentic browser runs inside a person's own session, with their cookies and their subscription, so it can read what that person can read. The Columbia Journalism Review's analysis of how AI browsers get past blockers and paywalls, and testing reported by Cybernews, both described agentic browsers retrieving subscriber-only article text, including a subscriber-exclusive piece from MIT Technology Review.

That capability does nothing for the problem this page is about. A browser reading your gated report for one logged-in customer is not an index, and the engine that answers a cold question from a buyer who has never heard of you is working from crawled, retrievable pages. Being readable by an agent that a customer already pointed at you is not visibility, it is service.

The publisher fight over this is worth watching for a different reason: it tells you where access rules are heading. Crawler policy, licensing deals and pay-per-crawl schemes are all moving, and any of them could change what a gate means in a year. What has not moved is the retrieval mechanic, which is that an engine can only cite text it has fetched.

For a company selling anything other than news, the practical read is simple. Treat AI browsers as a nice property of your product documentation and make the decisions on this page for the crawler case, because the crawler case is the one that decides whether you get named to a stranger.

What else do people ask about gated content and AI search?

Can ChatGPT read content behind a lead form?

No. A crawler requests a URL and reads what comes back, and it cannot type an email address into a form, click submit, or wait for a download link. What ChatGPT reads is the landing page in front of the gate, which is usually a title, three benefit bullets and a form. The report itself is never fetched, so nothing in it can be quoted or attributed to you.

Does a paywall stop AI engines citing you?

A hard paywall stops the crawlers that build AI answers. The Paywall Penalty study published by Everything-PR and 5W Public Relations in July 2026 tested 40 queries and found that seven major paywalled publications, including The Wall Street Journal, the Financial Times, Bloomberg and The New York Times, captured zero citations, while open-web publishers took 91.3% of the citation inventory on the same queries.

Should I ungate all of my content for AI visibility?

No. Ungate the argument and keep the artefact. The claims, the method and the numbers belong on an open HTML page an engine can read and cite. The template, the calculator, the benchmark dataset and the assessment are worth a form because people will trade an email for a tool they cannot get elsewhere, and losing a citation on a spreadsheet costs nothing.

Do gated PDFs get cited by AI engines?

A gated PDF is not fetched at all, so it cannot be cited. An ungated PDF can be read by most crawlers, but it is a weaker candidate than an HTML page: headings are less reliable, the text may sit in columns or images, and the file rarely carries the question wording a buyer used. Publishing the same material as an HTML page and offering the PDF alongside it is the version engines can quote.

What about AI browsers that read paywalled pages for the user?

An AI browser acting inside a logged-in session is a different system from a crawler building an index. The Columbia Journalism Review and Cybernews both reported in 2025 and 2026 that agentic browsers retrieved subscriber-only article text, including a subscriber-exclusive piece from MIT Technology Review. That behaviour helps one reader at a time and does nothing for the index that decides who gets named when a buyer asks a question cold.

Related guides

The Assessment

Find out who gets named when buyers ask about your category.

A free visibility assessment runs your buyer questions across the four engines, records who is cited and from which page, and shows where your own pages are being passed over.

Request a free visibility assessment →
SIGNALS · A BlackSig Systems company