SIGNALS
Engine behaviour

How long does it take for ChatGPT to index a new page?

The short version

Indexing takes days on a site the crawlers already visit often and weeks on a site they visit rarely, and being quoted in an answer is a later, separate event. Keep three things apart: a crawl is a bot fetching the page, indexing is the page entering a retrievable store, and citation is an answer quoting it. The crawler that decides whether you appear in ChatGPT search is OAI-SearchBot, not GPTBot, and OpenAI's documentation states they are controlled independently, so a robots.txt rule written as a training-data policy can leave citations intact or remove them entirely depending on which name was used. According to DigitalApplied's 2026 study of site logs, GPTBot revisited high-traffic pages at a median of about every 2.4 days and returned roughly 47% faster to pages serving a fresh last-modified header. Confirm the crawl in your server logs before concluding anything from asking the engine.

How long does it take for ChatGPT to index a new page?

Indexing a new page usually takes days rather than hours for a site the crawlers already visit regularly, and can take several weeks on a site they visit rarely. Being quoted in an answer is a separate event that happens later, and the gap between the two is where most of the confusion about this question lives.

Three things have to happen in order, and each has its own timescale. A crawler has to fetch the page. The fetched page has to enter the index that ChatGPT's search feature queries. And a buyer's question has to retrieve that page and the model has to choose to quote it. A page can clear the first two in a week and never clear the third.

Domain history is the strongest predictor of the first step. Sites that publish often and are already fetched daily tend to get new URLs picked up within a day or two; sites that publish quarterly can wait weeks, because nothing tells the crawler there is a reason to come back. Nothing about the quality of the new page changes this, which is why "we published it last Tuesday and ChatGPT has never heard of it" is usually a crawl-frequency story rather than a content story.

The practical implication for planning is to stop treating publication as the milestone. A page published on the first of the month and crawled on the ninth has not been live in any sense that matters to an answer engine for the first eight days, so a two-week check for citations is measuring a page the engines may only just have read.

What is the difference between being crawled, indexed and cited?

Crawled means a bot fetched the page, indexed means the page is in a retrievable store, and cited means a generated answer quoted it and showed the link. Keeping the three apart is the single most useful habit in this area, because the fixes for each failure are different and people regularly apply the third fix to the first problem.

A crawl is visible to you and nothing else. Server logs show the user agent, the URL and the timestamp, which makes crawling the only one of the three you can confirm directly without asking an engine anything. A page can be crawled hundreds of times and be absent from every answer.

Indexing is invisible from outside and has to be inferred. Nobody publishes an index status tool for ChatGPT, so the available evidence is indirect: whether the page can be surfaced at all when you ask a question narrow enough that only your page could answer it, and whether its content appears in the engines that share an index.

Citation is the only one of the three that has commercial meaning and the one most affected by things other than timing. Retrieval has to select your page out of a candidate set, and the model has to prefer it to the others, which depends on how well the page's language matches the question and whether any section can be lifted out and still make sense. Our page on how AI engines choose which sources to cite covers that selection step.

The reason this matters for diagnosis is that each stage fails differently. Crawl failures are robots.txt rules, blocked user agents and firewall rules. Index failures are usually rendering: the bot fetched something with no readable content in it. Citation failures are vocabulary, structure and competition.

Which OpenAI crawler decides whether you appear in ChatGPT?

OAI-SearchBot is the crawler that decides whether you appear in ChatGPT's search answers, and it is not the one most companies have been thinking about. According to OpenAI's own documentation on its crawlers, GPTBot collects content that may be used to train its models, OAI-SearchBot indexes pages for ChatGPT's search and citation features, and ChatGPT-User fires only when a live user asks ChatGPT to fetch a specific URL.

The consequence is a trap that has cost companies real visibility. Blocking GPTBot in robots.txt, which plenty of organisations did as a training-data policy, does not block ChatGPT search. Blocking OAI-SearchBot does, and OpenAI's documentation states that sites opted out of it will not be shown in ChatGPT search answers. A legal or brand decision about training data can therefore be implemented correctly and leave citations untouched, or implemented carelessly and remove you from the answers entirely.

Checking which of the three has visited you is the fastest useful diagnostic on this whole topic. A log full of GPTBot and empty of OAI-SearchBot means your content may be in a training corpus and absent from the search index, which looks like invisibility and is a crawl-level problem. Our page on whether you should block AI crawlers like GPTBot goes through the policy decision in full.

ChatGPT-User appearing in your logs means something different again and is frequently misread as evidence of indexing. A fetch by ChatGPT-User means somebody, possibly you, asked ChatGPT to read that URL in that moment. A page can be fetched this way all day and remain absent from the index that answers questions nobody pasted a link into.

How often does GPTBot come back to a page?

GPTBot returns on a schedule set by how much reason it has to, which in practice means frequently for pages on busy, frequently updated domains and rarely for static deep pages. According to DigitalApplied's 2026 agentic crawler study, a 30 day analysis of site logs, GPTBot revisited high-traffic pages at a median of about every 2.4 days, and returned roughly 47% faster to pages serving a fresh last-modified header.

The crawl volume behind that is large. A Vercel and MERJ analysis of AI crawler traffic, published in December 2024, measured GPTBot at roughly 569 million requests a month across the Vercel network and ClaudeBot at about 370 million, against 4.5 billion from Googlebot over the same period. Volume at that scale is not evenly spread: it concentrates on domains that already produce a lot of fetched-and-used content.

Crawling is bursty, not periodic. Logs commonly show a cluster of fetches over a few days, then silence for weeks, then another cluster. Averaging those bursts into a single interval produces a number that describes no actual behaviour, which is why a one-week log sample tells you much less than a 60 day one.

Other engines behave differently. PerplexityBot tends to crawl more often because it is powering live retrieval instead of periodic training collection, which is part of why the same new page shows up in Perplexity before ChatGPT on many sites. Our comparison of which AI engine to optimise for first covers what those differences mean for sequencing.

What you can influence is the reason to return. A sitemap with accurate last-modified dates, internal links from pages that already get fetched often, and a genuine publishing cadence all raise crawl frequency. Nothing you can do forces a recrawl on demand.

Why can a page be crawled and still never be cited?

A crawled page goes uncited for three common reasons, and only one of them is about content quality. The first is that the crawler fetched the URL and found nothing readable in it. The second is that the page's language does not match how the question arrives. The third is that it is simply beaten by other pages in the candidate set.

Rendering is the failure that produces the most confusing logs, because the fetches look healthy. The Vercel and MERJ analysis cited above found no evidence of JavaScript execution by the major dedicated AI crawlers, including GPTBot, ClaudeBot and PerplexityBot, so a page whose content is injected client-side is fetched successfully and read as empty. Our page on whether AI crawlers read JavaScript covers how to check what they actually receive.

Structure decides whether anything on the page is usable once it is readable. The ConvertMate GEO Benchmark 2026, an observational study of 12,500 queries across 8,000 domains, found 68.7% of cited pages using a strict H1 to H2 to H3 hierarchy, and 83% of citations coming from pages outside Google's organic top 10. A section that opens with "This means that" cannot be lifted into an answer, because the fragment has lost its subject.

Variance is the last reason and the one most often mistaken for a failure. Researchers at the University of St. Gallen found in "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv:2604.07585, April 2026) that identical prompts re-run at the same moment overlapped in their cited sources only 32% to 43% of the time. A page absent from one answer may well be present in the next, so a single check is not evidence of anything.

What makes a new page get indexed faster?

Getting indexed faster comes down to giving the crawlers a reason to come back and making sure what they fetch is readable when they do. The levers below are ordered by how much they change, and the first two are worth more than everything underneath them.

Lever What it changes How much it matters
Server-side rendered HTML Whether the fetched page contains readable content at all Decisive: a client-rendered page is read as empty
Allowing OAI-SearchBot in robots.txt Whether the page is eligible for ChatGPT search answers Decisive: blocking it removes you from answers
Accurate last-modified headers and sitemap dates How quickly crawlers return after a change Large: fresh headers correlated with faster revisits in log studies
Internal links from frequently crawled pages Whether the new URL is discovered at all Large on sites with deep, rarely visited sections
A real publishing cadence Baseline crawl frequency for the whole domain Compounding, over months and not weeks
Submitting URLs to Bing and IndexNow Speed into an index that several assistants draw on Moderate, and cheap enough to be worth doing
Schema markup Machine-readable facts about the page Small for indexing speed, useful for other reasons

Note what is absent from that list. There is no submission endpoint that forces ChatGPT to index a URL, no paid expedite, and no setting that guarantees a recrawl date. Anyone selling one of those is selling something that does not exist, and our page on whether schema markup helps AI citations covers the one item on the list that is most often oversold.

How do you check whether ChatGPT has seen your page?

Checking whether ChatGPT has seen your page is done in your server logs first and in the product second, in that order, because only the logs give you a fact. Filter access logs for the OpenAI user agents by name and look at which URLs were fetched, by which bot, and when. A page with no OAI-SearchBot fetch has not entered the search index, whatever else is true.

The product-side test needs care to be meaningful. Asking ChatGPT a question so specific that only your page could answer it is a reasonable probe, but pasting your URL into the chat is not: that triggers a live fetch by ChatGPT-User and proves only that the page is reachable right now. Many teams have concluded they are indexed on the strength of exactly that test.

Cross-checking against Bing is worth doing because several assistants lean on that index. If a page is missing from Bing entirely, the likeliest explanations are the same ones that would keep it out of an AI search index: blocked crawling, client-side rendering, a noindex left in place from staging, or a canonical pointing somewhere else.

Repeat any product-side check before believing it. The St. Gallen variance figures quoted above mean a single negative result is weak evidence, and the honest version of this test is the same narrow question asked several times across a few days. Our guide on checking whether AI can read your website sets out the full sequence, including what the crawlers receive rather than what a browser shows.

One more check is worth running on any site more than a year old: look for rules nobody remembers adding. Blanket AI crawler blocks added during a policy discussion in a previous year, and never revisited, are a common and entirely silent cause of total absence.

When does being indexed turn into being cited?

Being indexed turns into being cited on a timescale set by the engine rather than by the page, and the engines differ by weeks. Our page on how long AEO takes to work covers the measured curves in detail: in Semrush research published in December 2025 tracking 81 newly published test pages, Google AI Mode cited 36% of them the day after publication, while ChatGPT search reached 42% only by day 30.

Freshness then works in your favour for a while and against you later. Roughly half of the content AI engines cite is less than 13 weeks old, according to research Lily Ray presented at Tech SEO Connect in 2025. A Seer Interactive study of 7,683 pages and 47,097 citations, published in 2026, found that 75% of AI-cited pages had been updated within the previous year, so recent publication is rewarded and a page nobody has touched in a year is not.

The decay side is worth planning for instead of being surprised by. According to a 2026 Profound analysis of 240 million citations, the median half-life of a cited source is roughly 4.5 weeks, with 40% to 60% of the domains cited for a given query changing month to month. A page that wins a citation in week three can lose it by week eight without anything on it changing, and our guide on measuring AI citation decay covers how to tell that apart from noise.

What this adds up to for a new page is a sensible checking schedule instead of a daily refresh. Confirm the crawl in the logs within the first two weeks. Start asking the engines in week three, several times rather than once. Judge the page on a rolling two to four week window from there, which is the window the St. Gallen work recommends for reading per-brand trends at all.

What else do people ask about ChatGPT indexing a new page?

Can you submit a URL to ChatGPT to get it indexed faster?

No. There is no submission endpoint, no paid expedite and no setting that forces a recrawl on a given date, and any tool claiming to guarantee one is describing something that does not exist. What you can do is make the page worth returning to: accurate last-modified headers, internal links from pages that are already fetched often, and a real publishing cadence on the domain. Submitting to Bing and IndexNow is cheap and reaches an index several assistants draw on.

Does pasting my URL into ChatGPT help it get indexed?

Pasting a URL triggers a live fetch by ChatGPT-User, which proves the page is reachable at that moment and does not put it in the search index. According to OpenAI's documentation, ChatGPT-User fires only when a user asks ChatGPT to fetch a specific URL, while OAI-SearchBot is the crawler that indexes pages for search and citation. Teams regularly mistake a successful paste for evidence of indexing.

Why does Perplexity find my new page before ChatGPT does?

Because PerplexityBot crawls to serve live retrieval while OpenAI's crawling is split between training collection and a search index that updates on its own schedule. The practical effect is that the same new page often appears in Perplexity days or weeks before ChatGPT search will quote it. A gap in that direction is normal and is not evidence that anything is wrong with the page.

We blocked GPTBot last year. Are we invisible in ChatGPT?

Not necessarily, because blocking GPTBot blocks training collection, not search. OpenAI's documentation states the crawlers are controlled independently, so the question is whether your robots.txt also blocks OAI-SearchBot, which is the one that decides eligibility for ChatGPT search answers. Check the file for both names, and check your firewall and CDN rules too, since blanket bot blocks often live there and not in robots.txt.

How long should we wait before deciding a new page has failed?

At least a month, and judge it on repeated checks rather than one. Confirm the crawl in your logs in the first fortnight, start asking the engines in week three, and read the result over a rolling two to four week window, which is the window the University of St. Gallen study on measuring AI search visibility recommends. In Semrush research published in December 2025 tracking 81 test pages, ChatGPT search only reached 42% cited by day 30, so a verdict at two weeks is premature by design.

Related guides

The Assessment

Find out what the engines can actually see on your site.

A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews, records who is named and from which page, and reports where your own pages are being passed over.

Request a free visibility assessment →
SIGNALS · A BlackSig Systems company