Neither answer is clean. OpenAI documents its own crawler for this: OpenAI's crawler documentation describes OAI-SearchBot as the bot used to surface websites in ChatGPT search results, and says sites opted out of it will not be shown in ChatGPT search answers. Bing has supplied index results to ChatGPT, and how much of the current consumer retrieval path still runs through Bing is contested in published measurement rather than settled. The measured overlap points in both directions at once: Seer Interactive published an analysis titled "87% of SearchGPT Citations Match Bing's Top Results", while Ahrefs' 2026 overlap study of 15,000 queries found only around 11% of AI assistant citations appeared in Google and Bing's top 10 for the prompt asked. The practical consequence is that there is no single index to optimise for, and the work that survives either answer is being readable, being in the major indexes, and being written in the words of the prompt.
ChatGPT search runs through OpenAI's own crawler and search surface, with Bing having served as an index supplier, and the current balance between the two is not something either company has published in detail. Anyone telling you flatly that ChatGPT is Bing, or that it no longer touches Bing at all, is simplifying past the evidence.
What is documented is the crawler. OpenAI's crawler documentation sets out OAI-SearchBot as the agent that surfaces sites in ChatGPT search results, separately from GPTBot, which is used for model training, and separately again from ChatGPT-User, which fetches pages in response to something a user asked for. Each of the three has its own robots.txt token, so a site can allow one and refuse another.
What is contested is how much of the retrieval path is still Bing's. Industry measurements through 2026 disagree, with some reporting Bing as the primary real-time source and others reporting it as largely absent from the consumer path while OpenAI builds its own index. We have not found an official OpenAI statement resolving it, which is the honest position to take on a question this commercially loaded.
Google's position is cleaner because Google owns both ends. AI Overviews and AI Mode run on the standard Googlebot user agent, so they draw on the same crawl and index as classic search, and Google-Extended is a robots.txt token about training and Gemini app grounding rather than a crawler with an index of its own.
The reason this matters less than it feels like it should is that index membership is table stakes rather than strategy. Being in an index is what makes you eligible. What decides whether you are quoted happens after retrieval, on the page, which is the subject of how AI engines choose which sources to cite.
ChatGPT's citations overlap heavily with Bing's results on one published measurement and barely at all with Google's on another, and the gap between those two findings is the most informative thing in this article. Both can be true because they measure different things.
On the Bing side, according to Seer Interactive's published analysis, titled "87% of SearchGPT Citations Match Bing's Top Results", the great majority of what ChatGPT cited could be found in Bing's results. We could not open the full study to confirm its sample or date, so treat the headline as directional.
On the Google side, Ahrefs' 2026 overlap study, run by data scientist Xibeijia Guan across 15,000 long-tail queries, found about 12% of URLs cited by ChatGPT, Gemini and Copilot ranked in Google's top 10 for the same prompt, and around 11% when Google and Bing's top 10 were taken together. Perplexity was the exception at closer to one in three.
The reconciliation is the prompt. Ahrefs compared citations against the results for the question as the user asked it, while an engine searches its own rewritten sub-queries, so a page can be absent from the top 10 for the prompt and sitting in the top 10 for the search the engine actually ran. A separate Ahrefs analysis of 118,931 fan-out queries reported that 83.39% of ChatGPT's chosen results did not appear in Google's results for the same fan-out queries, which is a stronger claim than the first and in the same direction.
What a marketer should take from the disagreement is to stop treating any single overlap figure as a strategy. Index presence in both Google and Bing is cheap and worth having. Beyond that, what moves the result is on the pages rather than in guessing which index is being read this quarter.
ChatGPT sends its own rewritten queries to the index rather than the sentence you typed, which is why your keyword rankings and your citation results can disagree so sharply. One prompt becomes several searches, and the wording of those searches is chosen by the model.
The scale of that rewriting is measurable. AirOps' 2026 retrieval study found 89.6% of the 15,000 prompts it tested produced two or more internal sub-queries, and that 32.9% of cited pages appeared only in the results for those sub-queries rather than for the original prompt. A third of the pages that won were invisible to anybody tracking the question as asked.
The same study shows how much gets discarded after retrieval. Of 548,534 pages ChatGPT retrieved, only about 15% became a visible citation, so roughly 85% of what the engine read never reached the reader. Being in the index, and even being retrieved, is a long way from being quoted.
Our own logs show the input side of this from the other direction. The Bing Webmaster query log for this site is dominated by long conversational prompts rather than keyword phrases, including questions of twenty words and more about retrieval pipelines and citation behaviour. Those are people typing prompts, and the pages that match them are written as answers to questions rather than as keyword targets.
The writing implication is specific: cover the questions around the main one as well as the main one itself. A page that answers the question around the question appears in sub-query results its competitors miss, which is the mechanism query fan-out sets out in full.
Each AI engine documents its own crawler, and the index behind it is sometimes published and sometimes not. The table below separates what the vendors document from what is inferred, because the second category gets quoted as fact far too often.
| Engine | Documented crawler for search surfacing | Index behind the answers |
|---|---|---|
| ChatGPT search | OAI-SearchBot, documented by OpenAI as the bot that surfaces sites in ChatGPT search results, with GPTBot for training and ChatGPT-User for user-initiated fetches | OpenAI's own crawl plus index results that have come from Bing; the current split is not published |
| Google AI Overviews and AI Mode | Standard Googlebot, per Google's crawler documentation | Google's own index |
| Microsoft Copilot | Bingbot | Microsoft's Bing index |
| Perplexity | PerplexityBot for search surfacing, with Perplexity-User for in-platform requests | Its own continuous crawl, supplemented by third-party results |
| Claude with search | ClaudeBot, with a separate user agent for user-initiated fetches | A third-party search provider rather than a published Anthropic index |
Two practical rules follow from the table. Allow the search-surfacing crawlers even if you block the training ones, because they are separate tokens and blocking the wrong one removes you from answers rather than from training sets. And check your robots.txt against the current documentation rather than against a blog post, because these tokens and their descriptions have changed more than once. Our guide on whether you should block AI crawlers works through the trade-off.
The column that moves fastest is the third one. Index arrangements between these companies are commercial deals, and they change without announcement, which is a reason to build for retrievability in general rather than for one supplier's pipeline.
Being indexed in Bing is close to a precondition and nowhere near sufficient, which is the same relationship index membership has with every other engine here. It gets you into the pool. It does not get you out of it.
The arithmetic makes the point. AirOps' 2026 study found about 15% of the 548,534 pages ChatGPT retrieved became a visible citation, which means most of the index membership in the world still ends in being read and dropped. Whatever gets you selected happens after the search returns.
Where index presence does pay off is speed and eligibility for new pages. A page nobody has crawled cannot be retrieved by anything, so submitting your sitemap in Bing Webmaster Tools and wiring up IndexNow removes a delay you control, which is the subject of how long it takes for ChatGPT to index a new page.
What index presence cannot fix is a page that is not written as an answer. If the heading is a product name, the first sentence is context and the claims carry no sources, being perfectly indexed in three search engines changes nothing about whether a sentence gets lifted. Our AI citation checklist is the page-level version of this work.
A reasonable order of operations for a team starting from nothing: confirm the crawlers can reach the site, confirm the major indexes have the pages, then spend the remaining effort entirely on the pages themselves. The first two take an afternoon each. The third is the actual programme and it does not end.
Checking whether Bing has your pages takes two minutes with a site query and ten with the webmaster tools, and both are worth doing because they answer slightly different questions. Start by searching your domain with a site query in Bing and comparing the count against what you know you have published.
Then set up Bing Webmaster Tools, which gives you per-URL inspection, crawl errors and the query log for your own site. That log is the useful part and the part most teams never look at: it shows the actual phrasings people used to reach you in the Microsoft index, which is the vocabulary evidence that matters for the writing work.
Submit a sitemap and enable IndexNow while you are there. IndexNow pushes new and changed URLs to the Microsoft index rather than waiting for a crawl, which shortens the one part of the discovery delay you control directly.
Check the crawler side separately from the index side, because they fail independently. Fetch a few important pages with JavaScript disabled to see what a crawler receives, and read your robots.txt line by line against the current vendor documentation for each agent. Our walkthrough on how to check whether AI can read your website covers both tests in order.
Treat the whole check as hygiene with a date on it. Index coverage degrades quietly after a migration or a template change, so the right cadence is a quick look every month and a full pass after any structural release, which is also when a redesign can cost you visibility without anybody noticing.
Optimising specifically for Bing is the wrong conclusion to draw, and it is the one most commonly drawn from the overlap research. Bing index membership is a yes or no checkbox, and once it is yes there is very little further Bing-specific work with any leverage behind it.
The reason is that the ranking system and the citation system reward different things. A page can rank in an index and still be dropped at the writing stage, and the signals that decide the second step are about vocabulary, structure and sourcing on the page. Discovered Labs' analysis of 2 million AI citations found prompt-content alignment carried a standardised effect of +0.37, roughly three times the next strongest page-level signal, as reported in their citation research.
There is also a stability argument. Index supply deals change, engines build their own crawlers, and a programme designed around one pipeline inherits every change to it. Work that holds across all of them is a short list: be crawlable without JavaScript, be in the major indexes, write headings as the questions buyers ask, answer them in the first sentence, and source every number in its own paragraph.
Where engine-specific work does earn its place is in measurement rather than in optimisation. Running the same question set against each engine separately shows which ones already name you and which do not, and that comparison is what tells you where attention is worth spending, as which AI engine to optimise for first sets out.
If you want to see which engines currently name your company and from which pages, a free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews and reports the answer per engine rather than as a single score.
ChatGPT has drawn on Bing results and now also runs its own search crawler, and how much of the current retrieval path still goes through Bing is contested rather than settled. OpenAI documents OAI-SearchBot as the crawler that surfaces websites in ChatGPT search results, and independent measurements of the overlap between ChatGPT citations and Bing's top results disagree with each other. Treat any flat claim in either direction as a vendor simplification.
Blocking GPTBot and blocking ChatGPT search are two different decisions with two different robots.txt tokens. OpenAI's crawler documentation describes OAI-SearchBot as the bot used to surface websites in ChatGPT search results, and says sites opted out of it will not be shown in ChatGPT search answers. A company that wants to stay out of model training while remaining visible in answers can block one and allow the other.
Copilot is Microsoft's own product and Bing index membership is the baseline requirement for it, which makes Bing Webmaster Tools worth the setup time for two engines rather than one. The work beyond index membership is the same work as everywhere else: a page written in the buyer's vocabulary with answers that can be lifted out of it. Our guide to getting cited by Microsoft Copilot covers what differs.
Google-Extended controls whether content Google has already crawled may be used to train Gemini models and to ground answers in the Gemini apps, and Google documents that it is not a crawler and has no user agent of its own. Blocking it does not remove a site from Google Search or from AI Overviews, which run on the standard Googlebot user agent. The two decisions are separate and should be made separately.
IndexNow is worth wiring up because it is cheap and it addresses the one part of this you control directly, which is how quickly a new page becomes discoverable in the Microsoft index. It does not make a page citable and it does not shorten the time before an engine chooses to quote you. Treat it as publishing hygiene alongside a clean sitemap rather than as an AI visibility tactic.
A free visibility assessment runs your buyer questions across ChatGPT, Claude, Perplexity and Google AI Overviews and reports who is named and from which page, per engine, so the differences between them are visible rather than averaged away.
Request a free visibility assessment →