What Sources Does AI Use? How ChatGPT, Gemini and Google AI Overviews Choose Which Sites to Cite

What sources does AI use, and how does it choose them?
AI source selection is the process where an assistant turns your question into search queries, pulls candidate pages from an index and cites the passages that answer it most clearly. In short, AI uses sites it can reach in its index, that other sources corroborate, and that answer the question in a single passage.
Clients often ask me this, usually with a screenshot attached: what sources does AI use, and why did ChatGPT cite a competitor instead of us? This article covers the pipeline behind the citation, not which brands make the list. I looked at brand recommendations in which brands AI search engines recommend, and at practical visibility steps in how your brand shows up in ChatGPT and Gemini.
Below, I walk through the source pipeline for ChatGPT search, the Gemini app, Google AI Overviews, AI Mode, Perplexity and Microsoft Copilot. Then I separate the signals all platforms share, the published overlap research, and the levers you actually control. That way, you can see early which work pays off and which effort goes nowhere.
What is the difference between citations and training data?
A language model answers in one of two ways: from what it learned during training, or from web pages it retrieves at the moment you ask. Citations, however, come from the second path. Google calls this grounding, or retrieval augmented generation (RAG), and says it relies on its core Search ranking systems to retrieve pages from its own index.
That distinction matters because training data is a snapshot of the past. A page you publish today does not enter that snapshot until a new model finishes training. The retrieval pipeline, by contrast, runs again for every question. As a result, a well prepared page can enter the citation list without waiting for the next model release.
You manage the training side with separate controls: GPTBot for OpenAI and the Google-Extended token for Google. I explain how those bots relate to search visibility in my guide to AI crawlers. For the rest of this article, I focus only on retrieval, meaning the source list under the answer.
Where does ChatGPT search find its sources?
ChatGPT search draws on three pools: OpenAI's own crawler, OAI-SearchBot, third-party search providers, and content that partners supply directly. OpenAI named the providers and the partner content when it launched search. Its help page also lists the privacy policies of Microsoft and Shopify under its search providers.
Next, the critical step is query rewriting. According to OpenAI's ChatGPT search help page, ChatGPT rewrites your question into one or more targeted queries and sends them to its providers. After reviewing the first results, it may send narrower follow-up queries. It may also use your approximate location and, if memory is on, saved details when it rewrites the query.
On the access side, one key matters: OAI-SearchBot. OpenAI's publisher FAQ asks you not to block it if you want your pages to appear in summaries and snippets. If OpenAI learns the URL of a blocked page some other way, it may show only the title and the link. Visits from ChatGPT carry the utm_source=chatgpt.com parameter, so you can separate that traffic in analytics.
Which index does the Gemini app rely on?
The Gemini app takes its web information from the Google Search index. Google's crawler documentation says so directly in its definition of Google-Extended: grounding in Gemini Apps means providing content from the Google Search index to the model at prompt time. So you do not need to allow a separate Gemini crawler; a place in Google's index is the baseline.
Still, one detail causes a lot of confusion. If you block the Google-Extended token in robots.txt, your content stays out of training for future Gemini models and out of grounding in Gemini Apps. However, Google states that the token does not affect inclusion or ranking in Google Search. Therefore you keep appearing in AI Overviews and AI Mode, but you reduce your chances of becoming a source in the Gemini app.
In Gemini answers, sources sit behind a sources button under the response or in links within the text. Users can also double-check a response with Google Search. Google's help page points out that a link from that check is not necessarily the source Gemini used to write the answer. So do not treat the double-check screen as proof of source selection.
Why do Google AI Overviews and AI Mode pick different sources from one index?
AI Overviews and AI Mode rely on the same index and the same core ranking systems as classic search. Google's guide to optimizing for generative AI features says these features build on its core Search ranking and quality systems. Specifically, the eligibility rule is simple: Google must have the page in its index, and the page must qualify to show in Search with a snippet.
What separates AI source selection from classic search is query fan-out. Google describes it as issuing multiple related searches across subtopics and data sources. While the model writes the response, it also finds additional supporting pages. That is why Google says it can show a wider and more diverse set of links than a classic web search.
Google also notes that the two features may use different models and techniques, so their links will vary. An Ahrefs analysis from December 2025 measured the gap: for the same queries, only 13.7% of the URLs that AI Mode and AI Overviews cited overlapped. Yet 89.7% of the answer pairs scored above 0.8 for semantic similarity. In other words, both features say similar things but lean on different pages. I cover how to write for the summary itself in how to write content for AI Overviews.
Why does query fan-out make classic rankings insufficient?
Fan-out pushes the candidate pool beyond page one of the main query. The Ahrefs study of AI Overview citations from March 2026 examined 863,000 keyword results and 4 million citations. Only 38% of the pages that AI Overviews cited also ranked in the top 10 for the same query. Another 31.2% ranked between positions 11 and 100, and about 31% sat outside the top 100.
The same team's 2025 study put that share at 76%. Ahrefs notes that Gemini 3 has powered AI Overviews since January 2026 and suggests the new model may expand queries more aggressively or more broadly. In practice, the model now picks pages from side questions, not just from the main query. Ahrefs also reminds readers that AI Overviews are probabilistic, so citations change from query to query.
The practical lesson: ranking first for one head term no longer suffices. You also need to answer the sub-questions around a topic, such as cost, comparisons, steps and risks. On the other hand, Google says you don't need to break content into tiny pieces for AI. The goal is complete pages that cover the sub-questions, not a thin page for every variation.
How does Perplexity build its own index?
Perplexity runs its own index instead of leaning on another search engine's results. According to its Search API announcement, the index covers hundreds of billions of web pages, and the system processes tens of thousands of index update requests every second. PerplexityBot fills that index, and Perplexity states that the bot does not collect content for training foundation models.
Above all, granularity is the most interesting detail. Perplexity says it divides documents into fine-grained units and scores each sub-document unit against the query separately. So the paragraph that carries the answer competes, not the whole page. For that reason, paragraphs that make sense on their own and name their subject explicitly have an edge on Perplexity.
Perplexity-User is a different bot: it opens a page live when someone asks a question, and Perplexity's documentation says it generally ignores robots.txt. In Ahrefs' 2025 data, 28.6% of Perplexity's citations ranked in Google's top 10 for the same query. That was by far the highest share among the assistants in the study.
Why does Microsoft Copilot depend on the Bing index?
Copilot takes its web information from the Bing search service. Microsoft's Copilot documentation lays out the process: Copilot parses your prompt, picks the terms where web data would help, and sends Bing a search query of a few words. Copilot Chat also shows those queries in the citation section of the answer, so you can see which search surfaced which site.
The biggest change on the publisher side is the AI Performance report in Bing Webmaster Tools. In its February 2026 announcement, Microsoft said the report shows citations across Copilot, AI summaries in Bing and select partner experiences. It includes total citations, cited pages and grounding queries, meaning the key phrases the AI used when it retrieved content. In June 2026, Microsoft added intents, topics, citation share and period comparison as preview features.
The same announcement includes concrete advice: strengthen depth and expertise, improve structure with clear headings, tables and FAQ sections, and support claims with evidence. Microsoft also recommends IndexNow for freshness and up-to-date Bing Places listings for local businesses. Because Microsoft appears among OpenAI's search providers, your index health on Bing may matter for ChatGPT too.
How do the five source pipelines compare side by side?
The table below summarizes the pipelines from official documentation. Pay special attention to the access column, because that is where most problems start.
| System | Where candidates come from | Access | Query method | Official data |
|---|---|---|---|---|
| ChatGPT search | OAI-SearchBot crawl, third-party search providers, partner content | OAI-SearchBot | Rewrites the question into one or more targeted queries | utm_source=chatgpt.com in analytics |
| Gemini app | Google Search index | Googlebot and Google-Extended | Grounds answers in the Google index at prompt time | Outside the Search Console report |
| AI Overviews and AI Mode | Google Search index and core ranking systems | Googlebot and snippet eligibility | Query fan-out across many related searches | Search Console generative AI report |
| Perplexity | Own index, hundreds of billions of pages | PerplexityBot | Scores sub-document units against the query | Referral traffic in analytics |
| Microsoft Copilot | Bing index | Bingbot, IndexNow for notification | Generates a short Bing query | Bing Webmaster Tools AI Performance |
The most practical lesson is that every system has its own door. A single line in robots.txt can cut you off from one system while leaving the others open. Instead of writing rules by hand, I recommend defining each bot separately with a robots.txt generator. After a change, be patient: OpenAI and Perplexity both say their systems can take up to 24 hours to reflect new rules.
What sources does AI use on each platform, and why do they differ?
Short answer: each platform gathers candidates from a different index, with different queries and different merging rules. The Ahrefs overlap study from August 2025 ran 15,000 long-tail queries through search engines and assistants. Only 12% of the links that ChatGPT, Gemini and Copilot cited appeared in Google's top 10 for the same prompt.
In the same study, about 80% of citations did not rank anywhere in Google for the original query. Overlap with Bing's top 10 averaged roughly 10%, with Copilot highest at 14%. Ahrefs explains this mainly through fan-out: assistants generate many variations of a query, then blend the results with methods that merge several rankings. These are the main factors that widen the gap.
- Different indexes: Google, Bing and Perplexity run separate indexes, while ChatGPT combines its own crawl with provider results.
- Different queries: each system splits the same question into different sub-queries.
- Personal context: location, language and memory can change the rewritten query.
- Probabilistic output: the same question can return a different source list in a different session.
- Answer length: according to Ahrefs, AI Mode answers run about four times longer than AI Overviews and mention more entities, so they leave room for more sources.
What sources does AI use most: Wikipedia, Reddit or YouTube?
Each platform also has its own taste in sources. A Profound analysis of 680 million citations, published in June 2025, found that Wikipedia made up 47.9% of ChatGPT's top 10 most cited sources. On Perplexity, Reddit led the same list with 46.7%. AI Overviews spread more evenly: Reddit 21%, YouTube 18.8%, Quora 14.3% and LinkedIn 13%.
Moreover, these shares do not stay put. In Ahrefs' March 2026 data, YouTube became the most cited domain in AI Overviews. Among cited pages that did not rank in the top 100 for the same keyword, 18.2% were YouTube URLs. In the brand recommendations article, I used Semrush data to show how sharply the Reddit and Wikipedia shares fell in ChatGPT within a few weeks.
My takeaway: instead of copying the format one platform likes, build a solid presence in every format. Video explanations tend to land in AI Overviews, encyclopedic consistency in ChatGPT, and community discussion in Perplexity. However, these preferences can shift within months, so betting everything on one channel is risky.
Which authority signals do all platforms share?
Despite different indexes and different queries, some signals work on every platform. What they have in common is that they make a claim verifiable through sources independent of you. These are the six signals I run into most often in client work.
- Third-party mentions: trade publications, press coverage and expert lists describe your brand for you.
- Review and comparison sites: in Profound's data, G2 held a 6.7% share of ChatGPT's top 10 most cited sources.
- Wikipedia and Wikidata: your name, founding details and official website stay consistent in one place.
- Forums and communities: real user experience adds evidence that vendor copy cannot.
- Freshness: assistants tend to cite newer pages than organic results do.
- Clear passage-level answers: a page that answers the question in one paragraph stands out during retrieval.
One warning belongs here. Google's guide states plainly that chasing inauthentic mentions across the web is not as helpful as it might seem, and that other systems block spam. So the goal is to earn the signal, not to manufacture it. The next sections look at each signal through that lens.
Why do third-party mentions and review sites carry so much weight?
When an assistant answers a “which one is best” question, it cannot treat a brand's own website as a neutral source. That is why claims that independent sites confirm tend to rise to the top of answers. It is also the pattern I see most often in the field: brands with strong websites but no mentions anywhere else fall behind on category questions.
Official guidance, likewise, points the same way. Google writes that Merchant Center feeds and Google Business Profiles can help your products and services show up in AI responses and other Search results. Microsoft likewise recommends keeping Bing Places listings current for local businesses. I cover the local side in detail in my Google Maps SEO guide.
With reviews, the only rule is authenticity. Reviews you buy or incentivize break platform policies and, once discovered, damage trust for a long time. Instead, build a process that asks satisfied customers for a review after delivery. In B2B, keep your profiles on the review and comparison sites of your industry complete and current.
How does being on Wikipedia and Wikidata affect source selection?
Wikipedia ranks among the most cited sources, especially in ChatGPT. However, getting in is not your decision. Wikipedia's notability guideline says a topic qualifies for an article only when reliable sources independent of the subject cover it in depth. The guideline does not count press releases, advertising or the company's own website as independent.
The conflict of interest guideline also strongly discourages you from editing an article about your own company directly and asks you to propose changes on the talk page instead. Wikidata sets a lower bar: you can create an item for a clearly identifiable entity that serious, publicly available references describe.
My advice is not to reverse the order. First, collect genuine mentions in independent publications. Then describe your name, logo and social profiles consistently on your site with Organization markup and the sameAs property, which you can build in minutes with a schema generator. A Wikipedia article arrives on its own once you earn it; if you force one, you risk deletion.
Why do forums and community content show up as sources so often?
Forums offer something vendor copy cannot: user experience in the user's own words. A forum thread often tells you how a product behaves after six months, how the support team responds, or where setup gets stuck. That is why AI answers often turn to community pages, especially for comparison and troubleshooting questions.
The Ahrefs comparison of AI Mode and AI Overviews also shows that this pattern varies by feature. Quora receives 3.5 times more citations in AI Mode than in AI Overviews, while Reddit stays at a similar level in both. As a result, the payoff of forum visibility changes from platform to platform.
So what can a brand do? Answering questions openly under your own name works; posting praise from fake accounts does not. Helpful expert answers also attract independent users who mention you over time. I outline the broader trust framework in my E-E-A-T guide, and your behavior in forums is a direct extension of it.
How much does freshness affect source selection?
Freshness matters, but less mechanically than many people assume. In July 2025, Ahrefs examined close to 17 million citations. Pages that AI assistants cited averaged 1,064 days old, while pages in organic results averaged 1,432 days. That means assistants chose content about 25.7% fresher. ChatGPT cited the newest pages, while AI Overviews leaned toward older pages, much like organic results.
Ahrefs adds an important warning in the same article: updating the publish date without changing the content does not help. Refreshing a low quality page every day will not have a magic effect either. The average citation age of roughly three years also shows that long-lived content still carries value.
In practice, I update content at the level of facts. When a price, date, version, regulation or statistic changes, I revise the page and note the change in the text. Then I notify Bing through IndexNow; on the Google side, a current sitemap and internal links do the job. I describe this routine step by step in my content freshness guide.
Why do clear passage-level answers decide so much?
Retrieval systems may read the whole page, but what they carry into the answer is usually a passage. Perplexity states this openly and scores sub-document units against the query one by one. Microsoft also recommends clear headings, tables, FAQ sections and claims backed by evidence. Google, meanwhile, stresses that you don't need to chop content into small pieces.
These three positions do not conflict. Instead of splitting the page, write every section so it makes sense on its own. This is the checklist I use.
- The first sentence under each H2 answers the question directly.
- Each paragraph names its subject; write the method's name instead of “this method”.
- When you give a number, add its source and date in the same sentence.
- Move comparisons into tables and steps into numbered lists.
- Every section makes sense without relying on another section.
This structure also matches the logic of featured snippets. A paragraph you craft well for classic search also becomes ready for quotation in an AI answer.
Which parts of source selection can you influence?
The area you can influence is wider than people assume, although it consists of fundamentals. When my team reviews a site's potential as an AI source, we start with these six areas.
- Access: can OAI-SearchBot, PerplexityBot, Googlebot and Bingbot reach your pages? You can run a first scan with our AI visibility checker.
- Index: are your pages in both Google's and Bing's index? Search Console and Bing Webmaster Tools give the only reliable answer.
- Entity consistency: do your brand name, address and description match across your site, Business Profile, Bing Places, LinkedIn and directories?
- Sub-question coverage: do you answer the side questions that fan-out opens up around your topic?
- Evidence: does each claim rest on a dated source, original data or a real example?
- Independent mentions: do trade articles, expert quotes and genuine customer reviews keep building up?
None of this requires a new trick. Google's guide also says you don't need llms.txt style files, special schema markup or Markdown versions to appear in Google Search. That is why, in our AI SEO and GEO services, we first strengthen the foundation and only then measure visibility.
Which factors are completely outside your control?
Knowing what you cannot control is the fastest way to avoid wasted effort. No matter how hard you work, you have no direct lever over the following factors.
- Which search provider ChatGPT consults and how it rewrites the question.
- The user's location, language and saved memory.
- Probabilistic output, meaning the same question can surface different sources on each attempt.
- What the model's past training data says about your brand.
- OpenAI's partnerships with news and data providers.
- Wikipedia editors' notability decisions.
- Whether an AI summary appears for a query at all.
So don't change strategy based on a single screenshot. Repeat the same questions on different days and across platforms, and watch the trend. If the trend turns against you, go back to what you can control, the six areas in the previous section. More often than not, the cause sits there.
Which official reports can you use to track source selection?
Three official data sources exist, and each measures something different. Google added a generative AI performance report to Search Console in June 2026 and opened it to all sites by the end of August 2026. The report shows impressions in AI Overviews and AI Mode by page, country, device and date, but it does not offer a query breakdown.
Meanwhile, the AI Performance report in Bing Webmaster Tools shows citation counts and grounding queries. The third source is your own analytics: you can isolate visits from ChatGPT through utm_source=chatgpt.com, and other assistants through their referring domains. I explain the setup in my guide to detecting AI traffic in GA4.
Measurement protocol goes beyond this article, but one principle deserves emphasis. Put simply, impressions, citations and clicks are different metrics. A page that earns many impressions in AI Overviews but few clicks is normal. In that case, read which question the page answers and strengthen the next step with internal links.
Where should you start in the first 30 days?
If you're starting from scratch, split the work into four weeks. The plan below condenses the order I follow in client projects.
- First week, access: verify crawler access in robots.txt, your CDN and firewall settings; open a Bing Webmaster Tools account and turn on IndexNow.
- Second week, questions: list the ten most important questions for your business and note, platform by platform, which sites each answer cites.
- In the third week, rewrite your five strongest pages at passage level; lead with the answer, source every number and move comparisons into tables.
- Finally, in week four, align your brand details across all profiles and plan mentions in the independent publications you saw in the source lists.
After four weeks, you will know where you stand and which source type is missing. If you prefer, my team can run this process with you; we handle SEO and AI visibility in one plan. One last reminder: the answer to what sources does AI use keeps changing, but sites that stay accessible, verifiable and clear remain on the list.




