Artificial Intelligence

Why Is Your Website Not Cited by AI? 12 Common Mistakes and a Diagnostic Checklist

Talha Aslan 19 min read

Why is your website not cited by AI?

A website that is not cited by AI usually has a break in one of six layers: bots cannot reach the page, the page is missing from Google or Bing, the text depends on JavaScript, snippet controls restrict it, the content buries the answer, or no independent source confirms the brand. Diagnose the layers in that order.

I have worked on SEO and paid search projects since 2012. Over the last two years, one question has come up more than any other: "ChatGPT recommends our competitor, so why does it never mention us?" This guide turns that question into a diagnostic checklist. For each mistake you get the symptom, the check and the fix, and a summary table closes the article.

First, an honest caveat: no method guarantees a citation. Even Google says AI Overviews has no extra technical requirements beyond being in the index and eligible for a snippet. So the goal is not a secret setting. Instead, you are looking for the weak link in the chain, and it usually sits somewhere more technical and more ordinary than you expect.

How do you confirm the problem before diagnosing it?

Before you conclude that your site is not cited by AI, prove the symptom is real. A single chat session is misleading, because answers vary by account, location and phrasing. So collect evidence from several sources:

  • A prompt set: Write down 15 to 20 questions your customers actually ask. Run them with the same wording in ChatGPT, Gemini, Perplexity and Copilot, then log the cited sources with dates.
  • Bing data: In February 2026, Bing Webmaster Tools opened its AI Performance report as a public preview. It shows citation counts across Copilot, AI summaries in Bing and select partner experiences, plus your cited pages and the grounding queries behind them.
  • Google data: Google counts sites that appear in AI Overviews and AI Mode inside the Web search type of the Search Console Performance report. So look for page-level impression shifts rather than a separate AI number.
  • A quick scan: Our AI visibility checker gives you a first look at how your brand appears in AI answers.

To measure the visits side, use my guide to detecting AI traffic in GA4. Once the symptom is clear, move on to the decision tree.

Not cited by AI: in which order should you check?

Order matters, because a blocker in a lower layer wipes out every effort above it. If a bot cannot open your page, there is no point debating content quality. So follow this decision tree from top to bottom and stop at the first "no".

  1. Access: Do search and user bots receive a 200 status code? If not, review robots.txt and your CDN rules (mistakes 1 and 2).
  2. Index: Does the page appear in the Google and Bing indexes? If not, fix indexing first (mistake 3).
  3. Rendering: Does the raw HTML contain the main text? If not, move to server-side rendering (mistake 4).
  4. Permission: Do snippet and indexing rules leave the content open? If not, correct the tags (mistake 5).
  5. Content: Does the page answer the question early, accurately and originally? If not, strengthen structure and substance (mistakes 6, 9, 10 and 11).
  6. Trust: Does the brand have a clear definition, and do other sites confirm it? If not, work on entity and reputation signals (mistakes 7 and 8).
  7. Myth check: Are you hoping a shortcut like llms.txt will fix everything? Finish the first six steps before you touch it (mistake 12).

In my experience, the technical layers are the fastest to fix. Content and trust, on the other hand, take weeks, sometimes months. So close out the first four steps within a few days and give the remaining time to content.

Mistake 1: Is robots.txt blocking the search bots?

Symptom: You rank well on Google, yet your name never shows up in ChatGPT, Claude or Perplexity answers. The usual cause is a broad rule that meant to stop training bots and ended up stopping search bots too.

Check: Open yourdomain.com/robots.txt in a browser and read every User-agent group. Then note the status of these bots:

  • OAI-SearchBot: According to OpenAI, it surfaces websites in ChatGPT's search features. Sites that opt out of it will not appear in ChatGPT search answers, although they can still show up as navigational links.
  • GPTBot: It crawls for model training. Blocking it tells OpenAI you do not want your content in training data; it does not decide your search visibility.
  • Claude-SearchBot and PerplexityBot: These feed the search side of Claude and Perplexity. Anthropic also runs ClaudeBot for training and Claude-User for user requests.
  • Googlebot and Bingbot: They build the classic indexes that AI Overviews, AI Mode and Copilot rely on. Google-Extended, by contrast, does not affect your visibility in Google Search.

Fix: Write a file that explicitly allows search bots and makes a separate decision about training bots. OpenAI's crawler documentation says robots.txt changes can take about 24 hours to reach its systems. I cover bot names and sample files in my AI crawler guide. The robots.txt generator also helps you build a clean file.

Mistake 2: Is your CDN or firewall silently turning bots away?

Symptom: Your robots.txt allows everything, but server logs show search bots receiving a 403 response, or no requests at all. In that case the blocker is not the file; it is the infrastructure in front of your site.

On July 1, 2025, Cloudflare announced it had become the first internet infrastructure provider to block AI crawlers without permission by default. The same release said more than one million customers had switched on its one-click block since September 2024. The current Cloudflare documentation splits bots into Search, Agent and Training classes. Its AI bot blocking targets Training and Agent, while Search stays open. Since September 15, 2026, the default for new domains blocks Training and Agent bots on pages that display ads.

Two details matter for diagnosis. First, the Training class also covers mixed-purpose crawlers that serve both training and search. Second, the Agent class includes bots such as ChatGPT-User that open pages in real time on a user's behalf. If you block that class, a user who asks ChatGPT to open your page gets nothing back.

Check and fix: On Cloudflare, review the action for each bot in the crawler table of AI Crawl Control. With any other WAF, inspect rate limits, country blocks and bot verification rules. Allow the search bots, and separate fake bots from real ones with the IP lists the providers publish.

Mistake 3: Is the page in the Google and Bing indexes?

Symptom: You published the page recently or moved it to a new URL, and it appears neither in classic search nor in AI answers. As a result, a page outside the index has practically no chance of becoming a source.

Google's AI features documentation is clear here. To appear as a supporting link in AI Overviews or AI Mode, a page needs to sit in the index and qualify for a snippet. The document adds that there are no additional technical requirements.

ChatGPT is murkier. OpenAI's help center says ChatGPT search draws on third-party search providers as well as content that partners supply directly. Back in May 2023, Microsoft announced Bing as the default search experience in ChatGPT. OpenAI does not list its providers today, so I treat the Bing index as a strong prerequisite rather than a hard rule. Copilot, on the other hand, runs directly on Bing.

Check and fix: Use URL Inspection in Google Search Console and the URL Inspection tool in Bing Webmaster Tools. Next, submit a current XML sitemap to both. Then set up IndexNow to notify Bing of new and updated pages instantly; Microsoft recommends it as the way to keep AI experiences fresh.

Mistake 4: Does your content only load with JavaScript?

Symptom: The page works fine in a browser and sits in Google's index, yet ChatGPT and Claude never quote it. I see this pattern a lot on single-page apps, tabbed product descriptions and FAQ blocks that expand on click.

In December 2024, Vercel and MERJ published an analysis of AI crawler traffic. It found that none of the major AI crawlers rendered JavaScript. The exceptions were Gemini, which uses Google's infrastructure, and AppleBot, which crawls with a browser-based renderer. The same study also counted 569 million GPTBot requests and 370 million Claude requests across Vercel's network in one month.

That snapshot dates from late 2024, and providers can change their capabilities. Still, the safe assumption today is simple: if your main text is not in the raw HTML, these bots cannot read it. Run three quick checks:

  1. Open the page source and search for a sentence from your first paragraph.
  2. Then disable JavaScript in your browser and reload the page.
  3. Finally, check server logs to see whether bots fetch only the HTML document or the JS files as well.

Fix: Put the main content, pricing tables and FAQ answers into the HTML through server-side rendering, static generation or prerendering. Good JavaScript optimization helps both speed and readability at the same time.

Mistake 5: Are nosnippet, max-snippet or noindex limiting the answer?

Symptom: The page sits in the index and bots can reach it, yet your content never appears in AI Overviews, or only your title shows up. In cases like this, page-level directives are my first suspect.

Google's robots meta tag specification is explicit. The nosnippet rule prevents content from serving as a direct input for AI Overviews and AI Mode. In addition, max-snippet limits how much text those features may use. Sections you mark with data-nosnippet stay out of snippets, while noindex removes the page from results entirely.

Bing has its own trap. According to Bing's September 2023 announcement, content with a noarchive tag stays out of Bing Chat answers, which today means Copilot, and receives no link there either. Nocache, on the other hand, still allows the URL, title and snippet in the answer. Yet many sites added noarchive years ago only to hide cached copies.

Check and fix: Read the robots meta tag in the page source and the X-Robots-Tag header in the server response. Google also reminds you that when robots.txt blocks crawling, Googlebot never sees these rules at all. Remove restrictions you added by accident; for sensitive passages, mark only that section with data-nosnippet instead of the whole page.

Mistake 6: Where does the answer sit on the page?

Symptom: Bots can read the page and it ranks, yet the AI answer quotes someone else's sentence. When you open your page, the answer appears in the fourth paragraph after a long introduction.

AI search systems do not pick a whole page; they pick the passage that best fits the question. If a clear definition answers the question right under the heading, that passage is easy to quote. Microsoft makes the same point in its AI Performance announcement: clear headings, tables and FAQ sections make content easier for AI systems to reference accurately.

For example, picture a service page that answers "How long does the project take?" One page gives the timeline and the three factors behind it in the first two sentences. Another starts with company history and then its values. For an answer engine, quoting the first page is both easier and less risky.

Check: Read only the headings and the first sentence under each one. If that skeleton does not answer the question, a model will struggle too.

Fix: Use question-style headings, write a 40 to 60 word direct answer beneath them, and add detail afterwards. Also move comparisons into tables and steps into numbered lists. I explain this structure with examples in my AEO guide.

Mistake 7: Is AI confusing your brand with something else?

Symptom: The model mentions your brand but mixes it up with the wrong city, an old product or another company with the same name. Sometimes it does not recognize you at all and instead falls back on a generic category answer.

This is an entity ambiguity problem. If your brand name is a common word, or it appears in different forms on your site, social profiles and directories, the system cannot tell which record belongs to you. Faced with an uncertain entity, a model will naturally pick a competitor it can identify. Personal brands run into the same issue through name overlap, so consistently pairing your name with your field, city and company helps the model pick the right person.

Check: Ask ChatGPT and Gemini, "What is [brand], what does it do and where does it operate?" Then compare the answer with your About page, your Google Business Profile and your social profiles.

Fix: Use the same name, address and category description everywhere. In your Organization schema, connect your official profiles through the logo, contactPoint and sameAs properties; my schema markup guide explains the logic. Local businesses should also keep Bing Places for Business up to date, which Microsoft itself recommends. Finally, state founder, team and service facts in plain sentences on a single About page.

Mistake 8: Does anyone besides you vouch for you?

Symptom: Your site is technically clean and the content is strong, yet recommendation prompts keep returning the same three competitors. What those competitors usually share is that people talk about them outside their own websites.

My observation is that AI systems lean more on information that independent sources repeat than on what a site says about itself. Industry publications, comparison lists, genuine customer reviews, professional associations and forum discussions all supply that confirmation. I cover the mechanism in detail in how your brand shows up in ChatGPT and Gemini.

Check: Search Google for your brand name while excluding your own domain. Count the independent sources that mention you on the first two pages. Then run the same search for the competitor that AI keeps citing and note the gap.

Fix: Focus on durable work such as digital PR, data contributions to industry reports, genuine reviews and expert commentary. Fake reviews or paid listicles may work briefly, but they damage trust for the long term. Also remember that unlinked brand mentions still add context, so make sure press and industry pages describe what you do accurately in one sentence.

Mistake 9: Has your content gone stale?

Symptom: AI answers cover the topic with current data, while your page still shows prices from two years ago, old product names or menu paths that no longer exist. Citing such a page would force the model to risk passing on wrong information.

Microsoft's AI Performance announcement is explicit on this point: regular updates help AI systems reference the most current version of your content. Old names are also a separate issue. For example, a page that still says Bard or Bing Chat makes a weak candidate for a question about Gemini or Copilot.

Check: Read the visible date, the dateModified value in your structured data and the numbers in the text together. A fresh date on old information means the problem is still there.

Fix: Update the information rather than the date: new figures, new screenshots, new examples. After each update, notify Bing through IndexNow and request recrawling in Search Console. Then set a quarterly review calendar for your critical pages.

Mistake 10: Does the page add anything new?

Symptom: The page sits in the index and its structure is fine, but it repeats what your competitors already say in different words. Therefore, an answer engine has no particular reason to choose it.

A thin page is not just a short page. A 2,000-word article is thin too if it carries no original data, experience or examples. The sentence a model quotes should hold something it cannot find elsewhere: your own measurement, a clear price range, process steps, a comparison table or an expert view.

Check: Read the page and ask yourself, "Which sentence here could only we have written?" If there is no answer, the page is thin. In the content audits my team and I run, this original contribution is the first thing we look for. Without it, adding words does not solve anything; it only dilutes the page further.

Fix: Merge similar weak pages into one strong page. Make author details, real experience and sources visible; my E-E-A-T guide explains why these signals matter.

Mistake 11: Is the content behind a login or a paywall?

Symptom: Your most valuable content sits behind a membership, payment or form step. Logged-in users see everything, but a logged-out visitor sees only the headline and a short teaser.

Crawlers do not log in and do not fill out forms. Anthropic states plainly that its bots do not try to bypass CAPTCHAs. So text a bot cannot see cannot become a source in an answer either. The same applies to overlays that cover the whole page and never put the text into the HTML.

The problem also goes beyond paywalls. Geo-redirects, age gates and templates that load text only after cookie consent produce the same result. That is why you should check what the bot sees in the first HTML response before debating content quality.

Check: Open the page in a private window without logging in, then confirm that the main text appears in the page source.

Fix: Write a public summary for every locked piece that makes sense on its own. If you sell premium content, use the structured data Google recommends for paywalled content, so Google can tell a paywall apart from cloaking.

Mistake 12: Are you waiting for llms.txt to fix it?

Symptom: The team blames the missing llms.txt file for the lack of citations and keeps postponing the real issues. Over the past year, this has been the most common misdiagnosis I have seen.

Google's official documentation says you do not need new machine-readable files, AI text files or special markup to appear in AI Overviews or AI Mode. In press-reported comments, Google's John Mueller compared the file to the old keywords meta tag. Gary Illyes also said Google does not support it and has no plans to.

Check: Has llms.txt jumped ahead of one of the first eleven mistakes on your roadmap? If so, reorder the list.

Fix: You can keep the file as a harmless extra, and our llms.txt generator builds one in minutes. Just do not expect more citations from it. Put your effort into access, indexing, content and trust instead.

That said, the file can have a narrow use. For instance, on multi-page developer documentation it helps users feed context into their own AI tools. However, that benefit has little to do with earning citations in search answers.

Not cited by AI: the 12 mistakes at a glance

Use the table below as a shared checklist with your team.

MistakeSymptomCheckFix
robots.txt blockVisible on Google, absent from chat answersRead every User-agent groupAllow the search bots
CDN or WAF block403 responses or no bot requestsReview bot settings and logsKeep the Search class open
Missing from the indexNo appearance in any searchUse URL Inspection in both toolsSubmit sitemaps, use IndexNow
JavaScript dependencyPresent on Google, absent from ChatGPTSearch the source for your main textUse SSR or prerendering
Restrictive directivesYour text never shows up in answersRead meta tags and X-Robots-TagRemove unneeded restrictions
Buried answerThe answer quotes a rival's sentenceScan headings and first sentencesMove the answer to the top
Entity ambiguityThe model confuses your brandAsk the model about your brandStandardize the name, add schema
No third-party proofThe same rivals win every recommendationSearch your brand off your own siteEarn PR and genuine reviews
Stale contentOld figures and product namesCompare dates with the numbersUpdate the substance, not the date
Thin pageThe page repeats competitorsLook for a sentence only you could writeAdd data and experience
Login or paywallText disappears when logged outCheck the source in a private windowPublish an open summary
llms.txt expectationsThe team postpones real fixesReview the order of your roadmapFix the first eleven mistakes first

How do you measure whether the fixes worked?

Do not expect results the day after a technical fix. OpenAI says robots.txt changes take about 24 hours to reach its systems, while Google and Bing need to recrawl the page. So once the fix is live, request indexing in Search Console and submit the changed URLs through IndexNow.

  • Week one: Watch the server logs to confirm that OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot requests return a 200 status.
  • Month one: Re-run your original prompt set with the same wording and compare citations against the earlier log.
  • Ongoing: Track citation counts and cited pages in the Bing AI Performance report, and AI referral visits in GA4.

Above all, look at the trend rather than a single chat result. Answers vary from session to session; what matters is whether citations and the number of citing platforms grow over the following weeks.

When should you bring in outside help?

In practice, your own team can often fix access and directive mistakes quickly. Work that touches infrastructure needs more coordination, however. Moving a JavaScript-heavy site to server-side rendering, rewriting CDN rules or running an entity and reputation program are typical examples.

In the AI SEO and GEO services my team and I provide, we run all twelve checks from this list alongside log analysis. We first prove which layer is breaking, then plan the fixes by impact. If you want the broader strategic frame, read my GEO guide.

Whichever route you take, keep the diagnosis in writing. Note which step you checked, when you checked it and what you found. Then, at the next change, you will know what worked from records rather than guesswork.

In short, where should you start today?

In short, a site that is not cited by AI rarely suffers from one big mistake; it usually suffers from several small breaks that add up. Take these three steps today:

  1. Check your robots.txt file and your CDN bot settings on the same day.
  2. Confirm that the main text sits in the raw HTML of your ten most important pages.
  3. Save your prompt set and repeat the same test one month from now.

These three steps show you which layer you are losing in. After that, the rest is steady work in the right order. Once you find the source of the problem, the fix often takes less time than you feared.

Frequently Asked Questions

Do I need to allow GPTBot to appear in ChatGPT?
No. For ChatGPT search, the bot that matters is OAI-SearchBot. According to OpenAI, GPTBot crawls for model training, while OAI-SearchBot surfaces websites in ChatGPT's search features. If you do not want to contribute training data, you can block GPTBot and keep OAI-SearchBot open. OpenAI says robots.txt changes take about 24 hours to apply.
I rank first on Google, so why doesn't AI Overviews cite me?
A top organic ranking does not guarantee a citation in AI Overviews. Google's baseline requirement is that the page sits in the index and qualifies for a snippet, while the selection itself depends on which passage best answers the query. Check your nosnippet and max-snippet rules first, then make sure the answer sits clearly at the top of the page.
What should I check if my site runs on Cloudflare?
Review the Search, Agent and Training bot classes in AI Crawl Control. Cloudflare's AI bot blocking targets Training and Agent and leaves Search open, but it counts mixed-purpose crawlers as Training. Also confirm in your server logs that requests from OAI-SearchBot, Claude-SearchBot and PerplexityBot return a 200 status code.
Will an llms.txt file make AI cite my site?
No, there is no strong evidence that it does today. Google states in its official documentation that you do not need new machine-readable files or AI text files for AI Overviews and AI Mode. You can keep the file as a harmless extra, but fix access, indexing, rendering and content first if you want more citations.
How soon will I see results after fixing these issues?
Technical fixes usually show an effect within days, while content and reputation work takes weeks. OpenAI says robots.txt changes reach its systems in about 24 hours. Google and Bing have to recrawl your pages, so request indexing in Search Console and use IndexNow for Bing to speed things up.
  • AI visibility
  • AI Overviews
  • ChatGPT search
  • OAI-SearchBot
  • Cloudflare bot blocking
  • GEO
  • technical SEO
Share:
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.