Artificial Intelligence

AI Crawlers Explained: How to Manage GPTBot, ClaudeBot and PerplexityBot with robots.txt and llms.txt

Talha AslanTalha Aslan 18 min read 2 views

What is an AI crawler and why does it visit your site?

An AI crawler is an automated bot that an AI company sends to read web pages; some collect content to train models, some index pages so AI search can cite them, and some open a page the moment a user asks a question. That is why blocking them all with one rule is a mistake.

When I review server logs, this is where I see the clearest change of recent years: next to Googlebot and Bingbot, names like GPTBot, ClaudeBot and PerplexityBot now show up regularly. Moreover, most site owners add a block to robots.txt that either shuts out every bot or lets every bot in, without knowing what each one does. In this guide I summarise what the official documentation says about each AI crawler, and then I walk you through building a policy with robots.txt and llms.txt.

How do AI crawlers differ from classic search bots?

A classic search bot crawls your page, indexes it and shows it as a link in the results. In return you get traffic, so the trade is clear. With AI crawlers, however, that trade works in three different ways, and each one gives you something different back.

For example, a training bot reads your content, but because the content blends into a model's general knowledge, it does not send you a direct link. A search bot, on the other hand, can show your site as a source card in products like ChatGPT Search or Perplexity. An agent triggered by a user simply opens one page to answer the question at hand. Therefore the question "should we allow AI?" is too broad; the right question is "which purpose are we allowing?"

I covered the wider technical effects of this shift in my article on technical SEO after AI. Here I focus only on the bots, the robots.txt rules and the llms.txt file.

Which three jobs should you use to group AI crawlers?

When you put the official documents side by side, you see that the companies split their bots into almost the same three categories. I recommend you build your policy on those categories too, because each one has a different business outcome.

  1. Training bots: They collect content to train future models. GPTBot and ClaudeBot belong here, while Google-Extended and Applebot-Extended handle the same decision as permission tokens rather than bots.
  2. Search bots: They index pages so AI powered search can cite them. OAI-SearchBot, Claude-SearchBot and PerplexityBot belong here.
  3. User agents: They open a page on demand when a user asks a question or shares a link. ChatGPT-User, Claude-User and Perplexity-User belong here.

In short, closing a training bot does not automatically close your visibility in search. However, if you close a search bot, you seriously reduce your chance of being cited in that product's answers.

What do OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User do?

OpenAI describes each bot separately in its official bot documentation. GPTBot collects content for training generative AI foundation models. OAI-SearchBot, in contrast, works so that sites can appear in ChatGPT search results. Both follow robots.txt, and OpenAI notes that it can take about 24 hours from a robots.txt update for its systems to adjust.

ChatGPT-User is different: it opens pages during actions a user starts inside ChatGPT or custom GPTs. OpenAI adds two important notes about it. First, because a user initiates these actions, robots.txt rules may not apply. Second, ChatGPT-User is not used to decide whether content may appear in search.

In practice I draw this conclusion: if you want to be cited as a source in ChatGPT Search, you should allow OAI-SearchBot. Meanwhile, you can make the GPTBot decision separately, based on your policy about training use. The documentation also lists OAI-AdsBot, which checks ad landing pages; if you do not advertise there, you are unlikely to meet it.

One common mistake is worth noting as well: some sites misspell a bot name or list only an outdated one. You need to name each bot exactly as the documentation does, in its own User-agent line; otherwise the rule never applies to the bot you had in mind.

Why does Anthropic run ClaudeBot, Claude-SearchBot and Claude-User separately?

Anthropic defines three bots on its support page. ClaudeBot collects web content that could contribute to training its generative models. Claude-SearchBot navigates the web to improve the quality of search results and the relevance of Claude's answers. Claude-User, finally, lets Claude access websites when a user asks a question.

Anthropic states clearly that all three follow robots.txt and that you can block each one separately. In addition, it supports the non-standard Crawl-delay extension, so if your server struggles under load you can set a wait time between ClaudeBot requests.

There is one more important detail: Anthropic reminds you to edit the robots.txt file for every subdomain you want to restrict. For example, blog.yoursite.com and www.yoursite.com read different files. Consequently, on setups with many subdomains you have to write the policy for each one.

What is the difference between PerplexityBot and Perplexity-User?

Perplexity defines two agents in its developer documentation. PerplexityBot is designed to surface and link websites in Perplexity search results, and the company stresses that it is not used to crawl content for AI foundation models. This bot follows robots.txt.

Perplexity-User, by contrast, visits a page when a user asks a question, so it can give an accurate answer. According to Perplexity, this agent generally ignores robots.txt rules because a user initiates the fetch. If you want to manage that access, Perplexity suggests using its published IP lists in your firewall.

So if you want to be cited in Perplexity, you should not block PerplexityBot. In fact, the documentation advises site owners to allow this bot and whitelist its IP ranges. If you use a firewall or a CDN bot protection layer, make sure it does not block these requests by accident.

Why are Google-Extended and Applebot-Extended not real bots?

These two names are the most misunderstood part. In its crawler documentation, Google writes that Google-Extended has no separate HTTP user agent string, that crawling happens with existing Google user agents, and that the name works in robots.txt only as a control token. With this token you decide whether crawled content may be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI.

On the same page, Google states that Google-Extended does not affect inclusion in Google Search and is not used as a ranking signal. In other words, closing this token does not hurt your organic rankings. Note, however, that it does not stop Googlebot either: Googlebot still crawls and indexes your site. Therefore trying to leave Gemini training by blocking Googlebot is the fastest way to lose organic traffic.

Apple works in a similar way. According to Apple's Applebot page, Applebot-Extended does not crawl web pages; it decides whether data collected by the standard Applebot may be used to train Apple's generative models. Furthermore, rules for Applebot-Extended are not considered in ranking for search.

AI crawler comparison table

In the table below I gathered only what the official documents say. I suggest you go through it row by row and mark your own policy.

NameCompanyJobrobots.txt
GPTBotOpenAIModel trainingFollows
OAI-SearchBotOpenAIChatGPT search resultsFollows
ChatGPT-UserOpenAIUser actionMay not apply
ClaudeBotAnthropicModel trainingFollows, supports Crawl-delay
Claude-SearchBotAnthropicSearch qualityFollows
Claude-UserAnthropicUser questionFollows
PerplexityBotPerplexityPerplexity search resultsFollows
Perplexity-UserPerplexityUser questionGenerally ignores
Google-ExtendedGoogleToken for Gemini training and groundingToken only, no separate bot
Applebot-ExtendedAppleToken for Apple model trainingToken only, does not crawl

The clearest lesson from the table is this: you cannot fully control user agents with robots.txt. For content that truly must stay private, you should therefore use a login wall or server side access control instead.

Does robots.txt really stop AI crawlers?

robots.txt is not a lock; it is a list of invitations and refusals. Even though RFC 9309 wrote the Robots Exclusion Protocol down, compliance depends on the bot's owner. The companies above officially state that they follow robots.txt, but that statement covers only bots that use their real names honestly.

That is why I recommend using robots.txt for two purposes. First, to tell compliant companies your intent clearly. Second, to manage your crawl budget and server resources. On the other hand, robots.txt alone does not protect sensitive data, customer portals or paid content; moreover, because the file is public, you also announce the names of the folders you want hidden.

Caching matters too. OpenAI says changes can take about 24 hours to take effect, and other bots also reread the file at intervals. In short, after you change a rule you should measure its effect the next day, not the next minute.

Can you block model training and still appear in AI search?

Yes, according to the official documentation this split is possible, and for many brands it is the most balanced option. You write Disallow for training bots and Allow for search bots. As a result, your content stays out of future training data, but you keep your chance of being cited in ChatGPT Search, Claude and Perplexity answers.

Still, I should be honest about one thing: nobody can measure precisely what staying out of training data does to long term brand awareness. A model's knowledge of your brand partly comes from training data. Hence, for a new brand chasing awareness, allowing training bots can make sense, whereas for a publisher selling original research or paid content, blocking them is the more defensible choice.

Tracking how AI answers mention your brand is a separate job; I explained it in my article on how your brand shows up in ChatGPT and Gemini.

How do you write robots.txt for three different policies?

I prepared the three templates below around the scenarios I discuss most often with clients. Instead of writing the file by hand, you can use the robots.txt generator, pick the rules and copy the output.

Scenario 1, fully open: You write no special rule for any AI crawler, so your general User-agent: * rule applies. This suits new brands aiming for awareness.

Scenario 2, training closed, search open: You write separate blocks like these.

  • User-agent: GPTBot, followed by Disallow: /
  • User-agent: ClaudeBot, followed by Disallow: /
  • User-agent: Google-Extended, followed by Disallow: /
  • User-agent: Applebot-Extended, followed by Disallow: /
  • Allow: / for OAI-SearchBot, Claude-SearchBot and PerplexityBot

Scenario 3, selective: You keep your blog and guides open, and you close folders such as price lists, customer areas or internal search results to every bot. This way AI products read your most valuable public content, while low value or private pages stay out.

In every block, use the bot name exactly as the documentation spells it. For instance, if you write "ClaudeSearchBot" instead of "Claude-SearchBot", the rule matches no bot at all.

What should you watch for with Crawl-delay and subdomains?

Crawl-delay is not among the core rules of RFC 9309, so not every bot reads that line. Anthropic says it supports the extension for its bots. Google, however, ignores Crawl-delay. Thus you cannot solve server load with this line alone; if needed, you should set rate limits at the CDN or web server level.

With subdomains, I often see the same mistake: a company closes training bots on the main site, but the help center or blog subdomain publishing the same content stays open. Each subdomain reads the robots.txt at its own root. So after writing the policy, list all your hostnames and check each file one by one.

Also make sure your robots.txt returns a 200 status code. If the file throws a server error, bot behaviour can change and you may end up with a result you did not want.

How do you tell fake AI crawler requests from real ones?

A user agent string is a label anyone can copy. Not every request in your logs that says "GPTBot" really comes from OpenAI. Aggressive scrapers in particular imitate well known bot names so they do not get blocked.

To verify, you use the IP ranges the companies publish. For example, Perplexity shares IP ranges for PerplexityBot and Perplexity-User in separate JSON files, and OpenAI links to IP lists for its bots on its bot page. Google, for its part, recommends verifying its crawlers with reverse DNS and its published IP ranges.

  • Does the user agent name match the official documentation exactly?
  • Is the request IP inside the range the company publishes?
  • Do the request rate and page pattern match expected behaviour?

You can treat requests that fail these checks as fake and block them in your firewall. That way you stay open to real bots while keeping resource hungry imitators out.

What is llms.txt and why is it not a standard?

llms.txt is a proposed file at the root of a website that summarises the site's most important pages for large language models in Markdown. Jeremy Howard proposed it in September 2024 on llmstxt.org, and the site itself calls the text a proposal.

I stress this for a reason: llms.txt is not an IETF standard like robots.txt, and it is not a permission mechanism either. It blocks no bot and allows no bot. Moreover, none of the bot documents from OpenAI, Anthropic, Perplexity or Google that I cited above say their crawling behaviour changes based on llms.txt.

So why do I still prepare one? Because it costs little, does no harm, and gives a clean map of your site to coding assistants, documentation tools or agents that a user points directly at the file. In short, think of llms.txt as a tidy introduction document, not a ranking factor.

How do you structure an llms.txt file?

The proposal defines a simple Markdown order. You place the file at your root as /llms.txt and follow this layout.

  1. An H1 heading with the name of the site or project; this is the only required section.
  2. A blockquote that explains in a few sentences what the site does.
  3. Optional paragraphs with more detail.
  4. File lists separated by H2 headings, each line holding a Markdown link and an optional short note.
  5. By convention, a section named "Optional" for secondary information.

The proposal also encourages offering clean Markdown versions of informative pages by adding .md to the same URL. This part is entirely optional and not required for most business sites.

Rather than writing the file from scratch, you can fill in the title, summary and link groups in the llms.txt generator and take the output. Then, before publishing, check that every link returns 200.

Which mistakes should you avoid when writing llms.txt?

In the llms.txt files I review on client sites, I see the same mistakes again and again. Most of them come from misunderstanding what the file is.

  • Writing permission rules: Adding a "Disallow" line to llms.txt blocks nothing; blocking is the job of robots.txt.
  • Dumping the whole site: Copying a thousand links from the sitemap defeats the purpose. The proposal asks for a curated, annotated list.
  • Marketing language: Filling the summary with slogans tells a model nothing. Write in plain sentences what you sell, whom you serve and which page covers what.
  • Forgetting updates: If removed pages stay in the file, agents hit 404s. Review the file together with your sitemap updates.
  • Expecting miracles: Publishing llms.txt does not raise traffic by itself; it does not replace content quality.

If your sitemap is already clean, preparing llms.txt becomes much easier. For that reason I suggest you first tidy your map with the XML sitemap generator, and then pick the most important pages for llms.txt.

How do you read AI crawler traffic in server logs?

Before you write a policy, you need to see the current state. By filtering the user agent field in your web server access logs, you can find which AI crawler requests which pages and how often. JavaScript based tools such as Google Analytics usually do not show these bots, because most of them do not run the scripts on the page.

I suggest you look at three things in the logs. First, status codes: if search bots get 403 or 429, your security layer may be blocking them without you noticing. Next, the most requested pages: if bots wander through filter and parameter pages instead of your valuable content, you should fix your crawl structure. Finally, request volume: if one bot strains your server, rate limits or Crawl-delay come into play.

After this analysis, your robots.txt decisions rest on data rather than guesses. Furthermore, you can pull the same data a few weeks later and compare the effect of the change.

How do you test a robots.txt change before it goes live?

One wrong line in robots.txt can close your whole site to bots you wanted. That is why I recommend testing a change on a copy before you write it to the live file. Accidentally adding Disallow: / to the User-agent: * block is one of the most expensive mistakes I have seen.

During testing you first check which group applies to each bot. Under RFC 9309, a crawler follows the most specific group that matches its name, and once it finds that group it ignores the general star block. Consequently, if you wrote a separate block for GPTBot, you need to repeat the folder restrictions from the general block inside that block too.

Then you can use the robots.txt report in Google Search Console to confirm that Google reads the file. For other bots, the most reliable test is to check your server logs a day or two after the change and see which paths that user agent requests. This way you know the rule works in practice, not just in theory.

What should you do when a new AI crawler is announced?

The bot list is not fixed. As companies launch products, they announce new user agents, rename old ones or change their jobs. For example, Anthropic and OpenAI have over time split search and training work into separate bots. So a robots.txt you write once and forget can be incomplete within months.

My advice is to keep your written policy category based. When a new bot appears, you first read its official documentation, then decide whether it is a training bot, a search bot or a user agent. After that you apply the same rule you use for the other bots in that category. As a result, you do not have to restart the debate for every new name.

I also recommend a log review every quarter. If you see an unknown, high volume user agent, search for its name in the company's official documentation to confirm its identity; if you cannot find it, caution is the better choice.

Which AI crawler policy fits which kind of site?

There is no single right policy; your business model drives the decision. In the consulting projects my team and I run, we generally use this framework.

  • Local service businesses and B2B firms: Staying open to all bots usually makes sense. Being recommended in AI answers is worth more than keeping content out of training.
  • E-commerce sites: Keep product and category pages open to search bots, and close cart, account and filter combinations.
  • Publishers and original research producers: Closing training bots while keeping search bots open is often the most defensible balance.
  • Paid content and member areas: Protect these with a login wall, not with robots.txt.

Whichever group you are in, I suggest you put the decision in writing. Then, when a new bot is announced, you can add a rule quickly with the same logic.

Does crawl access alone make you visible in AI answers?

No. Opening the door to AI crawlers is only a precondition. For an AI product to cite you, your content has to answer the question clearly, carry trust signals and be technically readable. For instance, if you load important text only with client side JavaScript, bots that do not run scripts may see an empty page.

I described this content work under GEO, generative engine optimization. If you wonder how SEO, GEO and AEO differ, this comparison is a good start. For the role of structured data, you can also try the schema generator.

In short, robots.txt and llms.txt organise the door; the quality of your content decides what visitors find inside.

A 30 minute AI crawler audit checklist

After reading this guide, you can work through the list below in order. For most sites half an hour is enough; on large e-commerce sites or setups with many subdomains it can take longer, so you may split the work into several sessions.

  1. Open the robots.txt of every domain and subdomain and confirm it returns 200.
  2. Check whether the file contains a deliberate decision for GPTBot, ClaudeBot, Google-Extended and Applebot-Extended.
  3. Make sure OAI-SearchBot, Claude-SearchBot and PerplexityBot are not blocked by accident.
  4. Confirm that bot rules in your CDN or firewall do not return 403 to these bots.
  5. Filter the last two weeks of logs and group AI crawler requests by page and status code.
  6. Prepare your llms.txt, or verify the links in your existing file.
  7. Wait at least 24 hours after the change and compare the logs again.

If you would rather not run these steps yourself, my team and I can handle the whole process, from log analysis to robots.txt policy, as part of our SEO consulting service. If you want to review the technical basics in general, see my technical SEO tips as well.

Frequently Asked Questions

If I block GPTBot, will I disappear from ChatGPT Search?
No, you can still appear. According to OpenAI's documentation, GPTBot works for model training, while OAI-SearchBot works for ChatGPT search results. If you disallow GPTBot but keep OAI-SearchBot open, your content stays out of training, yet you keep your chance of being cited in ChatGPT Search answers. The change can take about 24 hours to apply.
Will blocking Google-Extended hurt my Google rankings?
No, it will not. Google officially states that Google-Extended does not affect inclusion in Google Search and is not used as a ranking signal. The token only controls whether crawled content may be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI. It is not a separate bot, and Googlebot keeps crawling as usual.
Does an llms.txt file improve SEO rankings?
There is no official evidence that it does. llms.txt is a proposal that Jeremy Howard published on llmstxt.org in 2024, not a standard. The bot documents of the major AI companies do not say their crawling changes based on this file. Still, because it costs little, you can prepare one to introduce your site neatly to AI agents and tools.
Can robots.txt stop all AI access to my site?
Not completely. Official bots follow robots.txt, but OpenAI says the rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt. There are also scrapers that ignore the rules entirely. You should protect content that must stay private with a login wall or server side access control rather than robots.txt.
I see GPTBot in my logs; is it really OpenAI?
Not always. A user agent string is easy to fake. You should verify by comparing the request IP with the IP lists the company publishes. OpenAI and Perplexity both link to IP ranges in their bot documentation. If the address is not on the list, you can treat the request as fake and block it in your firewall without affecting real bots.
Can I limit the crawl rate of ClaudeBot?
Yes, you can. Anthropic says it supports the non-standard Crawl-delay extension, so you can add a Crawl-delay line to the User-agent: ClaudeBot block in robots.txt to set a wait between requests. However, not every bot reads this line, and Google ignores it. For general load problems, rate limiting at the CDN or server is more reliable.
#ai crawler#GPTBot#ClaudeBot#PerplexityBot#robots.txt#llms.txt#technical SEO
Share:
Talha Aslan
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.

WhatsApp Call Now