What Is Robots.txt? How to Create and Test It

What is robots.txt and what does it do?
Robots.txt is a plain text file in your site's root directory that tells search engine crawlers which URLs they may access. Also, its main job is to manage crawl traffic. It does not hide a page from Google; for that you need a noindex directive or password protection.
So think of the file as a doorman. A crawler arrives, reads the file first, and then walks only through the areas you allow. So robots.txt works like a sign, not a lock. Also, well-behaved bots respect it, while malicious bots usually ignore it.
Most sites use the file for three reasons, so let us go through them. First, you keep your server from answering pointless requests. Second, you steer crawl activity toward the pages that matter. Third, you close off areas that generate endless URLs, such as internal search results.
How does a robots.txt file work?
When a crawler visits your domain, its first request goes to yourdomain.com/robots.txt. If it finds the file, it reads the rules and picks the group that matches its own name. If the server returns a 404, crawlers treat the whole site as open.
Then you write the rules in groups. Each group starts with a user-agent line and carries allow or disallow lines beneath it. According to Google's documentation, a crawler follows only the single most specific group that matches it. So rules from your general group do not flow into a specialized group automatically.
The server response code matters too. First, a 200 means the rules apply. With 4xx errors, crawlers usually behave as if no restrictions exist. With 5xx errors, Google may pause crawling until it can read the file again, which is why you should monitor the file's uptime.
For sources, read Google's robots.txt introduction and the RFC 9309 standard that formalizes crawler behavior.
Which directives can you use in robots.txt?
The syntax rests on four directives, and each one is simple. You write each one as a field: value pair, one per line. Directive names ignore case, but the path values do not.
- User-agent: names the crawler the rule applies to. An asterisk covers all crawlers.
- Disallow: gives a path the crawler should not visit. An empty value blocks nothing.
- Allow: opens a specific path inside a folder you have disallowed.
- Sitemap: states the full URL of your XML sitemap.
You also get two wildcards. The asterisk matches any sequence of characters, and the dollar sign marks the end of a URL. For example, Disallow: /*.pdf$ covers every address that ends in .pdf.
Also, comments start with a hash sign. Use them to leave notes for your team, because six months from now nobody remembers why a rule exists.
How do you create a robots.txt file?
Open a plain text editor, save the file as robots.txt with UTF-8 encoding, and upload it to your site's root directory. Google's documentation says a site can have only one robots.txt file, and it must sit directly under the domain.
So here is the sequence our team follows on new projects.
- List the areas you do not want crawled: admin panel, cart, filter results, internal search.
- Decide which crawlers each area applies to.
- Write the rules in user-agent groups.
- Add the sitemap line at the end.
- Upload the file to the root and open it in a browser.
- Validate it with the robots.txt report in Search Console.
If you prefer not to write it by hand, our robots.txt generator builds a draft from your choices. Still, review the output against your own site structure.
Where do you upload a robots.txt file?
The file must sit at the root of each host, including the protocol and subdomain. For example, https://www.example.com/robots.txt controls only that exact origin. Then a file in a subfolder gets ignored by crawlers.
Also, every subdomain needs its own file. So blog.example.com gets one file and example.com gets another. Likewise, Google treats the http and https versions separately, which is why redirecting everything to one version makes your life easier.
On platforms like WordPress, the system sometimes generates a virtual file for you. Once you upload a physical file, the virtual one stops responding. So open the address in a browser to see which one answers.
Google processes roughly the first 500 KiB of the file, and anything beyond that limit gets ignored. So if your file approaches that size, simplify your rules.
What are some robots.txt examples?
First, the examples below show common scenarios. Adapt each one to your own paths, because copying and pasting often blocks the wrong areas.
A basic file that allows everything. This is often enough for a small business site:
- User-agent: *
- Disallow:
- Sitemap: https://www.example.com/sitemap.xml
An e-commerce example. It closes the cart, checkout, and internal search pages:
- User-agent: *
- Disallow: /cart/
- Disallow: /checkout/
- Disallow: /search/
- Sitemap: https://www.example.com/sitemap.xml
An admin area example. It appears often on WordPress sites. Allowing the admin-ajax file keeps some themes working:
- User-agent: *
- Disallow: /wp-admin/
- Allow: /wp-admin/admin-ajax.php
One detail deserves attention: paths begin with a slash. When you close a folder, keep the trailing slash, because /search without it also matches unrelated addresses like /search-results. Teams misread this small detail more often than you might expect.
Blocking the entire site. Also use Disallow: / only on staging environments. If it reaches the live site by accident, you lose traffic fast.
What is the difference between robots.txt and noindex?
Robots.txt controls crawling, while noindex controls indexing. Also, they act at different stages. A bot first crawls the page, then decides whether to index it. Robots.txt affects the first stage and noindex affects the second.
Here is the critical point. If you block a page in robots.txt, Google never fetches it, so it never sees the noindex tag on that page. The page can still show up in results if other sites link to it. Google's guide on blocking indexing explains why a noindex page must stay crawlable.
| Feature | Robots.txt | Noindex |
|---|---|---|
| What it controls | Crawling | Indexing |
| Where you set it | File in the root directory | Page meta tag or HTTP header |
| Removes a page from results? | No | Yes, after Google recrawls |
| Must the bot read the page? | No | Yes |
| Best use | Reducing crawl load | Removing a page from results |
Why does a blocked page still show up in Google?
Because Google can index a URL without reading its content. Links from other sites are enough. In that case the result often appears without a description, or with a note that no information is available for the page.
This is the most common misunderstanding we see in the field. So site owners disallow a folder and expect the pages to vanish within days. In reality, blocking does not delete anything from the index; it only stops Google from refreshing the content.
The right order looks like this. First, leave the pages crawlable and add noindex. Then confirm in Search Console that the pages dropped out of the index. After that, you can use robots.txt to cut crawl load if you still want to. For sensitive data, even noindex is not enough, so use password protection or remove the page entirely.
Which pages should you block with robots.txt?
First, block areas that have no value in search results and waste crawl activity. Every site differs, but these groups show up again and again.
- Admin and login pages.
- Cart, checkout, and account pages.
- Internal search result pages.
- Filter and sort parameters that create endless combinations.
- URLs that carry session IDs or tracking parameters.
- Staging and test areas.
That said, every restriction carries risk. For example, some filter pages answer real search demand and belong in the index. So check Search Console data and organic traffic per page before you write a rule.
Here is a field example. Imagine a store that creates a separate URL for every color and size combination. That means thousands of near-duplicate pages, and bots waste time on them. You would consider closing those parameters, but first you must measure which filters attract real searches.
Our team also studies crawl statistics at this point. Seeing which folders Googlebot requests most often makes the list of blockable areas clearer. For crawl frequency problems, read our article on why Googlebot crawls less.
Which pages should you not block with robots.txt?
Do not block your CSS, JavaScript, or image files. Google renders your page the way a browser does. If it cannot reach your styles and scripts, it may misread the page, and both mobile usability and content evaluation suffer.
Also avoid closing any page you want to rank. Category, product, blog, and service pages must stay crawlable. A broad path in a rule can cover these areas without you noticing.
Blocking pages that carry a canonical tag causes trouble as well. Google must read the page to see the canonical signal. So manage duplicate content with canonical tags, not with crawl blocks.
Finally, closing an address that also appears in your sitemap sends a conflicting signal. Search Console warns you in that case, and the indexing status of those pages turns inconsistent.
How do robots.txt and your sitemap work together?
They complement each other. Robots.txt tells bots where not to go, and the sitemap hands them a list of important URLs. Adding a sitemap line at the end of robots.txt lets bots find your map on their first visit.
Write the sitemap address as a full URL, because relative paths do not work. If you have several maps, put each on its own line. On large sites, one line pointing to the sitemap index file is enough.
The map should contain only addresses you want indexed, that return a 200, and that use the canonical URL. Remove blocked, redirecting, or noindex addresses from it. That way you send Google a consistent signal.
When building the map by hand gets tedious, try our XML sitemap generator. For the date fields, our sitemap lastmod guide gives the details. Finally, submit the map in Search Console too.
How do you test a robots.txt file?
After you publish the file, run three checks. First, open the address in a browser and confirm the text looks right. Second, check that the server returns a 200. Third, look at the robots.txt report in Search Console.
That report shows when Google last read the file and which lines caused problems. Google's guide to creating a robots.txt file also points to its open-source robots.txt library for testing.
After a critical change, test your key pages one by one with the URL Inspection tool. For example, if a category page shows a "blocked by robots.txt" warning, your rule is broader than you intended. New to the tool? Our Google Search Console guide walks you through it.
Test with both the mobile and desktop crawlers as well. Some rules affect only one bot group, and you see that difference only with two separate checks. Finally, keep the file under version control so you can roll back fast.
What are the most common robots.txt mistakes?
The most common mistake is carrying the staging rule Disallow: / over to the live site. Typos in paths, missing slashes, and case confusion follow close behind. We covered each fix in our post on common robots.txt mistakes, so we will not repeat it here.
Still, a short checklist helps:
- Does the file sit in the root, not in a subfolder?
- Is the encoding UTF-8?
- Is the sitemap line a full URL?
- Are the CSS and JS folders open?
- Did you leave noindex pages open to crawling?
Another frequent error is assuming one file covers all subdomains. A shop subdomain needs its own file, and rules on the main domain never reach it.
Also, writing "noindex" inside robots.txt does nothing. Google does not support that directive. Use the page tag or the HTTP header instead.
How do you manage AI bots with robots.txt?
You manage AI bots by writing a separate group for each user-agent name. For example, you open one group for GPTBot and another for ClaudeBot, then give each an allow or disallow rule. These bots say they follow the rules, but compliance is their own choice.
Think about your business goal first. If you want your content cited in AI answers, blocking access may lower that chance. If copyright and data policy come first, blocking makes sense. So this is a strategic decision, not a technical one.
For the current bot names, which bots train models and which serve search, and the setup we recommend, read our guide to AI crawlers. Also check each vendor's official documentation, because names change over time.
How does robots.txt affect indexing problems?
Robots.txt does not control indexing directly, but its indirect effects are large. Pages in a folder you closed by mistake cannot refresh, and new pages get discovered late. So robots.txt should be one of your first stops when you investigate "not indexed" reports.
If Search Console shows "blocked by robots.txt", first ask whether that was a deliberate choice. If the page should stay in the index, narrow the rule and request a recrawl. Our guide to finding unindexed pages covers the full process.
E-commerce sites are more complex. Product pages can miss the index for many reasons, and we gathered them in why product pages are not indexed.
How often should you review your robots.txt file?
Review the file after every major site change and at least once a quarter. When you add a new section, migrate the site, or swap a theme or plugin, old rules can lose their meaning.
Google usually caches the file and refreshes it about once a day. So a change may not show up instantly. If you fixed an urgent error, request a recrawl in Search Console.
In migration projects, the robots.txt check belongs among the first items on your launch list. For the other steps, see our website migration SEO checklist.
On a business site, tie this work to a routine. If you want a team to run technical audits for you, our SEO consulting service covers these checks.
What are the best practices for robots.txt?
The best practice is to keep the file short and readable. Every line needs a reason. If you cannot explain a rule, deleting it is often safer.
- Write down what you want to block before you touch the file.
- Keep rules as narrow as possible.
- Explain each group with a comment.
- Test changes on a staging site first.
- Check the Search Console report after you publish.
- Add the sitemap address to the file.
Also, do not treat robots.txt as a full SEO strategy. Title tags, canonical structure, and site architecture matter just as much. You can draft titles with our meta tag generator, and for the wider picture, read our technical SEO guide.
What should you do when your robots.txt keeps changing?
If the file changes often, the problem usually sits in the process, not the file. When several people edit it without any record, unexpected blocks appear.
So name an owner and log every edit with a short note. Put the file under version control. If you set up a monitoring alert for critical rules, you learn right away when the file changes unexpectedly.
Some content management systems and plugins write the file automatically. In that case, decide which tool has the final say. Otherwise a plugin update may silently erase the rules you wrote by hand.
How do user-agent groups work inside robots.txt?
Each user-agent line opens a new rule group. If you write a separate group for Googlebot, Googlebot drops the general asterisk group entirely and reads only its own. So you must copy your general rules into the specific group too.
Google runs several crawlers. The image bot, the news bot, and the ads bot use separate names. For example, AdsBot ignores the general asterisk group, so you must name it directly. If you run ads, knowing this matters.
Picture a practical case. You closed the /search/ folder in the general group. Then you added a Googlebot-Image group that lists one image folder. The image bot no longer sees the /search/ rule, because it applies only its own group.
In short, keep groups to a minimum. Before you open a group for a special bot, ask whether you truly want different behavior. Extra groups make maintenance harder and invite mistakes.
How do wildcards and matching rules work in robots.txt?
When several rules match the same address, Google picks the most specific one, meaning the rule with the longest path. If an allow and a disallow tie in length, the less restrictive allow wins. Knowing this logic helps you read complex files.
For example, say you wrote Disallow: /products/ and Allow: /products/popular/. The popular folder stays open, because the allow path is longer and more specific. The other product folders stay closed.
However, the asterisk can surprise you. The line Disallow: /*? closes every address that contains a question mark. That helps when you target filter parameters. However, it may also hit real pages that open through pagination or campaign parameters.
So test every wildcard rule together with sample addresses it affects. Do not publish a rule before you see exactly which URLs it covers.
How does robots.txt affect crawl budget?
Robots.txt improves crawl budget indirectly by steering Googlebot's requests toward valuable pages. Still, most small sites will not feel this effect. Google says crawl budget mainly concerns very large sites that update often.
So assess the scale of your own site. On a business site with a few hundred pages, fine-tuning robots.txt will not change rankings. On a store that produces hundreds of thousands of filter combinations, however, this control makes a clear difference.
Watch your crawl data. The crawl stats report in Search Console shows how many requests Google makes each day and how fast your server answers. If one folder receives a disproportionate share, consider closing it.
Finally, remember that blocking a page does not automatically move its budget to other pages. If your server is slow or your architecture is weak, fix those first.
How do you edit robots.txt on a WordPress site?
If WordPress finds no physical file, it serves its own virtual robots.txt. You can edit that response through your theme or an SEO plugin. Many SEO plugins offer a file editor in the dashboard, so you can write rules right there.
You can also upload a physical file to the server. When a physical file exists, WordPress does not serve the virtual one. If both a plugin and a file exist, open the address in a browser to see which one wins.
Watch the setting that asks search engines not to index your site. That option adds a rule that closes the whole site. Leaving it ticked on a live site is a common accident.
Lastly, check the file after each plugin update. Plugins sometimes rewrite default rules and erase your custom lines.




