Robots.txt Mistakes: The Most Common Errors and How to Fix Them

What are the most common robots txt mistakes?
Robots txt mistakes are errors in the crawl instructions you give search engine bots: blocking CSS and JavaScript, writing a Disallow rule that is far too broad, using robots.txt instead of noindex, ignoring case sensitivity, a broken Sitemap line and exceeding the 500 KiB limit. As a result, the usual outcome is a quiet loss of traffic.
The file only holds a few lines, so most teams write it once and forget it. However, a single slash can close an entire site to Googlebot. I have run technical audits since 2012, and some of the most expensive problems I have found lived in this tiny text file.
This guide does not cover blocking AI crawlers; I wrote about that separately in my guide to AI crawlers and llms.txt. Instead, I focus on the errors that break classic search crawling. Every point below rests on Google's official documentation and on the RFC 9309 standard.
What does a robots.txt file actually do, and what does it not do?
A robots.txt file is a plain text file in the root directory of your host. Specifically, it tells crawlers which paths they may request. In other words, it manages crawling, not indexing. Once you miss that distinction, half of the common errors follow on their own.
Google's introduction to robots.txt says it plainly: robots.txt is not a mechanism for keeping a web page out of Google. Also, the rules carry no enforcement. Each crawler decides whether to obey them, and malicious bots simply ignore the file.
- Crawl control: it tells compliant crawlers which URL paths they should not request.
- Sitemap discovery: it announces the location of your sitemap through the Sitemap line.
- No index removal: it cannot remove a page from search results.
- No security: it does not protect a private folder; in fact, it shows that folder's address to anyone who opens the file.
There is also the question of standards. For decades, robots.txt lived as an informal convention. In September 2022, RFC 9309 turned the Robots Exclusion Protocol into a formal standard. Google's parser follows it, so most of the rules below apply to other major search engines too.
So robots.txt is not a security tool. You need passwords for admin areas and staging sites. With that frame in place, it becomes much easier to see why the following errors keep coming back.
Why is blocking CSS and JavaScript files such a serious mistake?
Googlebot does not read a page as raw HTML only. Instead, it renders the page the way a browser does and looks at the result. Therefore, if you block theme files, stylesheets or key JavaScript bundles, Google cannot see what your visitors see.
Old habits still cause trouble here. For example, some older themes and outdated guides recommended closing whole plugin or content folders. Yet those folders often hold the files that build the menu, the product grid or the price box. As a result, Google sees an empty or broken page.
Google takes a balanced view on this. According to the official guide, you can block unimportant images, scripts or style files. That said, if their absence makes the page harder to understand, you should not block them. In practice the line is hard to draw, so my default is to leave resource files open.
Likewise, the same logic applies to images. If you want product photos to appear in image search, keep those folders open. On the other hand, blocking truly worthless files, such as old backup folders or temporary downloads, is a sensible choice. What matters is that every block reflects a deliberate decision.
To check, run a live test in the URL Inspection tool in Search Console and look at the rendered screenshot. If the page looks bare, list the resources Googlebot could not load. If you want to know how heavy scripts affect speed, my article on how JavaScript affects site speed is a good companion.
How can one wrong Disallow line shut down a whole site?
A Disallow rule works as a path prefix. So when you write Disallow: / you block every URL that starts at the root, which means the whole site. An empty Disallow line says the opposite: nothing is blocked. In short, the difference is a single character.
I see this most often on launch day. A developer adds Disallow: / to keep the staging site out of search, and the file then travels to production with everything else. Nobody notices, because the site works fine for visitors. Then, a few weeks later, the Search Console graph starts to slide.
The prefix logic also creates subtler errors. For instance, Disallow: /blog blocks the blog folder, but it also blocks /blog-news and /blogger, because both paths start with the same letters. If you mean the folder, end the rule with a slash: Disallow: /blog/ states your intent far more precisely.
- Read every Disallow line and write down which paths it actually covers.
- End the rule with a slash when you target a folder.
- Test a few real URLs whenever you use the wildcard (*) or the end anchor ($).
- Add a "robots.txt root rule" item to your launch checklist.
Can robots.txt remove a page from Google's index?
No. This is the most widespread misconception. When you block a page in robots.txt, Google stops crawling it; but if other pages link to that address, Google can still index the URL. In that case, you see a bare result with no description.
Google's introduction describes this scenario directly: a disallowed page can still end up in the index if other sites link to it. In Search Console, you will usually notice it as a status saying the page got indexed even though robots.txt blocks it.
Writing a noindex line inside robots.txt does not help either. Google stopped supporting that unofficial rule on 1 September 2019 and announced the change on the Search Central blog. Put simply, Google now supports only four fields: user-agent, allow, disallow and sitemap.
If you want to keep a page out of search results, you have three options:
- Add a noindex meta tag to the page and keep crawling open.
- Use the X-Robots-Tag HTTP header for non-HTML files such as PDFs.
- Protect private areas with a password, or remove the page entirely.
Why does combining noindex and Disallow backfire?
Because Google has to crawl a page to see its noindex tag. If you mark a page with noindex and also block it in robots.txt, Googlebot cannot get in, so it never reads the tag. The page can then stay in the index, which is exactly what you wanted to avoid.
This mistake usually starts with good intentions. A team wants to clean up a set of pages and thinks two safeguards beat one. However, the two safeguards cancel each other out. First add the noindex tag, then let Google recrawl the pages and drop them.
After the pages leave the index, you can add a robots.txt rule to save crawl capacity if you wish. Still, the order matters: noindex first, then observation, then the block. Otherwise, you create "ghost" URLs that linger for months.
Filter and sort parameters are where this error shows up most. For example, online shops often block thousands of parameter URLs and tag them with noindex at the same time. When you plan parameter handling, the decision flow in my URL parameters guide will help.
Which robots txt mistakes come from case sensitivity?
Two different rules apply inside the file. Field names such as user-agent, allow and disallow are case insensitive. Path values, however, are case sensitive. So Disallow: /Products/ does not cover /products/.
In practice, this matters most on sites where old and new URL structures overlap. For example, a previous version of a site used capitalised folder names, and the new version switched to lowercase. The robots.txt file, however, kept the old spelling. As a result, the area you meant to close stays fully open.
The reverse can also happen. If a CMS serves the same content under both uppercase and lowercase addresses, blocking only one version leaves the other crawlable. In that case the real fix is not robots.txt; it is a redirect to one canonical spelling.
My check is simple. I pull the paths Googlebot actually requests from the server logs and compare them with the robots.txt lines letter by letter. You can do the same with the log file analyzer. That way, you find robots txt mistakes from real crawl data rather than from guesswork.
Which rule wins when Allow and Disallow conflict?
Google picks the rule with the longest matching path. In other words, the most specific rule wins. If two rules match with equal length, Google applies the least restrictive one, which is the allow rule.
This logic surprises teams who think in terms of order. Many people assume the line that comes first takes priority. For Google, though, line order does not matter; the length of the match decides.
For example, if you write Disallow: /account/ and Allow: /account/login together, the login page stays crawlable. That is because the Allow rule matches a longer path. Used on purpose, this pattern is handy: you can close a folder and open a single page inside it.
Here is a trickier case. You might block a sort parameter with Disallow: /*?sort= and keep pagination open with Allow: /*?page=. Once a URL carries both parameters, you have to work out which rule matches the longer string. That calculation goes wrong more easily than it looks.
Wildcards make the maths harder. So if your file contains overlapping rules, test the outcome on several real URLs. If you want to rebuild your rule set from scratch, the robots.txt generator gives you a clean starting point; then review the conflicts by hand.
Why do user-agent groups merge in unexpected ways?
Robots.txt rules work in groups. Each group starts with one or more user-agent lines, and the rules below apply to those crawlers. A crawler follows only the group that matches it most specifically and ignores the rest.
Here is the common trap. If the file has a general asterisk group and a separate group for Googlebot, Googlebot reads only its own group. Any important block you placed in the asterisk group no longer applies to Googlebot. Many teams add a Googlebot group without realising they just lost their general rules.
On the other hand, if the file contains two separate groups for the same user agent, Google merges them into one. In complex files, this behaviour leads to surprises. It happens most on sites where several plugins or several teams add lines to the same file.
- Keep one group per crawler.
- If you open a dedicated Googlebot group, copy the general rules into it as well.
- Add comment lines between groups to record who added what and why.
Which robots txt mistakes appear in the Sitemap line?
The Sitemap line is the easiest way to tell crawlers where your sitemap lives. Even so, it attracts its own set of robots txt mistakes. The most common one is a relative address. Google expects the Sitemap value to be a full URL that includes the protocol.
The second mistake is pointing to a sitemap that moved or no longer exists. When a site gets rebuilt, the sitemap address often changes while robots.txt keeps the old one. The third is a mix of HTTP and HTTPS, or of www and non-www versions. The sitemap address should match the canonical version of your site.
There is some good news as well. According to the official documentation, you can list several Sitemap lines, and they do not belong to any user-agent group. Also, the sitemap does not have to sit on the same host as the robots.txt file.
- Check that the Sitemap line holds a full URL.
- Open the address in a browser; you should get a 200 response and valid XML.
- Keep URLs you block in robots.txt out of the sitemap, because the two signals contradict each other.
If you need to regenerate the file, the XML sitemap generator speeds things up. For the date field, see my article on the sitemap lastmod tag.
What problems do the 500 KiB limit and file format cause?
Google reads the first 500 kibibytes of a robots.txt file and ignores anything beyond that. RFC 9309 also requires crawlers to parse at least 500 KiB. For a small site the limit feels distant; however, files that software writes automatically can grow fast.
I see this most on sites that add a separate Disallow line for thousands of individual URLs. For example, a plugin that adds one line for every deleted product keeps inflating the file. Once it crosses the limit, the rules at the end stop working, and nobody knows which lines Google skipped.
The fix is to write patterns instead of single URLs. If the pages share a folder or a parameter, one rule can replace hundreds of lines. Also, handle deleted pages with a 404 or 410 status rather than with robots.txt.
Format has two rules too. Save the file as UTF-8, and separate lines with CR, CR/LF or LF. Smart quotes pasted from a word processor, invisible characters or another encoding can break parsing. If you want the full detail, read the RFC 9309 text.
How does the server's status code change the way Google treats robots.txt?
Google looks at the HTTP response for the file as closely as at its content. On a 2xx response, it processes the file as provided. On 4xx responses other than 429, it acts as if no file exists and crawls without restrictions.
The real danger lies in 5xx errors. According to the robots.txt specification, when the request returns a server error, Google stops crawling for the first 12 hours. For the next 30 days, it uses the last good cached copy. After 30 days, if the site itself is reachable, Google behaves as if there were no restrictions.
So even a brief server error on the robots.txt URL can slow the crawling of the whole site. I see this most when a firewall, a CDN or a maintenance mode setting blocks the request. If your crawl rate dropped for no clear reason, check this status code alongside the steps in my article on why Googlebot crawls less.
One more detail: 429 differs from the other 4xx codes. Google reads it as "slow down" rather than "no file", and it reduces the crawl rate. Keep that in mind when you write rate limiting rules that touch Googlebot.
Redirects have a limit as well. Google follows at least five redirect hops, then treats the file as a 404. Also, Google generally caches the file for up to 24 hours, so your change may not take effect right away.
Why do subdomains, protocols and ports catch teams out?
A robots.txt file applies only to the host, protocol and port where it sits. So the file on example.com does not cover shop.example.com. Every subdomain needs its own robots.txt file in its own root.
Companies that run a shop, a blog or a help centre on a separate subdomain often forget this. They polish the main site's file, while the subdomain has no file at all or keeps the platform default. That default sometimes leaves useless areas open and sometimes closes important pages.
Location matters too. Robots.txt works only in the root directory; crawlers never look for a file inside a subfolder. A service on a non-standard port needs its own file as well. HTTP and HTTPS also count as separate origins, so make sure your redirect chain sends the robots.txt request to the right place.
A practical habit helps here. List every domain and subdomain you own, open each robots.txt address and note what it contains. This small inventory quickly reveals a forgotten staging subdomain in the index or a shop area that stays closed.
When you restructure a site, my website migration SEO checklist keeps these details from slipping.
Why do Crawl-delay and other unsupported lines do nothing?
Google supports only four fields in robots.txt: user-agent, allow, disallow and sitemap. It ignores lines such as Crawl-delay, noindex, nofollow and host. They can sit in the file, but they have no effect on Googlebot.
The problem is that teams rely on them. For example, a team that adds Crawl-delay when server load spikes may think the issue is solved, while Googlebot's crawl rate has not changed at all. Some other search engines may honour the line; Google does not.
If Googlebot strains your server, temporary 500, 503 or 429 responses tell Google to slow down. A lasting fix, however, means more server capacity, better caching and fewer useless URLs. For these technical foundations, my technical SEO tips are worth a look.
Common robots.txt errors compared with the correct approach
The table below puts every error from this guide side by side. Open your own file and match each line against it.
| Mistake | What happens | Correct approach |
|---|---|---|
| Blocking CSS and JS folders | Google cannot see the rendered page | Keep resource files open |
| Disallow: / left on the live site | Google stops crawling the whole site | Add it to the launch checklist |
| Using robots.txt to deindex | The URL can stay indexed without a snippet | Use noindex or X-Robots-Tag |
| Combining noindex and Disallow | Google cannot read the tag | Noindex first, block later if needed |
| Wrong letter case in a path | The rule misses the target folder | Match the real spelling on the server |
| Relative Sitemap address | The sitemap reference fails | Write a full URL with protocol |
| File above 500 KiB | Rules at the end get ignored | Write patterns, not single URLs |
| 5xx on the robots.txt URL | Crawling pauses or uses an old copy | Monitor the status code |
Use the table as an audit template. For each row, check whether your file contains a matching rule. Mark the rows that apply, then return to the relevant section of this guide and apply the fix. That way you review the whole file in one sitting.
Every row shares one trait: the error rarely produces a visible failure. The site works and pages load, yet crawling breaks quietly. That is why regular checks are worth more than a one-off fix.
Which tools help you find robots txt mistakes?
Start with Search Console. The robots.txt report under Settings shows when Google last fetched your file, which response it received and which warnings it found while parsing. The URL Inspection tool then tells you whether robots.txt blocks a single address.
If you are new to it, my Google Search Console guide explains the core reports. Still, Search Console only gives you Google's summary. To see what the crawler actually requested, you need the server logs.
- Server logs: show which response Googlebot got for robots.txt and whether it still requests blocked areas.
- Crawling tools: crawl your site while obeying robots.txt and export a list of blocked URLs.
- General SEO checks: the SEO checker gives you a quick page level review.
In the logs, look at two things in particular. First, check whether Googlebot requests robots.txt regularly and which status code it receives. Second, see how often it visits the areas you plan to block. Blocking an area Googlebot never visits gains you nothing; it only clutters the file.
My method is to read all three sources together. For instance, if Search Console reports a blocked page, I check in the logs when Googlebot last requested that path. That tells me roughly how long the error has been active.
What should you check before a launch or migration?
Many robots.txt errors appear on the day a new version goes live. The design changes, the CMS changes, the server changes; meanwhile the file either stays as it was or arrives as a copy of the staging version. So run a short checklist on launch day.
- Open the live robots.txt URL and confirm there is no Disallow: / line.
- Confirm the file returns 200 and does not sit behind a redirect chain.
- Check that the Sitemap line points to the new sitemap.
- Compare folder names in the new URL structure with the robots.txt lines letter by letter.
- Run a live URL Inspection test on key templates and look at the rendered view.
- Watch Googlebot requests in the server logs during the first week.
The list takes ten minutes, yet it prevents a visibility loss that can last weeks. To protect SEO across the whole project, combine it with the steps in my guide on protecting SEO during a website redesign.
How do you build a process that prevents robots txt mistakes?
A one-off fix is not enough, because robots.txt is a living file. Plugins add lines, teams change and new subdomains appear. The file needs an owner, and every change needs a record.
The process I recommend has three parts. First, keep the file in version control and write a short note for each change. Second, review the Search Console robots.txt report and the count of blocked pages every month. Finally, run the checklist above for every major release.
Measure the result of each change as well. After you add or remove a rule, watch the page indexing report and the crawl stats in Search Console for a few weeks. If the number of blocked URLs does not move the way you expected, go back to the rule. You cannot know whether a fix worked unless you measure it.
In the SEO consulting work my team and I deliver, robots.txt sits near the top of every technical audit. That is because one error in this file can cancel out the effect of content and link work. We make sure crawling is healthy first, and then we move on.
In short, robots txt mistakes rarely make noise, yet their effects last. Open your file today, compare it with the table above and make sure you can explain every line in one sentence. A line you cannot explain is probably one you do not need.




