What Is Googlebot? How It Works and How to Verify It

What is Googlebot and what does it do in Google Search?
Googlebot is the generic name for the web crawlers that Google Search uses to find, download and process web pages. It is not one robot but two: a smartphone crawler and a desktop crawler. The pages it fetches then move on to rendering, indexing and ranking.
In this guide we share the framework our team uses in technical SEO audits. Also, we base the technical facts on Google Search Central documentation. So every claim here traces back to Google's own publications.
Think of Googlebot as a courier, not a judge. In practice, its job is to fetch your pages, not to decide how they rank. However, if the courier cannot reach your door, no ranking system can work with your content. That is why crawl access is the first requirement of SEO.
What is Googlebot's main job, and which types exist?
According to Google's documentation, Googlebot is the name for two crawler types. Googlebot Smartphone imitates a mobile user, while Googlebot Desktop imitates a desktop user. Both obey the same product token, so your robots.txt rules apply to both in the same way.
Google also runs other crawlers for other products. For example, there are separate variants for images, videos and news. In addition, Google Storebot handles shopping pages, and Google-InspectionTool powers the testing tools in Search Console.
This split matters in practice. Not every Google-labelled request in your server logs is the crawler that decides your rankings. Therefore, you should separate crawler types before you draw conclusions from log data.
- Smartphone crawler: the main crawler for mobile-first indexing.
- Desktop crawler: the crawler for the desktop version of a page.
- Image, Video and News variants: crawlers that collect content for those search types.
- Google Storebot: the crawler for shopping and product pages.
- Google-InspectionTool: the crawler behind Search Console and rich result tests.
What is Googlebot compared with Google's other crawlers?
Google sorts its crawlers into three groups. First, common crawlers feed the search index and always respect robots.txt rules. Second, special-case crawlers work under agreements with site owners. Finally, user-triggered fetchers make a single request when a person asks for it.
The table below puts the three groups side by side. The AdsBot example deserves attention. It checks your ad landing pages and can ignore the global user agent (*) rules when you have agreed to that. As a result, many site owners wonder why a robots.txt block does not stop it.
| Group | Examples | robots.txt behavior |
|---|---|---|
| Common crawlers | Googlebot, Storebot, Google-InspectionTool | Always respect rules for automatic crawls |
| Special-case crawlers | AdsBot, Google-Extended | Can follow agreed rules instead of the global ones |
| User-triggered fetchers | Google Site Verifier | Run once, when a user starts an action |
Source: Google's overview of crawlers and fetchers.
How does Googlebot discover new pages?
Googlebot finds pages in two main ways. First, it follows links on pages it already knows. Second, it reads the URLs in the sitemaps you submit. So your internal links and your sitemap directly shape how fast it discovers new content.
For example, imagine you publish a blog post and link to it from nowhere. Then Googlebot can only find it through the sitemap. Besides, being in a sitemap does not guarantee crawling or indexing. A sitemap is a suggestion list, not a command.
When we publish new content, our team follows this order:
- We link the post from at least one relevant service or category page.
- Then we add it to the sitemap and set lastmod to the real update date.
- Next, we confirm in Search Console's URL Inspection that the page is reachable.
- Finally, we check crawl and index status a few days later.
For the sitemap side, read our guide on the sitemap lastmod tag. To build a sitemap from scratch, our XML sitemap generator does the job.
What is the difference between crawling, rendering and indexing?
First, crawling means Googlebot downloads the page from your server. Second, rendering means a browser engine runs the page and its JavaScript. Finally, indexing means Google stores the page's content in its database. However, each step can fail on its own, and each failure needs a different fix.
The most common mistake we see is mixing these three up. For instance, if Search Console shows pages that were crawled but not indexed, crawling is not your problem. Googlebot fetched the page and chose not to index it. So you look for the answer in content quality and signals, not in server settings.
The reverse also happens. Pages that Google found but did not crawl yet usually point to a crawl-side issue. To tell these cases apart, use our guide on finding unindexed pages in Search Console.
How does Googlebot process JavaScript?
Google processes JavaScript pages in three phases: crawling, rendering and indexing. Googlebot first downloads the HTML and extracts the links. Then pages that return a 200 status code join a render queue. There, an evergreen version of Chromium runs the JavaScript and passes the resulting HTML to indexing.
The wait in the render queue can be a few seconds, or it can take longer. Therefore, if you load critical content only through JavaScript, indexing may lag. Server-side rendering or static generation removes that delay.
Two rules matter here:
- If robots.txt blocks a URL, Googlebot never requests it, so it does not render it. Also, content from blocked JavaScript files stays invisible.
- If the original HTML contains a noindex tag, Google may skip rendering and JavaScript execution.
So blocking CSS and JavaScript files in robots.txt can make a page look incomplete to Googlebot. We also covered the speed side in our article on how JavaScript affects site speed. Source: Google's JavaScript SEO basics.
How does mobile-first indexing change what Googlebot sees?
Mobile-first indexing means Google uses the mobile version of your pages for indexing and ranking. As a result, the main crawler visiting your site is usually Googlebot Smartphone. Content that exists on desktop but not on mobile can effectively vanish for Google.
However, the consequences are simple but harsh. A product table you hide on mobile, a category text you trim, or a schema block you skip will not reach the index. Likewise, resources blocked on the mobile version make it harder for Googlebot to see the page correctly.
In every audit our team opens the mobile version first and then compares it with desktop. For the design logic behind this, our guide on mobile-first design is a useful read.
What is Googlebot's crawl budget, and who needs to care?
Crawl budget is the number of URLs Googlebot can and wants to crawl on your site. It has two parts: crawl capacity and crawl demand. First, capacity reflects how much load your server can take. Second, demand reflects how interested Google is in your pages.
According to Google, this topic is not for everyone. It applies mainly to these sites:
- Large sites with more than 1 million unique pages whose content changes about once a week.
- Medium or larger sites with more than 10,000 unique pages whose content changes daily.
- Sites that see many URLs listed as discovered but not yet indexed in Search Console.
If you run a small business site, crawl budget is rarely your daily problem. Your priority is to build pages well and link them clearly. For a deeper case, see our article on why Googlebot crawls less.
What decides crawl capacity and crawl demand?
First, crawl capacity follows your server's health. If response times grow, or the server returns 5xx errors and 429 signals, Googlebot slows down. If the server answers quickly and steadily, capacity can rise.
Crawl demand depends on your perceived inventory, your popularity and how stale your content is. In other words, pages that change often and earn links get crawled more often. Pages that sit unchanged get crawled less.
| Factor | What raises crawling | What lowers crawling |
|---|---|---|
| Server health | Fast, steady responses | 5xx errors, 429 signals, slow responses |
| Content freshness | Real updates and honest lastmod dates | Unchanged pages that always look new |
| Site structure | Clear internal links and merged duplicates | Many parameter and duplicate URLs |
| Popularity | Links from other sites | Pages with no links anywhere |
In short, there is no magic switch for crawl budget. You speed up the server, cut duplicate URLs and push valuable pages forward. Source: Google's crawl budget guide.
Which mistakes waste crawl budget?
Googlebot's time is limited, so waste hurts. If you make it spend that time on worthless URLs, your important pages wait longer. These are the most common sources of waste we see in the field.
- Endless URL combinations from filter and sort parameters.
- Duplicate URLs that carry session IDs or tracking parameters.
- Removed pages that still return a 200 code, which creates soft 404 errors.
- Long redirect chains and redirect loops.
- Tag and archive pages that get links everywhere but add no value.
Google's advice is clear. Return 404 or 410 for pages you removed for good. Fix soft 404 errors and keep your sitemaps current. Also, if your server answers conditional requests with a 304 code, Googlebot does not download the same content again.
Be careful here. Do not try to fix this by editing robots.txt for a short time. According to Google, the crawl share you free up in robots.txt does not move to other pages unless your site is already at its capacity limit.
What is Googlebot's relationship with robots.txt?
The robots.txt file sits in the root of your site and tells crawlers which paths they may visit. Googlebot reads it before it requests a page, and it does not request blocked paths. So one wrong line can pull a large part of your site out of crawling.
The file manages crawling. It does not manage indexing, and people often miss that difference.
These are the basics our team checks:
- Does the file sit at the domain root under the right name?
- Does any rule block CSS or JavaScript files by accident?
- Also, does the file list your sitemap address?
- Did a blanket Disallow rule from a test environment leak into production?
If you start from zero, our robots.txt generator helps. For the logic, read what is robots.txt. For typical errors, see common robots.txt mistakes.
What is the difference between robots.txt and noindex?
Robots.txt controls crawling, while noindex controls indexing. Googlebot can only see a noindex tag if it is allowed to crawl and read the page. So blocking a page in robots.txt and marking it noindex at the same time is a contradictory combination.
According to Google, you use noindex to keep a page out of search. If you want to keep both users and crawlers out, you need password protection. For crawl budget, noindex is not a fix, because Google keeps requesting the page anyway.
| Goal | Right tool | What to watch |
|---|---|---|
| Page should not be crawled | robots.txt | The page can still enter the index through links |
| Page should not appear in search | noindex tag | The page must stay open to crawling |
| Nobody should access the page | Password protection | Blocks both users and crawlers |
For more detail, see our comparison of noindex and nofollow.
How much of a file does Googlebot read?
Googlebot processes the first 2 MB of supported file types and the first 64 MB of PDF files. Also, resources such as CSS and JavaScript follow the same limits. Google measures these limits on uncompressed data. For Google's crawlers and fetchers in general, the documentation states a 15 MB ceiling.
Most pages never come near these limits. However, pages with huge inline JavaScript, giant JSON blocks or embedded base64 images can push important content past the limit. Content beyond it does not reach the index.
So we suggest keeping key text, headings and schema data near the top of the HTML. Moving large scripts into separate files also helps speed and maintenance.
How do you verify that a visitor is the real Googlebot?
The Googlebot user agent header is often spoofed. Therefore, looking at that header alone is not enough. Google recommends two checks: a reverse DNS lookup and a match against its published IP ranges.
For a one-off check, follow the manual method:
- Take the requesting IP address from your server logs.
- Run a reverse DNS lookup on it and confirm the domain is googlebot.com, google.com or googleusercontent.com.
- Then run a forward DNS lookup on that domain name.
- Finally, confirm that the result matches the original IP address.
To check many requests at once, match them against the IP range files Google publishes. Use common-crawlers.json for common crawlers and special-crawlers.json for special-case crawlers. Source: Google's guide to verifying Googlebot.
How do you deal with fake Googlebot traffic?
Malicious bots often call themselves Googlebot so that nobody blocks them. This traffic can strain your server, copy your content or probe for weak spots. So you should never trust a request just because it says Googlebot.
So the approach we recommend is simple. First, automate the verification above. Then limit or block requests that fail it. Finally, never block requests that pass it.
Firewall rules hide a common trap. If you add rate limits and forget to exempt the real Googlebot IP ranges, you throttle crawling without noticing. As a result, your server stays safe while your indexing suffers. Watch the Crawl Stats report after every new rule.
How can you test what Googlebot sees on your site?
The most reliable method is the URL Inspection tool in Search Console. It shows how Googlebot fetched the page, the rendered HTML and any blocked resources. A live test also shows the current state of the page.
Then compare the page source with the DOM in developer tools. If critical content exists only in the DOM, Googlebot sees it after rendering. In that case, weigh the risk of delay.
Also remember that Googlebot does not wait forever the way you might in your browser. Content that loads very late, scripts that run very long, or sections that need a click can look empty. For example, text that appears only after someone clicks a show more button is probably invisible to Googlebot.
- URL Inspection: crawl, render and index status for a single page.
- Crawl Stats report: total requests, response time, host issues and file type split.
- Server logs: the raw data on which URLs were requested and how often.
- Page indexing report: which pages stay out and why.
If you are new to the interface, our guide on Google Search Console covers the first steps. For raw log analysis, try our log file analyzer.
How does Googlebot react to HTTP status codes?
The status code your server returns is the first and clearest message Googlebot gets about a page. The right code shapes crawling and indexing. However, a wrong one makes Google misread the page.
| Status code | Meaning | Result for you |
|---|---|---|
| 200 | The page loaded fine | It joins the render queue and can enter the index |
| 301 and 302 | The page moved | Googlebot follows the redirect; long chains raise crawl cost |
| 304 | Content did not change | Googlebot skips the download and saves resources |
| 404 and 410 | The page is gone | Googlebot drops it from the index over time |
| 429 and 5xx | The server is busy or broken | Googlebot slows its crawl rate |
During maintenance, showing a Maintenance page with a 200 code is a risky habit. Googlebot may treat it as real content. Instead, a suitable 5xx code is safer for temporary situations.
The same goes for redirects. Send visitors from the old URL to the new one in a single step. Chained redirects slow users and waste Googlebot's time.
What is Googlebot's difference from AdsBot in Google Ads?
They are not the same. AdsBot is one of Google's special-case crawlers, and it checks the quality of your ad landing pages. Googlebot, however, works for the search index. Their rules and robots.txt behavior differ.
According to Google, AdsBot can ignore global user agent (*) rules when the site owner agrees. So even if your robots.txt blocks all bots, AdsBot may still reach your ad page. For advertisers, that is an important detail.
If you have access problems on a landing page, run two separate checks. First, see whether Googlebot can crawl it. Second, see whether AdsBot can reach it. If you want both handled together, our Google Ads management service reviews both areas.
What common myths surround Googlebot?
Here are the myths we hear most in the field, with short answers. Also, each one creates extra technical work or a wrong expectation.
- "If I submit a sitemap, my page gets indexed." No. A sitemap only helps discovery.
- "I can remove a page from search with robots.txt." No. The file manages crawling, so the page can still enter the index through links.
- "Googlebot crawls my site every day." It depends. Crawl rate follows server health and content change.
- "If the user agent says Googlebot, it is Googlebot." No. The header is often spoofed, so verification is a must.
- "Crawl budget is every site's problem." No. According to Google, it mainly concerns large, fast-changing sites.
Clearing these myths early shortens our audits. When you ask the right question, you open the right report right away.
How should a Googlebot-friendly site structure look?
A Googlebot-friendly site is one where pages sit a few clicks away, links stay consistent and responses come fast. This structure makes the crawler's work easier, and it helps users too. So technical SEO and user experience often point in the same direction.
When we review a structure, we look at these principles:
- Important pages sit only a few clicks from the home page.
- Every page gets a link from at least one related page, so no orphan pages remain.
- Also, descriptive anchor text tells Googlebot what the target page covers.
- Category, tag and filter pages stay closed to crawling when they add no value.
- Finally, each piece of content has one canonical URL.
For example, e-commerce filters can create thousands of URLs. Because most of them show users nothing new, they add no value. At that point, a conscious decision about which combinations to crawl protects your crawl budget.
We covered this for big sites in our guide on category structure for large websites.
What practical Googlebot checklist can you use?
Instead, the list below sums up what our team checks on day one of a technical audit. We suggest reviewing all of it on a regular schedule, for example every month.
- Is robots.txt reachable, and does it hold an unexpected blanket block?
- Is the sitemap current, and does it list only the URLs you want indexed?
- Can internal links reach every important page?
- Does the mobile version carry the same main content as desktop?
- Are CSS and JavaScript files open to Googlebot?
- Are server response times steady and 5xx errors low?
- Do deleted pages return 404 or 410, and are there soft 404 errors?
- Does the Crawl Stats report show an unusual spike or drop?
The answer to what is Googlebot comes down to this list in the end. Googlebot must reach your site, see the page correctly and find your important URLs easily. Once you meet those three conditions, the rest of your SEO work pays off far more.
This list is a crawl-focused summary of the general frame in our guide on technical SEO tips.
In what order does our team audit Googlebot access?
Every project differs, but our order rarely changes. First, we check access. Then we check crawlability. Last, we check indexability. This way we look for the problem in the right layer and avoid needless changes.
- Access: server responses, redirects, firewall and CDN rules.
- Crawlability: robots.txt, internal links, sitemap and parameter URLs.
- Rendering: JavaScript dependence, blocked resources and mobile content parity.
- Indexing: noindex, canonical, duplicate content and page quality.
After that, we rank the findings by impact and effort. If you want support on the technical side, look at our SEO consulting service. To understand the scope, read what SEO consulting includes.




