Tools
Log File Analyzer for SEO
Use this log file analyzer to see what Googlebot and AI crawlers really fetch from your server and how often. Catch fake bots with the official IP lists, then find crawl budget waste and sitemap gaps. Your log is processed in your browser and never uploaded.
Result
How to use the Log File Analyzer for SEO
- Download your access logs
Export at least 7, ideally 30 days of access logs from your host, server or CDN. Raw files, rotated .gz archives and several files at once all work.
- Drop the files into the tool
Drag the files onto the upload area or choose them. The tool reads them in your browser piece by piece, so even very large logs stay on your device.
- Check format and verification
The tool detects the format on its own. Keep IP verification on: the tool then loads only the official crawler IP lists and marks every bot as verified, fake or unverified.
- Read the four tabs
Start with Overview. Then open Googlebot for status codes and crawl waste, AI crawlers for training, search and user triggered bots, and Sitemap for gaps.
- Export and act
Download any table as CSV, fix the redirects and errors that Googlebot keeps hitting, and repeat the analysis after each release.
How does the tool calculate its figures?
Every number comes from the lines in your log. The tool removes fake bots before it counts the Googlebot and AI figures.
bot requests / all requestsIP address inside the owner's published CIDR listbot name in the user agent, IP outside every list of that ownerthe owner publishes no IP list, or verification is off(3xx without 304 + 4xx + 5xx) / Googlebot requestsGooglebot requests with ? in the URL / Googlebot requestsA 304 answer does not count as waste: Google's crawl budget guide asks sites to support 304, so Googlebot can reuse its cached copy.
What does the tool conclude from single log lines?
The rows below come from the sample data. Each one shows a request, the result the tool gives and the next step.
| Log line (short) | Result | What to do |
|---|---|---|
| 66.249.66.1, Googlebot smartphone user agent | Verified: the IP sits in Google's common crawlers list | Nothing; count it as real crawl activity |
| 192.0.2.44, Googlebot user agent | Fake: no Google list contains the IP | Rate limit or block it at the firewall, not in robots.txt |
| 20.171.207.12, GPTBot user agent | Verified OpenAI training crawler | Allow or disallow GPTBot in robots.txt as your policy requires |
| 104.210.139.200, ChatGPT-User user agent | Verified user triggered fetch | Treat the page as one that people ask ChatGPT about |
| 198.51.100.5, YandexBot user agent | Unverified: Yandex publishes no IP list | Use a reverse DNS lookup if the traffic matters |
| Googlebot GET /old-page/ with status 301 | Crawl waste | Point internal links and the sitemap to the final URL |
The sample log is fictional. Visitor and impostor addresses come from documentation ranges (RFC 5737, RFC 3849), so none of them belongs to a real person.
Which AI crawlers and tokens should you know?
The table lists the AI crawlers that the tool recognises. Two rows are robots.txt tokens only: Google-Extended and Applebot-Extended never appear as a user agent in your log.
| Name | Owner | Category | robots.txt token | IP list |
|---|---|---|---|---|
| GPTBot | OpenAI | AI training | GPTBot | Yes |
| OAI-SearchBot | OpenAI | AI search | OAI-SearchBot | Yes |
| ChatGPT-User | OpenAI | User triggered | ChatGPT-User (may not apply) | Yes |
| ClaudeBot | Anthropic | AI training | ClaudeBot | Yes |
| Claude-SearchBot | Anthropic | AI search | Claude-SearchBot | Yes |
| Claude-User | Anthropic | User triggered | Claude-User | Yes |
| PerplexityBot | Perplexity | AI search | PerplexityBot | Yes |
| Perplexity-User | Perplexity | User triggered | Perplexity-User (generally ignored) | Yes |
| CCBot | Common Crawl | AI training | CCBot | Yes |
| Meta-ExternalAgent | Meta | AI training | meta-externalagent | No |
| Amazonbot | Amazon | AI training | Amazonbot | Web page only |
| Google-Extended | Token only | Google-Extended | Not a crawler | |
| Applebot-Extended | Apple | Token only | Applebot-Extended | Not a crawler |
Talha Aslan and the team checked names, roles and lists against the official pages of OpenAI, Anthropic, Perplexity, Google, Apple, Common Crawl, Meta and Amazon on September 29, 2026. ByteDance publishes no documentation for Bytespider, so the tool lists it under other bots.
What is a log file analyzer and why does it matter for SEO?
A log file analyzer reads the access log of your web server and shows which bots requested which URLs, when, and with what result. Search Console reports a sample of crawl activity; your server log, however, records every single request, including the ones that never show up in a Google report.
That makes the log the most honest source for technical SEO. For example, you can see whether Googlebot spends its visits on product pages or on filter URLs, redirects and errors. In addition, the log reveals which AI crawlers read your content, and whether a visitor that calls itself Googlebot really comes from Google.
Talha Aslan and the team use log files in every technical audit within our SEO consulting work. This tool brings the same first pass to your browser. It never uploads the file; the analysis runs on your device, so you can also check logs that you may not share.
Where do you get your server access logs?
Most hosts keep an access log for every site. On shared hosting with cPanel, open Metrics and then Raw Access; there you download the current log or the .gz archives. Turn the archive option on first, because otherwise cPanel deletes old data after each statistics run.
On your own server, Apache usually writes to /var/log/apache2/access.log and Nginx to /var/log/nginx/access.log. Older days end in .1 or .gz, and you can drop all of them at once. Behind a CDN, however, your server only sees the traffic that passes through. In that case, export the edge logs: Cloudflare Logpush writes JSON lines, and AWS stores ALB and CloudFront logs in S3.
The tool detects combined and common formats, IIS W3C, ALB, CloudFront, Cloudflare and Caddy on its own. Also, keep at least a week of data. A single day hides patterns such as the weekly crawl of your sitemap.
How does the log file analyzer spot fake Googlebots?
Anyone can write Googlebot into a user agent. Scrapers and vulnerability scanners do exactly that, because many sites let Google through without limits. So the user agent alone proves nothing.
Google describes two checks on its verification page. First, a reverse DNS lookup of the IP must return a host under googlebot.com, google.com or googleusercontent.com, and a forward lookup of that host must return the same IP. Second, for large logs, Google publishes JSON files with its IP ranges, and you match each address against them.
This log file analyzer uses the second method. Our server downloads the official lists of Google, Bing, OpenAI, Anthropic, Perplexity, Apple, DuckDuckGo, Common Crawl and Ahrefs once a day, and your browser matches every bot address against them, IPv4 and IPv6. A match means verified. A bot name from an address outside the lists of its owner means fake. Unverified applies to owners without a list, such as Yandex, Baidu or Meta; for those, use reverse DNS.
Which AI crawlers show up in your logs, and what do they want?
AI bots fall into three groups, and each group needs its own decision. Training crawlers such as GPTBot, ClaudeBot and CCBot collect pages for model training. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot build the index behind AI answers. User triggered agents such as ChatGPT-User, Claude-User and Perplexity-User fetch a page because a person asks a question right now.
The last group deserves special attention. A fetch by ChatGPT-User is not crawling; it is demand. Therefore the AI crawlers tab lists the pages that these agents open most often. OpenAI and Perplexity also state that user triggered requests may not follow robots.txt, since a person started them.
Two names never appear in a log. Google-Extended and Applebot-Extended are robots.txt tokens only, and Apple explains that Applebot-Extended does not crawl. To shape access, write your rules with our robots.txt generator, then read our guide to AI crawlers, robots.txt and llms.txt.
What counts as crawl budget waste in your logs?
According to Google's guide on crawl budget, crawl capacity and crawl demand decide how much Google crawls. The guide targets large sites; still, every site gains when Googlebot spends its visits on pages that matter. Your log shows where those visits go.
The tool counts redirects (3xx), client errors (4xx) and server errors (5xx) as waste and shows their share of all Googlebot requests. A 304 answer is the exception: it tells Google to reuse its cached copy, so it saves resources. Server errors matter for another reason too; according to Google, 5xx answers and 429 signals make Googlebot slow down.
In practice, look at the waste column of the per URL table. Old URLs that Googlebot still requests usually sit in internal links or in the sitemap, so fix those links first. Next, test redirect chains with our redirect checker and review parameter URLs, because filters and sort orders can multiply the same page.
How do you compare the log with your sitemap?
Your sitemap lists the URLs you want in the index, while the log shows the URLs that Googlebot actually requests. The Sitemap tab puts the two side by side and returns two lists.
The first list holds sitemap URLs that Googlebot never requested in the log period. These pages may be new, weakly linked or of low value in the eyes of Google. Add internal links to them and check their status in the URL Inspection tool, which our Search Console guide explains.
The second list holds HTML pages that answered with 200 and got crawled, although your sitemap leaves them out. Some belong in the sitemap; others, such as internal search results, should rather drop out of crawling. If your sitemap is outdated, build a clean one with our XML sitemap generator. The tool reads up to 5,000 sitemap URLs through our server and compares paths, so the host name does not matter.
Is it safe to analyze logs that contain IP addresses?
Access logs contain IP addresses, and under the GDPR an IP address can count as personal data. That is why this tool processes the file only in your browser. The log never reaches our server; the only outgoing requests fetch the official bot IP lists and, if you ask for it, your sitemap.
Still, handle log files with care. Share them only with people who need them, keep them no longer than your privacy policy states, and mask the last part of visitor addresses before you send a log to an agency. For instance, 203.0.113.25 becomes 203.0.113.0. Bot addresses from the official lists belong to company networks, so you can keep those intact for verification.
If you want an expert to read the results with you, our team combines log data with Search Console and a crawl of the site. As a result, you get a short list of fixes instead of another report.
Common mistakes in log file analysis
- ✕MistakeTrusting the user agent alone✓Do this insteadVerify the IP address against the official list or with reverse DNS; spoofed Googlebots are common.
- ✕MistakeCounting 304 answers as waste✓Do this insteadKeep 304 separate: it lets Googlebot reuse its cached copy and saves resources.
- ✕MistakeBlocking ChatGPT-User to leave ChatGPT✓Do this insteadDecide per bot: GPTBot controls training, OAI-SearchBot controls search, and user triggered fetches may ignore robots.txt.
- ✕MistakeJudging from one day of logs✓Do this insteadAnalyse at least 7, ideally 30 days, so weekly patterns and the effects of releases show up.
- ✕MistakeSending raw logs by email✓Do this insteadMask visitor IP addresses first and share only the period and the fields that someone needs.
Frequently Asked Questions
Do your logs point to a crawling problem?
Talha Aslan and the team review your logs, Search Console data and site structure together and steer crawl budget to the pages that matter.





