Crawl Log Analyzer
Read a local web server access-log excerpt to see which User-Agent strings claim to be crawlers, which paths returned 4xx/5xx responses, and how requests changed by UTC day. The analyzer never contacts your server or a search engine.
Key features
- Classify Googlebot, Bingbot, Yeti, NaverBot and other bot-like User-Agent claims without certifying identity
- Summarize redirects, 4xx, 5xx and error paths by claimed UA and request method
- Plot request and error counts by UTC day, with a claimed-UA filter
- Calculate mean and nearest-rank p95 only for lines carrying an explicit rt=seconds field
- Export JSON aggregates or error-path CSV without IP, referrer, full UA or query strings
How to use
- Remove personal data and secrets from the access log before pasting or choosing a file.
- Use Apache/nginx Common or Combined lines; append rt=seconds only if your custom log format actually records request time.
- Analyze the local input, then compare claimed-UA groups, UTC daily activity and 4xx/5xx paths.
- Filter a group or download the aggregate report; inspect paths again before sharing an export.
Use cases
- Find 404 paths requested by User-Agents that claim to be Bingbot or Yeti.
- Check whether crawler-like requests reached an updated route after a release.
- Compare server errors against claimed crawler traffic by day.
- Review recorded response times when a custom access log includes rt=seconds.
Frequently asked questions
Does a Googlebot or Bingbot label prove the requester was genuine?
No. Anyone can send that User-Agent string. The tool labels a claim only and always reports identityVerified=false. Google and Bing describe separate IP or DNS verification methods; this browser tool does not perform them.
Why are response-time cells sometimes blank?
The standard Common and Combined formats do not contain request duration. This tool computes milliseconds only from an explicit trailing rt=seconds value in a custom log line and never substitutes bytes, timestamps or network estimates.
Which log formats are accepted?
This is a strict profile of Apache/nginx Common or Combined: host, ident, user, bracketed CLF timestamp with numeric offset, quoted HTTP request, status and bytes; Combined adds quoted referrer and User-Agent. A trailing rt=seconds is optional. Custom JSON, proxy/vhost prefixes and extra fields need conversion first. An invalid nonblank line stops the report and shows its line number.
Does the result tell whether Bing or Naver indexed a page?
No. An access log shows responses to requests that reached this server. It does not prove crawler identity, discovery of every URL, rendering, canonical selection, search inclusion or rank. Use each engine's webmaster tools for those questions.
What data appears in exports?
IP addresses, referrers, full User-Agent values, raw lines and query strings are not exported. Method and path are included because they are needed to investigate errors; path segments themselves may be sensitive, so sanitize the input and inspect downloads before sharing.
Privacy
Input stays in this browser tab and is not sent to an application API or saved to browser storage. The tool does not run DNS or fetch any URL. Prepare a sanitized excerpt before using it; the output still contains error-path segments. Exports omit raw lines, IPs, referrers, User-Agent text and query strings.
Comments & questions