Robots Policy Tester
Examine a proposed robots.txt policy before deploying it. This tool evaluates supplied text and URLs under a documented matching profile, showing the selected groups and highest-priority rule lines. It makes no requests to those URLs and does not determine whether a search engine will crawl or index a page.
Key features
- Exact bot-token selection with merged matching groups
- Case-sensitive path and query comparisons with explicit percent handling
- Wildcard and terminal-anchor matching with rule-line evidence
- Separate results for foreign origins and invalid input
- Cancellable bounded work without partial success
- Complete CSV and JSON exports independent of preview filters
How to use
- Enter the HTTP/S origin and one product token such as Bingbot or Yeti.
- Paste robots.txt and up to 200 absolute HTTP/S URLs or root-relative paths.
- Run the simulation and inspect parsing notices and the selected groups.
- Expand normalized targets and compare the strongest matching rule lines.
- Download the complete report; changing any input clears the previous result.
Use cases
- Compare a shared wildcard policy with an explicitly named bot group.
- Check how private-path exceptions and query-string rules interact.
- Review a proposed rule change together with its exact input and decisions.
Frequently asked questions
Does this reproduce a search engine or check whether a URL is indexed?
No. The teck-tani-rep-core-v1 profile is a bounded local rule simulator. Bot names are input tokens, not verified crawler identities. Fetching robots.txt, redirects, HTTP failures, caching, engine aliases, crawl-delay, meta robots and actual crawling or indexing are outside its scope. Unknown directives are reported and ignored.
Which groups and rules are used?
Bot tokens use letters, hyphens and underscores. Matching is exact and case-insensitive; all matching groups are combined, with wildcard groups used only as a fallback. Rules compare from the start of a case-sensitive path plus query. The longest normalized ASCII pattern wins, counting star and terminal dollar operators; Allow wins equal-priority conflicts. Every equal-best rule line is retained.
How are URLs and escaped characters handled?
The browser URL parser normalizes hosts, ports and dot segments. Only the supplied origin is assessed. The comparison target includes a query, including an empty question mark, and excludes fragments. Escaped unreserved characters decode; escaped reserved separators remain distinct. Raw Unicode becomes UTF-8 percent sequences. Encoded star and dollar are literals. Exact normalized /robots.txt without a query is implicitly allowed.
What happens with invalid lines or limits?
Parseable rules are used while ignored lines are listed. This stricter profile rejects malformed rule paths, rather than guessing an engine’s recovery. Its own limits are 512 KiB UTF-8 text, 20,000 lines, 2,000 rules and groups, 200 URLs, and 4,096-character input and normalized paths. An execution stops at 20 million matching operations or a 1 Mi-character expanded winning-rule budget. These application bounds are separate from the RFC crawler parser minimum of 500 KiB; they are not a full conformance claim.
Are cancelled or incomplete runs reported as allowed?
No. Cancellation or a run-level limit discards the complete run. A foreign-origin row is outside scope and an invalid URL is unassessed, not allowed. Parsing notices remain visible because a rule decision only reflects the accepted lines. An empty policy has no applicable groups; it is not evidence about a live site.
What is included in downloaded reports?
Both exports include every URL row regardless of filter or page. JSON also keeps the exact supplied input, selected groups and all parsing notices. CSV uses UTF-8 BOM, CRLF and quoted fields; cells starting like spreadsheet formulas receive an apostrophe. The preview shows 25 rows per page, 20 strongest rules per row, 40 groups with 20 agent labels each, and 100 notices. Exports retain all collected details.
Privacy
The supplied robots text, origin, bot and URL list are processed in page memory. This tool does not fetch those URLs, send these inputs to a server or save them automatically. Downloads are created only on request and may contain private paths or query values. Leaving the page removes unsaved results.
Comments & questions