Robots Policy Tester

Robots Policy Tester

Explain a supplied robots.txt policy using its groups and rule lines. Enter an origin, bot token and URL list to simulate the documented local profile.

No URL fetching, robots.txt changes or search-engine submission. Allow/disallow here describes only the supplied rule policy; it is not a crawl or indexing prediction.

Matching profile and limits

teck-tani-rep-core-v1: exact case-insensitive bot tokens; matching groups merge; wildcard groups are fallback only. Case-sensitive path + query. Longest normalized ASCII rule length wins, including * and terminal $; Allow wins ties. A nonterminal $ is literal. Only exact normalized /robots.txt without a query is implicitly allowed. Unknown directives do not split a group.

URL parsing uses the browser’s WHATWG URL normalization, including host case, IDN, default ports and dot segments. The displayed target is the actual comparison input. Percent-encoded unreserved characters decode; encoded reserved separators stay distinct. Raw * and $ in URLs compare with literal %2A and %24 rules. Fragments are excluded. Invalid escapes, controls, spaces, credentials and backslashes are rejected.

App bounds: 512 KiB UTF-8 robots text; 20,000 lines; 2,000 rules and groups; 200 URLs; 4,096-character input and normalized paths. Hard stop at 20 million matcher steps or 1 Mi characters of expanded winning-rule evidence. No truncation to an allowed result. These bounds differ from the RFC’s 500 KiB crawler-parser minimum.

Policy and URL inputs

HTTP/S scheme + hostname + optional port, with no path, query or credentials. Example: https://example.test

Letters, underscores and hyphens only. Presets are names to compare, not a claim to emulate those crawlers. Do not paste a full User-Agent header.

Blank text is an empty policy. Comments begin with #. BOM, CRLF, CR and LF are accepted; ignored lines are reported.

Up to 200 nonempty lines. Root-relative /paths use this origin. Absolute HTTP/S URLs from another origin are shown as outside scope. Plain relative and // URLs are rejected.

Enter a policy and URL list. Changing inputs clears old results.

Profile references

Reviewed 23 September 2026. These sources describe robots policies; the simulator’s exact contract and limits are stated above.

Comments & questions

Robots Policy Tester

Examine a proposed robots.txt policy before deploying it. This tool evaluates supplied text and URLs under a documented matching profile, showing the selected groups and highest-priority rule lines. It makes no requests to those URLs and does not determine whether a search engine will crawl or index a page.

Key features

  • Exact bot-token selection with merged matching groups
  • Case-sensitive path and query comparisons with explicit percent handling
  • Wildcard and terminal-anchor matching with rule-line evidence
  • Separate results for foreign origins and invalid input
  • Cancellable bounded work without partial success
  • Complete CSV and JSON exports independent of preview filters

How to use

  1. Enter the HTTP/S origin and one product token such as Bingbot or Yeti.
  2. Paste robots.txt and up to 200 absolute HTTP/S URLs or root-relative paths.
  3. Run the simulation and inspect parsing notices and the selected groups.
  4. Expand normalized targets and compare the strongest matching rule lines.
  5. Download the complete report; changing any input clears the previous result.

Use cases

  • Compare a shared wildcard policy with an explicitly named bot group.
  • Check how private-path exceptions and query-string rules interact.
  • Review a proposed rule change together with its exact input and decisions.

Frequently asked questions

Does this reproduce a search engine or check whether a URL is indexed?

No. The teck-tani-rep-core-v1 profile is a bounded local rule simulator. Bot names are input tokens, not verified crawler identities. Fetching robots.txt, redirects, HTTP failures, caching, engine aliases, crawl-delay, meta robots and actual crawling or indexing are outside its scope. Unknown directives are reported and ignored.

Which groups and rules are used?

Bot tokens use letters, hyphens and underscores. Matching is exact and case-insensitive; all matching groups are combined, with wildcard groups used only as a fallback. Rules compare from the start of a case-sensitive path plus query. The longest normalized ASCII pattern wins, counting star and terminal dollar operators; Allow wins equal-priority conflicts. Every equal-best rule line is retained.

How are URLs and escaped characters handled?

The browser URL parser normalizes hosts, ports and dot segments. Only the supplied origin is assessed. The comparison target includes a query, including an empty question mark, and excludes fragments. Escaped unreserved characters decode; escaped reserved separators remain distinct. Raw Unicode becomes UTF-8 percent sequences. Encoded star and dollar are literals. Exact normalized /robots.txt without a query is implicitly allowed.

What happens with invalid lines or limits?

Parseable rules are used while ignored lines are listed. This stricter profile rejects malformed rule paths, rather than guessing an engine’s recovery. Its own limits are 512 KiB UTF-8 text, 20,000 lines, 2,000 rules and groups, 200 URLs, and 4,096-character input and normalized paths. An execution stops at 20 million matching operations or a 1 Mi-character expanded winning-rule budget. These application bounds are separate from the RFC crawler parser minimum of 500 KiB; they are not a full conformance claim.

Are cancelled or incomplete runs reported as allowed?

No. Cancellation or a run-level limit discards the complete run. A foreign-origin row is outside scope and an invalid URL is unassessed, not allowed. Parsing notices remain visible because a rule decision only reflects the accepted lines. An empty policy has no applicable groups; it is not evidence about a live site.

What is included in downloaded reports?

Both exports include every URL row regardless of filter or page. JSON also keeps the exact supplied input, selected groups and all parsing notices. CSV uses UTF-8 BOM, CRLF and quoted fields; cells starting like spreadsheet formulas receive an apostrophe. The preview shows 25 rows per page, 20 strongest rules per row, 40 groups with 20 agent labels each, and 100 notices. Exports retain all collected details.

Privacy

The supplied robots text, origin, bot and URL list are processed in page memory. This tool does not fetch those URLs, send these inputs to a server or save them automatically. Downloads are created only on request and may contain private paths or query values. Leaving the page removes unsaved results.

Related Tools

Meta Tag GeneratorURL EncoderText CleanerWasm Module InspectorHreflang Matrix CheckerAST Query PlaygroundContainer Build GraphDependency Graph ExplorerSemver Range LabCron Schedule AuditorPatch Review WorkbenchSource Map ExplorerLocalization Catalog AuditorStructured Data ReviewerHTTP Archive AnalyzerWebhook Signature LabProtobuf Schema WorkbenchGraphQL Schema LabAvro Schema EvolutionLocal SQL WorkbenchSchema Form BuilderMesh Repair WorkbenchPipe Network LabRobot Arm Kinematics LabThermal Network LabBeam Response LabGear Train DesignerTolerance Stackup LabSensor Calibration FitPCB Stackup PlannerDigital Filter DesignerNetwork Reachability MapSun Shadow MapGPS Error SimulatorDigital Logic SimulatorAnalog Circuit LabMechanism Linkage LabAnalysis Mesh GeneratorOpenAPI Contract InspectorDatabase Migration PlannerDimensional Equation CheckerTruss Force LabBoolean Minimization LabControl Response LabQueueing Simulation LabGeofence Event SimulatorCoordinate Reference LabSurvey Traverse LabRaster Classification LabChoropleth Design LabMap Print ComposerRaster Reprojection LabElevation Contour MakerTerrain Viewshed LabWatershed DelineatorMap Tile PackagerText File Encoding WorkbenchFilesystem Portability AuditorSBOM License ExplorerFile Signature Auditornpm Lockfile Conflict ResolverSource Secret AuditorOffline Web Package BuilderCertificate Chain InspectorTorrent Metainfo InspectorChunked File PackagerEncrypted File VaultDuplicate File FinderArchive WorkbenchDesign Token ManagerSpacing Token DesignerResponsive Type SystemPackaging Dieline DesignerSVG Icon Sprite PackerFlex Layout PlaygroundCSS Grid PlaygroundRegex Equivalence LabMarkdown Repository AuditorLog Template MinerResponsive Layout AuditorEmail Template PreviewInternal Link GraphState Machine TesterPetri Net SimulatorGit History VisualizerCurl Request WorkbenchBinary Protocol DesignerHex File EditorBinary Patch WorkbenchFile Signature WorkbenchAPI Mock SandboxSchema Column MapperEvent Log SessionizerER Diagram DesignerTime Series Gap AuditorStratified Data SplitterData Lineage DesignerDecision Tree LabData Anonymization WorkbenchData Expectation RunnerJSON Schema ValidatorBasket Pattern AnalyzerSEO HTML AuditorAccessibility Structure AuditorSyndication Feed WorkbenchIndexNow Payload BuilderCrawl Log AnalyzerCSP Policy WorkbenchSearch Performance AnalyzerCSV Formula Risk AuditorCORS Response SimulatorCache Header LabCookie Policy InspectorWeb Vitals Trace LabSitemap Health AuditorBatch File RenamerFile Manifest VerifierFolder Space MapFolder Difference ReviewerRoute Order OptimizerGeoJSON Map EditorPolygon Overlay LabCartographic Label PlacerSpatial Table JoinGeoJSON Topology AuditorGPX Track AnalyzerTrack Privacy RedactorCSV Table JoinCSV Pivot WorkbenchScientific Data ProfilerTabular Cleaning WorkbenchRecord ReconciliationData Dictionary BuilderCanonical Graph AuditorRedirect Plan TesterHTTP response and ping reference testBrowser and System InformationJSON ↔ YAML ConverterXML ↔ JSON ConverterHTML FormatterJavaScript MinifierMock Data Generator.gitignore GeneratorLicense GeneratorUser-Agent ParserPassword Strength CheckerCode to ImageXML FormatterHTTP Status Code LookupMIME Type LookupJS & SQL String EscapeCSS Box Shadow GeneratorCSS Gradient GeneratorIndent ConverterNumber Base ConverterUnicode Escape ConverterUnicode InspectorJSON Structure DiffMarkdown Table GeneratorBase64 EncoderJSON FormatterSQL FormatterCron Expression GeneratorRegex TesterUUID GeneratorHash GeneratorTimestamp ConverterJWT DecoderHTML Entity ConverterMarkdown PreviewCSS MinifierJSON ↔ CSVCase ConverterImage to Base64
Explore all Dev Tools tools →Image/Media →Text/Convert →Life/Fun →