Sitemap Health Auditor
Review authored sitemap documents before publishing them. Assign each XML document its intended public URL to check loc fields, date rules, duplicate declarations, origin and directory scope. Add referenced XML documents yourself to inspect the supplied index graph. This tool never visits the listed pages or retrieves search indexing status.
Key features
- Manage explicitly located XML documents with optional previous snapshots
- Parse namespace prefixes and inheritance, character references, comments and CDATA while rejecting DTDs
- Check supported loc, lastmod, changefreq and priority rules with origin and directory scope
- Separate duplicate URLs, invalid or future dates and date reversals against supplied prior XML
- Inspect supplied index links, unavailable evidence, self references and cycles
- Export complete URL/document findings to CSV and all links and findings to JSON regardless of preview filters
How to use
- Enter a baseline origin and the calendar date used for future-date comparisons.
- Assign each document its published URL, then paste or read its UTF-8 XML.
- Add XML documents referenced by indexes; optionally provide previous XML for date comparisons.
- Run the local audit and inspect URL rows, findings and index links separately.
- Revise inputs, rerun the audit and save the complete CSV or JSON report.
Use cases
- Find a page accidentally repeated across separately generated sitemap files
- Inspect out-of-directory URLs in a sitemap published under a subfolder
- Compare XML snapshots to find lastmod values that moved backwards
- Review obsolete index references and cycles among supplied documents
Frequently asked questions
Does a row without findings mean a page is indexed?
No. Only the checked local rules passed. HTTP availability, robots, noindex, canonical declarations, ownership, crawling and indexing are not checked. Index targets that you did not supply remain unprovided evidence, never assumed healthy pages.
Which XML formats are supported?
Sitemaps0.9 urlset and sitemapindex. The bounded XML1.0 parser supports UTF-8 declarations, namespace prefixes and inheritance, standard entities, numeric references, comments and CDATA. DTD and external entity declarations are rejected. gzip, RSS, Atom, text sitemaps, extension semantics and complete XSD validation are outside scope; extensions receive an unchecked warning.
What does a reversed date compare against?
It compares lastmod for the same normalized URL against previous XML that you explicitly supplied for that document. Duplicate previous declarations or invalid prior dates are identified as insufficient evidence. Date-only values describe an entire day, so a time within that same day does not invent a reversal. An index file date is not compared as though it were a page modification date.
Which date and URL policies apply?
Supported lastmod forms are actual calendar YYYY-MM-DD or timezone-qualified YYYY-MM-DDThh:mm[:ss[.fraction]]Z/±hh:mm, up to9 fraction digits and offsets through±14:00. URLs must be absolute HTTP(S), shorter than2048 characters, without credentials, whitespace, bad percent escapes or fragments. Standard host/default-port/dot-segment normalization applies; path case, query order and trailing slashes stay distinct.
Are the work limits the same as search engine submission limits?
This tool accepts20 current documents,2MiB per XML,8MiB total current and previous text,10000 entries across both snapshots, XML depth32 and60000 tokens. These are local work bounds, separate from Sitemaps50000 URLs/50MiB. Naver also publishes a10MB feed condition. Submission, response speed and site ownership are not tested here.
What happens when I change inputs or load files?
Input changes and file loading invalidate the previous report. Published URLs must be assigned explicitly, not inferred from file names. Cancel or Escape stops work without accepting partial results. CSV and JSON downloads always include the complete report rather than the currently filtered page.
Privacy
XML, URLs and reports stay in the current page memory. Document URLs are never fetched, inputs are not uploaded or automatically saved. Download needed reports before closing the page.
Comments & questions