SiteTidy
Home/Tools/Website Crawler

Website Crawler

Connected Diagnostic

Website Crawler

Run a bounded same-origin crawl and see the evidence that matters: HTTP failures, missing metadata, canonical mismatches, noindex directives, duplicate titles/descriptions and where broken internal links were found.

Useful deliberately beats unlimited

This crawler is intentionally capped at 100 pages per run. It is designed for quick diagnosis of a section, launch, migration or small-to-medium site rather than pretending to replace a distributed enterprise crawler.

Start at the canonical HTTPS URL you want crawled. The crawler follows same-origin HTML links, removes fragments, limits response bodies and reports the source pages behind broken internal links. It does not execute JavaScript, so compare the results with a browser-rendered audit when a site depends heavily on client-side rendering.

What I would fix first

Prioritise failed internal URLs and accidental noindex signals before cosmetic metadata warnings. A duplicate title can be messy; a 404 in a purchase path or an important page excluded from indexing is materially worse.