View and validate any website's robots.txt file. See all crawl rules, test whether a specific URL is blocked for Googlebot or other crawlers, and check sitemap declarations.
robots.txt is a plain-text file at the root of a domain (like anonymiz.com/robots.txt) that tells search engine crawlers which parts of a site they're allowed to visit. It's a voluntary convention, not a security mechanism — well-behaved bots like Googlebot respect it, but it can't actually block access to anything.
A robots.txt file is built from simple rules: `User-agent` specifies which bot the rule applies to (or `*` for all), `Disallow` blocks a path, and `Allow` can carve out an exception within a blocked path. A `Sitemap` line pointing to your sitemap.xml is also commonly included to help crawlers discover your pages faster.
The most common mistakes are disallowing a path that's actually meant to rank in search results, forgetting that robots.txt is publicly viewable (so it shouldn't be used to “hide” sensitive paths — anyone can read it), and syntax errors like a missing colon or incorrect path formatting that cause the intended rule to be silently ignored.