Search 93+ free tools… (e.g. json, vpn, password) ⌘K
Link Tools Dereferer Hide Referrer Link URL Shortener Affiliate Cloaker PayPal Links PayPal DonationPayPal Links Privacy Tools Password Generator Cloudflare Resolver My Referrer Torrent Tools Magnet → Torrent Torrent → Magnet Torrent Editor Pirate Bay Proxies Movierulz Proxies ExtraTorrent Proxies Dev Tools Base64 Encoder Hash Generator HTTP Headers Disposable Email Checker Company Blog About Us Contact Anonymize Free
Tutorials

Robots.txt Complete Guide: How to Control What Google Crawls

JAY
JAY
Author
May 27, 2026 · 3 min read · 285 views · 1 (5)
Robots.txt Complete Guide: How to Control What Google Crawls

Learn how robots.txt works, how to write rules for specific bots, and common mistakes that accidentally block Google from crawling your site.

Your robots.txt file is one of the most powerful — and most misunderstood — files on your website. Get it wrong and you can accidentally deindex your entire site. Get it right and you save crawl budget and keep private pages private.

What Is robots.txt?

robots.txt is a plain text file at your site root at yoursite.com/robots.txt. It tells search engine crawlers which pages to crawl or skip, following the Robots Exclusion Protocol — a voluntary standard followed by Googlebot, Bingbot, and DuckDuckBot.

Important: robots.txt controls crawling, not indexing. A blocked page can still appear in search results if other sites link to it. Use a noindex meta tag to prevent indexing.

How Crawl Budget Actually Works

Crawl budget is Google's practical limit on how many pages it will request from your site in a given period, based on your server's response speed and the site's overall popularity — slow response times directly shrink the budget, since Googlebot backs off to avoid overloading a struggling server. Blocking low-value pages (search result pages, filtered category duplicates, admin areas) in robots.txt frees up that budget for pages that actually matter, which is the real mechanism behind "save crawl budget," not just a vague SEO buzzword.

Why Crawl-Delay Doesn't Reliably Work

The Crawl-delay directive, which asks bots to wait a set number of seconds between requests, is honored by Bingbot and some other crawlers but is explicitly ignored by Googlebot, which manages its own crawl rate through Search Console settings instead. Relying on Crawl-delay to slow down Google specifically won't work — for that, the actual controls are in Search Console's crawl rate settings, or reducing server load so Googlebot's own automatic throttling backs off naturally.

Common robots.txt Rules

To allow everything: set User-agent to * and Allow to / and include your Sitemap URL. To block a directory: set Disallow to /admin/ or /private/. To block AI training bots like GPTBot: set User-agent to GPTBot and Disallow to / then repeat for other AI crawlers.

Common Mistakes to Avoid

Generate Your robots.txt Free

Use the Anonymiz Robots.txt Generator to build a valid file without memorizing syntax. Choose directories to block, add your sitemap URL, and download the ready-to-upload file in seconds.

🆔
UUID Generator

Generate v4 UUIDs instantly, in bulk, free.

Generate UUIDs →
# Tutorials
Share on X
Rate this article
★ 1 / 5 from 5 ratings
Your rating is stored anonymously. You can rate once per post.
JAY
Written by
JAYVerified site owner
Site Owner & Founder
JAY founded Anonymiz in 2013 and has personally built and maintained every one of its 100+ privacy and web utility tools since — from the referrer-stripping dereferer engine to the DNS leak and WebRTC leak testers. All technical infrastructure, tool logic, and site content are handled directly

Related Articles

We Tested 47 Pirate Bay Proxies Over 14 Days — Here’s What’s Actually Working
We Tested 47 Pirate Bay Proxies Over 14 Days — Here’s What’s Actually Working
Jul 4, 2026 · JAY
Certificate Key Matcher: How to Check If Your SSL Certificate Matches Its Private Key
Certificate Key Matcher: How to Check If Your SSL Certificate Matches Its Private Key
Jul 3, 2026 · JAY
How to Check HTTP Headers of Any Website
How to Check HTTP Headers of Any Website
Jun 4, 2026 · JAY
← Back to Blog
Done!