Robots.txt Generator: How It Works
robots.txt controls what crawlers are allowed to request. It does not control what appears in search results — a distinction that causes more accidental de-indexing and more unwanted indexing than any other file on a website.
Crawling is not indexing
This is the central point. Disallow prevents a crawler from fetching a URL. It does not remove that URL from the index, and a disallowed page linked from elsewhere can still appear in results — typically with no description, marked as blocked.
Worse, because the crawler cannot fetch the page, it cannot see a noindex tag on it. Blocking a page you want removed guarantees the removal instruction is never read.
| Goal | Correct method |
|---|---|
| Keep a page out of search results | noindex meta tag or header — and allow crawling |
| Save crawl budget on low-value URLs | Disallow in robots.txt |
| Protect private content | Authentication — never robots.txt |
| Handle duplicate URLs | Canonical tags |
robots.txt is a public file at a predictable address. Listing a sensitive directory in it advertises exactly what you were hoping to hide.
Syntax
User-agent: * Disallow: /admin/ Disallow: /cart Allow: /admin/public-guide.html User-agent: Googlebot Allow: / Sitemap: https://example.com/sitemap.xml
- Groups are matched by
User-agent; a crawler obeys the single most specific group that matches it, and ignores the rest — so a Googlebot group that omits your general rules means Googlebot never sees them. - Paths are case-sensitive and matched by prefix.
Disallow: /cartalso blocks/cart-abandoned. *matches any sequence,$anchors the end:Disallow: /*.pdf$.Allowcreates exceptions within a disallowed path.- The file must sit at the domain root and applies only to that exact host and protocol.
The catastrophic line
User-agent: * Disallow: /
This blocks the entire site. It is the standard configuration for a staging environment, and pushing it to production is one of the most common ways a site disappears from search. If organic traffic drops sharply and suddenly, check this file before anything else.
What to actually block
- Internal search results pages, which generate unlimited low-value URLs.
- Faceted navigation with parameter combinations.
- Cart, checkout and account pages.
- Scripts and admin endpoints with no user-facing content.
Do not block CSS or JavaScript. Google renders pages to evaluate them, and blocking the resources needed to render produces a page it judges broken.
Crawl-delay and other directives
Crawl-delay is not supported by Google, though Bing and others honour it. Google's crawl rate is managed through Search Console and adjusts automatically to your server's response times. Noindex in robots.txt was never officially supported and stopped working entirely in 2019 — anyone still using it is not blocking anything.
After you generate it
Place the file at the root, confirm it returns a 200 status, and test specific URLs with Search Console's robots.txt report. A 500 error on robots.txt causes Google to pause crawling the site entirely, which makes the file's availability as important as its contents.