Picking crawl rules and getting robots.txt as text
robots.txt has a fixed line order and spacing, so one stray character written by hand can make a whole group be ignored. Enter the User-agent, the paths to disallow or allow, and optionally a Crawl-delay and sitemap URLs, and you get text ready to upload. Paths go one per line, and duplicates are collapsed automatically.
The assembly order is the User-agent line, then Allow, Disallow and Crawl-delay, with sitemap lines appended after a blank line. Every path must start with /, and a line holding a space or control character is rejected with its line number. Sitemaps must be full URLs, so anything not starting with http or https is refused, and the User-agent name passes only allowed characters.
This page produces file text only. It does not predict crawl budget or indexing outcomes, and it does not analyze the robots.txt you are already serving. Keep in mind too that Crawl-delay is not honored by every crawler. Written as of October 2026.
Frequently asked questions
A single blank Disallow: line is emitted, which means nothing is blocked and is read as allowing everything. Disallow: / is the opposite and blocks the whole site, so keep the two apart.
No. robots.txt asks crawlers not to fetch a path; it is not a removal tool. An already indexed URL can linger in results on the strength of links alone, so removing it needs page-level directives or each search engine removal request.
A space has to be percent-encoded as %20, and non-ASCII paths are safer written in encoded form too. This page rejects any line containing a space and tells you which line it was.