Skip to content
Toolbench

Robots.txt Tester

Paste a robots.txt and a URL, and find out which rule actually decides — longest match, not first match.

Runs entirely in your browser

Loading tool...

How to use

  1. Paste the robots.txt exactly as it is served.
  2. Enter the URL you are wondering about — full or just the path.
  3. Set the crawler; the wildcard group is only used when no specific group matches.
  4. Read which rule decided, and on which line.

Features

  • Longest-match precedence with Allow winning ties, as crawlers implement it.
  • The `*` and `$` wildcards, applied to the path and query.
  • Group selection by most specific user-agent — a specific group hides the wildcard one entirely.
  • Names the deciding rule and its line number.
  • Runs entirely in your browser; nothing is fetched.

Frequently asked questions

Why does my Disallow not block anything?
Usually because a longer Allow matches the same URL. `Disallow: /blog` and `Allow: /blog/public` together leave everything under /blog/public crawlable, because the longer path wins — regardless of which line comes first. The other common cause is a specific group: a crawler with its own `User-agent:` block never reads the wildcard rules at all.
Does robots.txt keep a page out of search results?
No. It asks well-behaved crawlers not to fetch the page. A blocked URL can still be indexed from links elsewhere and appear without a description — and because the crawler cannot fetch it, a `noindex` on the page itself is never seen. To keep something out of the index, allow crawling and use `noindex`, or require authentication.
Do all crawlers obey it?
The major search engines do. Scrapers and many AI crawlers do not, and there is no mechanism that forces them to — robots.txt is a request, not access control. Anything that must not be read needs authentication.

Build a robots.txt file, and see what it will actually do before you upload it.

Turn a list of URLs into a valid sitemap.xml — with the ampersands escaped and the off-site links called out.

Break a URL into its scheme, host, path, query parameters and fragment.

Robots.txt Tester — Is This URL Crawlable?