Skip to content

Robots.txt builder and tester

Write a robots.txt with the rules you need, including AI crawlers, or paste yours and see which rule allows or blocks a URL. The file is checked against RFC 9309.

  • Runs in your browser
  • Nothing is uploaded
  • Free, no sign-up
  • method v1.0.0
  • Updated Oct 2026
  • Data checked: Oct 8, 2026

Input

What do you want to do

Adds a group for the training crawlers in the list below, each blocked from the whole site.

Rules by crawler

Each group applies to the crawlers named in it. A crawler with its own group ignores the * group.

  1. group 1

    Crawler names separated by commas, or * for every crawler.

    One path per line, starting with /. End with $ to match the end of the URL.

    One path per line. A longer Allow rule overrides a shorter block.

One full address per line. Optional.

Updates as you type · nothing leaves this page

Result

Nothing on the bench yet

Name a crawler and add the paths to block, or load an example to see what a result looks like.

Builds a robots.txt from the crawlers and paths you choose, or tests a robots.txt you paste: it tells you whether a URL is allowed for a crawler and shows the rule that decided. For founders and marketers who need search and AI crawlers to reach the right pages.

How to use it

Choose Build a file to write a robots.txt from scratch, or Test a URL to check one you already have.

  1. Build: name the crawlers of a group, then list the paths to block and the paths to allow. Use Block AI training crawlers to start from the training crawlers in the list.
  2. Test: paste the whole robots.txt, enter the full URL and the crawler's name.
  3. Read the verdict and the rule that decided it, or copy the generated file.
  4. Fix the findings, then copy or download the file and upload it to the root of your site as /robots.txt.

How Reavlo tests this

The file is read and matched in your browser, following RFC 9309 (the Robots Exclusion Protocol).

Groups. A group starts with one or more User-agent lines. A crawler uses the group that names it, compared without regard to case; groups naming the same crawler are combined. When no group names it, the * group applies; when there is none, everything is allowed.

Matching. A rule matches when its path matches the start of the URL path and query. * matches any run of characters and a final $ anchors the end. The most specific rule wins: the longest path, counted in bytes. When an Allow and a Disallow have the same length, Allow wins. An empty Disallow: blocks nothing. Percent-encoded characters are compared after their hex digits are written in upper case.

Checks. Lines that aren't field: value, rules before the first User-agent, paths without a leading /, relative sitemap addresses, directives outside the standard (such as Crawl-delay) and a * group that blocks the whole site are reported with their line numbers. Only the first 30 groups of 20 crawlers each are accepted in build mode; test mode reads up to 500 KiB, the size every crawler must parse.

Crawler list. The suggested crawler names come from the ai-user-agents dataset. Each name was read from its vendor's own documentation and carries the source and the date it was read.

Fixtures. The matcher is tested against the worked examples of the RFC and the examples in Google's documentation of how it reads robots.txt (path matching and rule precedence).

Limits. The tool follows the standard, and crawlers differ in the details. Google also reads a few directives outside RFC 9309, and some crawlers don't honor robots.txt at all. Requests a person triggers, such as ChatGPT-User and Perplexity-User, may not follow it by design. A page blocked here can still appear in results as a bare address if other sites link to it; to keep a page out of results, use a noindex rule instead.

Questions

Is my robots.txt uploaded?

No. Building and testing both run in your browser, and nothing you paste is sent to a server or stored. If Google Analytics is enabled on this site, it only records that the tool ran, never your file or your URLs.

Which rule wins when Allow and Disallow both match?

The most specific one, which is the longest matching path. If they are the same length, Allow wins. RFC 9309 defines this, and Google documents the same order.

Does Disallow remove a page from Google?

No. It stops crawling, not indexing. A blocked page can still be listed by its address if other pages link to it. To keep a page out of results, let it be crawled and add a noindex rule.

How do I block AI crawlers but keep Google Search?

Give each AI crawler its own group with Disallow: /, and leave Googlebot out of it. A crawler that has its own group ignores the * group, so the rest of your rules keep applying to everyone else. The crawler list marks which tokens are for training, search or user requests.

Where do I put the file?

At the root of the host it applies to, for example https://example.com/robots.txt. Each subdomain and each protocol needs its own file.