thoughtasylumTOOLBOX Preview

Tools › Web

robots.txt Tester

Paste a robots.txt and test whether a crawler may fetch each address, with the rule that decides it. Mistakes are pointed out, and so are the crawlers, AI ones included, that are blocked from the whole site.

About this tool What it's for, how to use it and an example

What it's for

Test what a robots.txt file allows. Search engines and other crawlers read it before fetching a site’s pages, and a single misplaced rule can hide a whole site from search or let through crawlers you meant to block, such as the ones that gather text for AI.

For example, before publishing a new robots.txt that blocks AI crawlers, check here that Googlebot can still reach your pages and that GPTBot and ClaudeBot can’t.

How to use it

Paste the file, choose or type a crawler (the list includes the main search engines and AI crawlers; the choice is remembered), and add addresses or paths to test, one a line. Each is marked Allowed or Blocked with the rule that decided it, following the standard (RFC 9309): a crawler uses the group naming it, else *; the longest matching rule wins, and Allow wins a tie; * matches anything and $ marks the end.

Mistakes are listed above the results: rules before any User-agent, unknown fields, Noindex lines (which no longer work), and crawlers blocked from the whole site.

Example

Press Try an example, with the crawler set to Googlebot. The file blocks /admin/ but allows /admin/help, blocks addresses ending .pdf, and blocks GPTBot completely.

/admin/settings is blocked by Disallow: /admin/ (line 2), /admin/help/faq is allowed by Allow: /admin/help (line 3), /files/report.pdf is blocked but /files/report.pdf?v=2 is allowed (it doesn’t end in .pdf), and the page notes that gptbot is blocked from the whole site.

Good to know

robots.txt is a request, not a lock: well-behaved crawlers obey it, others don’t. Blocking a page doesn’t remove it from search results if other sites link to it; a noindex tag does that. A crawler that isn’t named exactly falls back to a group named with the start of its name (Googlebot-Image to Googlebot), then to *, as Google does; other crawlers may differ.

Private: this tool runs in your browser. Nothing you type, paste or choose leaves this page.

Saved you a few minutes? Say thanks with a coffee.

Something wrong with this tool, or missing from it? Report a bug or suggest a feature.

↑ ↓ move↵ openesc close