EN
Build my website
Home›AI SEO Tools›Robots.txt Tester
Free SEO tool · no sign-up

Robots.txt tester: can search engines crawl this page?

Type a page address and pick a crawler. The tester reads the site's robots.txt, finds the rule that decides, and tells you whether the page may be crawled.

Free · results in seconds · we only read public pages

This tool checks one thing. The AI Website SEO Checker checks your whole website: 23 checks, every public page, three pages read in depth and the keywords you can win.

Check your website SEO with AI →

What robots.txt is for

Robots.txt is a plain text file at the root of a website (yourbusiness.com/robots.txt). Crawlers read it before anything else to learn which parts of the site they may visit. It is a request, not a lock: well-behaved crawlers like Googlebot and Bingbot follow it, while bad bots ignore it, so never rely on it to hide private pages.

How this tester decides

  1. It finds the right group. Robots.txt is split into groups, each starting with one or more User-agent lines. A crawler follows the group whose name matches it most closely (Googlebot-Image follows a Googlebot-Image group, then a Googlebot group, then the group for every crawler, *).
  2. It finds the rules that match the address. A rule matches when the address starts with its path. A star (*) stands for any text and a dollar sign ($) marks the end of the address.
  3. The most specific rule wins. Of all the matching rules, the one with the longest path decides. When an Allow and a Disallow are equally long, Allow wins. No matching rule means the page may be crawled.

These are the rules Google publishes and follows, so the answer here is the answer Googlebot gets.

Common mistakes

  • Disallow: / left over from the development site, which shuts the whole website out of search.
  • Blocking CSS and JavaScript folders, so Google cannot render the page the way visitors see it.
  • Using robots.txt to remove a page from Google. A blocked page can still appear, without a description, if other sites link to it. Use a noindex tag instead, and let the page be crawled so Google sees it.
  • A robots.txt that answers with a server error. Google then pauses crawling the whole site.

Robots.txt is also the best place to name your sitemap with a Sitemap line. Check it with the sitemap checker, and see how Google reads your whole site with the AI Website SEO Checker.

Questions

Robots.txt Tester: questions and answers

What does robots.txt do?

It tells crawlers which parts of a site they may visit. It does not remove pages from Google: a blocked page can still be listed if other sites link to it. To keep a page out of results, use a noindex tag and leave it crawlable.

Which rule wins when two rules match?

Google uses the most specific rule, the one with the longest path. When an Allow and a Disallow rule are equally long, Allow wins. This tester follows the same rules.

What happens when a site has no robots.txt?

If robots.txt returns "not found", crawlers may crawl everything. If the server answers with an error (a 5xx code), Google stops crawling the site for a while, so a broken robots.txt is worse than none.

Should I block AI crawlers like GPTBot?

It is a business choice. Blocking them keeps your content out of some AI training and answers, but can also keep your business out of the answers people read in AI assistants.