Free tool

Robots.txt Checker

Fetch any site's robots.txt and test one URL against it: allowed, or blocked by which rule.

How it works

What is the Robots.txt Checker?

The Robots.txt Checker is a free tool that fetches a website's robots.txt file and tests one URL against its rules. It applies the matching Google uses, where the longest matching rule wins and Allow beats Disallow on a tie, and tells you whether a crawler following the User-agent: * rules may fetch that page.

How to use the Robots.txt Checker

  1. Paste a full page URL, such as https://example.com/blog/post, to test that page. A bare domain tests the homepage.
  2. Press View robots.txt. The tool fetches /robots.txt from the root of that site.
  3. Read the verdict: the URL is either blocked or not blocked for crawlers that follow the User-agent: * group.
  4. Read the file shown above the verdict for rules aimed at named bots, such as Googlebot or GPTBot, and for Sitemap: lines.

Does robots.txt stop a page from being indexed?

No. Disallow tells a crawler not to fetch a URL. It does not keep that URL out of the index. A disallowed page with links pointing at it can still appear in Google as a bare URL with no snippet, because Google knows the page exists and has been told not to read it.

A blocked page also can never have its noindex read, because reading it means fetching it. To remove a page from search, leave it crawlable and add a noindex tag. To stop it being fetched, use robots.txt, and accept that the URL may still show up.

How does a crawler decide which rule applies?

Crawlers do not read robots.txt top to bottom. For a given path, the rule with the longest matching pattern wins. When an Allow and a Disallow match at the same length, the Allow wins. So Disallow: /admin and Allow: /admin/public/ work together, and the order of lines in the file makes no difference.

* matches any run of characters and $ anchors a rule to the end of the path. Everything else is a plain prefix match. This checker applies those rules to the User-agent: * group, which is the one any crawler without its own group follows.

What else should a robots.txt file contain?

A Sitemap: line with an absolute URL. It is the one sitemap declaration every crawler can read, including AI crawlers that will never see your Search Console account. The checker shows the whole file, so you can see whether yours has one.

Google stops reading a robots.txt file after 500 KiB and ignores Crawl-delay. A Noindex: line has done nothing in Google since September 2019. If your file still has one, it is not removing anything.

Questions

How do I check a website's robots.txt file?
Open the site's root followed by /robots.txt, such as https://example.com/robots.txt, or paste the site into this checker. The checker also tests a page URL against the rules and tells you whether it is blocked for crawlers that follow the User-agent: * group.
Is a robots.txt file necessary?
No. If /robots.txt returns a 404, crawlers treat the whole site as open to crawling, which is the right default for most sites. What you lose is the Sitemap: line, so add a robots.txt file once you have a sitemap to declare.
Should I block AI crawlers in robots.txt?
You can. GPTBot, ClaudeBot, PerplexityBot and CCBot each use a named user-agent, and their operators say they follow robots.txt. Blocking them also keeps your pages out of what AI answer engines read and cite, which is the opposite of what most brands want.
What is the difference between robots.txt and llms.txt?
robots.txt is a standard (RFC 9309) that tells crawlers which URLs they may fetch. llms.txt is a proposed convention that lists the pages a site wants language models to read. It is not a standard, and it does not allow or block anything.
Is robots.txt case sensitive?
Partly. The paths in Allow and Disallow rules are case sensitive, so /Admin and /admin are different. Field names such as User-agent are not, and crawlers match their own name case-insensitively. The file itself must be named robots.txt in lower case.
Does robots.txt apply to subdomains?
No. robots.txt applies to one host and one protocol. blog.example.com needs its own file at its own root, and the file at example.com has no effect on it.

Technical SEO Tools

More in Technical SEO Tools

Opengraph Viewer

Paste a URL and see the card it produces when someone shares it — title, description, image.

View Opengraph

Open Graph Generator

Write the Open Graph tags, preview the card, copy the markup.

Generate Open Graph tags

Noindex Checker

Whether a page is set to noindex, checking the header and the meta tag both.

Check noindex

Ready when you are

Before your next buyer asks AI about you, ask AI about yourself

One check, four engines, about twenty seconds. We show you the answers before we ask for anything.