Free tool

AI Bot Crawlability Checker

Which AI crawlers your site allows and which it blocks — GPTBot, ClaudeBot and the rest.

How it works

What is the AI Bot Crawlability Checker?

The AI Bot Crawlability Checker is a free tool that tests whether 14 AI crawlers, including GPTBot, ClaudeBot and PerplexityBot, can reach your website. It reads your robots.txt rules for each one, requests your page with each crawler's user-agent string, and reports every bot as allowed, blocked or unreachable, with the reason.

How to use the AI Bot Crawlability Checker

  1. Enter your domain, such as example.com. The robots.txt verdicts are for the site root.
  2. Press Check website. The tool reads your robots.txt and sends one request to your site as each of the 14 crawlers.
  3. Read the score and each crawler's card: allowed, blocked or error, the HTTP status it got, its response time, and whether robots.txt or a meta tag blocked it.
  4. Check Page Analysis for a noindex or nofollow tag and whether robots.txt and llms.txt exist, then work through the recommendations, which name the robots.txt line to change.

What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

They do different jobs, and blocking the wrong one has the opposite effect to the one you intended. GPTBot collects data for training OpenAI's models. OAI-SearchBot builds the index ChatGPT search uses. ChatGPT-User fetches a page when someone in a conversation asks for it. Blocking GPTBot while allowing OAI-SearchBot is a common position: no training, still citable.

Other operators split the same way. ClaudeBot is Anthropic's crawler, and anthropic-ai and Claude-Web are older names still found in robots.txt files. Google-Extended and Applebot-Extended are not crawlers at all. They are robots.txt tokens that control whether content Googlebot or Applebot already fetched may be used for Google's and Apple's AI models, and Google says Google-Extended has no effect on Search.

Do AI crawlers respect robots.txt?

The major declared crawlers do. OpenAI, Anthropic, Google, Apple, Microsoft and Common Crawl all document that their crawlers follow robots.txt, and several publish IP ranges you can verify requests against.

User-triggered fetches are treated differently. Perplexity documents Perplexity-User as acting for a person and not generally bound by robots.txt. Separately, Cloudflare reported in 2025 that it saw Perplexity fetch pages from sites that had blocked PerplexityBot, using undeclared user agents. Perplexity disputed that account.

How do I block AI crawlers?

Add a User-agent group for each crawler with Disallow: / in robots.txt. That works for crawlers that follow the file, but robots.txt is a request, not a control, and a user-agent string is free to fake.

To stop a crawler for certain, block it at your CDN or server with a rule that matches the user agent and checks the request comes from the operator's published IP ranges. Cloudflare offers managed rules that block AI crawlers.

For most sites the better answer is to block little or nothing. Being fetched is how your pages get cited, and a site that blocks every AI crawler has opted out of every AI answer that would have named it.

Why does a crawler show as blocked when robots.txt allows it?

Because this checker counts a crawler as blocked for any of three reasons: a Disallow rule in robots.txt, an HTTP error status of 400 or above, or a noindex in the page's robots meta tag. The card shows which one applied.

A 403 or similar status with an open robots.txt usually comes from a firewall, CDN or bot-protection rule that matches the crawler's user agent. Check those settings, then run the test again.

Questions

Should I block GPTBot?
Block GPTBot only if you do not want your content used to train OpenAI's models. GPTBot is the training crawler, so blocking it does not stop ChatGPT search from finding your pages, as long as OAI-SearchBot is still allowed.
Does blocking GPTBot remove my site from ChatGPT?
Blocking GPTBot alone does not remove your site from ChatGPT search, because OAI-SearchBot builds the search index. Block both and ChatGPT search can no longer fetch your pages to cite them.
What is Google-Extended?
Google-Extended is a robots.txt token, not a crawler, so no request ever arrives with that user agent. It controls whether content Googlebot has already fetched may be used for Google's Gemini models. Google says it does not affect Search ranking or inclusion.
Does Cloudflare block AI crawlers?
Cloudflare can block AI crawlers, and since July 2025 it has offered blocking them by default when a new domain is added. If your site is behind Cloudflare and AI crawlers show as blocked here with a 403 status, check its AI crawler settings before your robots.txt.
How do I check a crawler is who it claims to be?
Verify a crawler by matching the request IP against the operator's published ranges, or by a reverse DNS lookup that resolves forward to the same address. OpenAI and Google both publish their crawler IP ranges as JSON files for this.

AI Visibility Tools

More in AI Visibility Tools

AI Brand Visibility Checker

Ask the answer engines about your brand and see which name you, which name a competitor, and which name nobody.

Analyze visibility

ChatGPT Visibility Checker

How often ChatGPT names your brand or your site when asked the questions your buyers ask.

Check ChatGPT visibility

Grok Visibility Checker

How often Grok names your brand or your site when asked the questions your buyers ask.

Check Grok visibility

Ready when you are

Before your next buyer asks AI about you, ask AI about yourself

One check, four engines, about twenty seconds. We show you the answers before we ask for anything.