Free tool

Which AI crawlers can
reach your site?

Enter a website or page address. HUB reads its robots.txt and shows which search, user-initiated and model-training crawlers are allowed, with the exact rule that decides.

Check a website

Free. HUB reads only the site’s robots.txt and tests the path you enter, or the homepage if you enter a domain. Nothing is saved.

Three kinds of crawler

Not every AI bot
does the same job

Blocking “AI bots” as one group can remove a store from AI search answers while leaving model training untouched, or the other way round. Providers publish separate crawler names so site owners can decide each case.

Search and answer engines

Googlebot, Bingbot, OAI-SearchBot, PerplexityBot, Claude-SearchBot and Applebot find pages to show and link in search and AI answers. Blocking them can make a store invisible there.

User-initiated fetches

ChatGPT-User, Claude-User and Perplexity-User visit a page because someone asked about it. Providers handle robots.txt differently for these; the table notes what each one documents.

Model training

GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot relate to training AI models. Google and Apple state that blocking their training tokens does not affect their search results.

Beyond robots.txt

CDNs, firewalls and bot-protection services can block crawlers that robots.txt allows. If a crawler is allowed here but your pages still do not appear, check those settings too.

Sources
Crawler purposes follow each provider’s documentation: OpenAI, Anthropic, Perplexity, Google and Apple. Rule matching follows Google’s robots.txt specification: the most specific path wins, and Allow wins a tie.

Questions

About the
crawler checker

Crawler access is one part of AI commerce readiness. The full check also covers product data, trust signals and performance.

Should I block AI training crawlers?

It is a business decision about how your content may be used. It is separate from search visibility: a site can allow search crawlers and block training crawlers, and many do.

Why does it read only robots.txt?

robots.txt is where crawler permissions are declared. Reading one file keeps the check fast and light on your server. Page-level noindex rules are covered by the readiness check.

What if robots.txt returns an error?

A missing robots.txt (404) means everything is allowed. A server error (5xx) is serious: Google stops crawling a site while its robots.txt keeps returning server errors.

Is anything saved?

No. HUB does not save the address or the result. Checks per visitor are limited with a short-lived hashed identifier to prevent abuse.

What does your software
need to do next?

A new build, a difficult codebase or a system that needs support. Let’s talk.

Discuss a project