Free Tool — AI Crawler Access Checker Can AI Actually
Read Your Site?

‌‌‌‌​​‌‌​‍‌​‌​‌‌‍‍‌​‍‌‌​​‌​​‌‌​‍ This checker reads robots.txt and compares live homepage responses. robots.txt is a polite request that AI crawlers may honour — it is not a lock, and access controls can apply separately. This tool requests your homepage using published crawler User-Agent strings and an ordinary browser User-Agent, then compares the responses. These requests come from our service, so results may differ from requests made by the vendors themselves. Free, no signup, no email.

Get This Fixed

    We make about seven requests to your site, read nothing but the homepage, and store none of it. Every result below is a live measurement — where something could not be measured, we say so rather than guess.

    HOW THIS WORKS

    robots.txt rules
    and live response checks

    1. We parse robots.txt properly. Not with a keyword search — to the actual standard (RFC 9309). Consecutive User-agent: lines are grouped so their shared rules apply to every agent in the group; the most specific group wins; within it the longest matching path wins, and on an exact tie Allow beats Disallow. An empty Disallow: means allow everything. We tell you which group matched and which line decided it, so you can check our working. We do that for twelve named agents.

    2. We compare requests using crawler User-Agent strings. Robots rules describe a site's preferences, while server access controls can refuse requests independently. Our service sends each published User-Agent string and reports the status received. A 403 means this request was refused. Server rules, firewall settings, and bot controls are possible causes; compare the results and check the site's configuration to investigate.

    3. We run a control, so the answer means something. We make the same request one more time with an ordinary browser User-Agent. If the browser request gets 200 and the GPTBot-labelled request gets 403, the service received different responses for those requests. That is a reason to investigate User-Agent-based filtering, while also checking the site's logs and configuration. If everything fails, we say the result is inconclusive rather than blaming AI blocking for a site that is simply down or refusing datacentre traffic. Without the control a 403 could mean anything, and we would rather report nothing than report something we cannot stand behind.

    4. Then the supporting checks. Whether the page is readable without JavaScript, whether your business is identified in structured data, whether an llms.txt exists — and whether a stray noindex is telling the world to ignore you. We found exactly that on a live law firm's site once: a dev-mirror config with Disallow: / had shipped to production and had been telling Google to go away for months. Nobody had looked, because nothing on the page looks broken.

    More than half the live business sites we have tested were quietly turning AI crawlers away — nearly always at the firewall, and nearly always without the owner knowing. We wrote up how that happens.

    THE CAST

    What These AI Crawlers Are,
    and When Blocking One Makes Sense

    If you have just found one of these names in your server logs, here is the plain-English version — what each one does, and how each crawler's purpose relates to your publishing and access preferences. The short rule: bots that put you into an answer are worth allowing; bots that only take content for model training are a business decision, and a legitimate one either way.

    ClaudeBot

    Anthropic

    Anthropic's crawler, and by a wide margin the AI bot people search for most after finding it in their logs. ClaudeBot collects web content that may be used to train Claude. Anthropic also runs Claude-User, which fetches a page because a person asked Claude about it right then, and Claude-SearchBot for search indexing.

    Blocking it: defensible for ClaudeBot if you do not want your content in training data. Blocking Claude-User is different — that one is a real person trying to read your page through Claude, and blocking it costs you a visitor.

    GPTBot

    OpenAI

    OpenAI's training crawler. Disallowing GPTBot tells OpenAI your content should not be used to train its foundation models. It is the single most commonly blocked AI user-agent on the web, and it is often blocked by accident — a managed "block AI bots" switch at the CDN catches it along with everything else.

    Blocking it: reasonable if training use is your concern. Just be clear that it is a separate decision from appearing in ChatGPT's answers — that is the next one.

    OAI-SearchBot

    OpenAI

    This is the one that decides whether you show up when someone asks ChatGPT a question. It is entirely separate from GPTBot: you can allow search and disallow training, and OpenAI documents that combination explicitly. Sites opted out of OAI-SearchBot do not appear in ChatGPT search answers.

    Blocking it: almost never a good idea. This is pure visibility — the AI equivalent of blocking Googlebot.

    PerplexityBot

    Perplexity

    Perplexity's search crawler. It exists to surface and link websites in Perplexity's answers, and per Perplexity's own documentation it is not used to collect content for training foundation models. Perplexity cites its sources prominently, so being indexed here tends to send real clicks.

    Blocking it: hard to justify. You lose the citation and the traffic that comes with it.

    Google-Extended

    Google

    Not a crawler at all — a robots.txt control token. Google already has your pages via Googlebot; Google-Extended only says whether that content may be used to train and ground Gemini. Disallowing it has no effect on your Google Search ranking, which is the detail most people get wrong.

    Blocking it: a clean, no-cost opt-out of AI training that leaves Search untouched. Plenty of publishers do exactly this.

    CCBot

    Common Crawl

    Common Crawl is a non-profit that publishes a free, open archive of the web. It does not build a product from your content — but nearly every AI lab has trained on its archive at some point, so in practice CCBot is an upstream source for many models at once.

    Blocking it: a broad, blunt way to stay out of many training sets. It also removes you from an open research archive, which some site owners value being in.

    THE BIGGER PICTURE

    Access Is Step One.
    Being Worth Quoting Is the Job

    We review crawler access, structured data, and the content search systems can reach. Then we track the visibility signals available to see how the work performs.

    1-800-765-2042

    While you are here: our free llms.txt generator builds a valid llms.txt from your real pages, and the GEO guide covers AI search visibility.