AI SEARCH & GEO

Is Cloudflare Blocking ChatGPT From Reading Your Website? Here’s How to Check

‌‌‌‌​​‌‌​‍‌​‌​‌‌‍‍‌​‍‌‌​​‌​​‌‌​‍We audited a client site whose robots.txt welcomed AI crawlers — while the firewall silently served every one of them a 403. ChatGPT literally could not read the site. Here’s the five-minute check, and the exact settings to fix.

AI Search & GEO
Blueprint-style technical drawing of robot crawlers stopped at a firewall gate in front of a website, with an inspection checklist alongside

This year we audited the website of an established business that wanted to show up in ChatGPT’s recommendations. Their robots.txt politely allowed every AI crawler. Their content was good. Our requests using several AI-crawler user-agent strings received 403 responses. The investigation identified a Cloudflare rule that conflicted with the owner’s intended policy. This article explains how to run those preliminary checks and what further evidence is needed to confirm actual crawler access.

Why this happens to well-configured sites

Cloudflare offers AI-bot blocking, and at various points it has been enabled by default for new zones and via one-click “Block AI bots” managed rules. It was pitched as protection against content scraping — reasonable for some sites. Review the active policies for search, training and agents separately. Access for OAI-SearchBot and ChatGPT-User can matter for different kinds of retrieval; a successful fetch does not guarantee inclusion in an answer.

Policies can overlap. In the account we audited, the new AI-bot policy panel said Allow — and an older, legacy “Block AI bots” rule scoped to all pages was quietly overriding it. Compare the effective rules with responses and security logs.

A preliminary request check

robots.txt is a request; the firewall is a decision. Test what the server does, not what your settings say. From any terminal (Mac: Terminal.app), run these against your own domain — each one pretends to be a specific AI crawler:

curl -s -o /dev/null -w "GPTBot: %{http_code}\n" \
  -A "GPTBot/1.0" "https://YOURDOMAIN.com/?cb=$(date +%s)"

curl -s -o /dev/null -w "OAI-SearchBot: %{http_code}\n" \
  -A "OAI-SearchBot/1.0" "https://YOURDOMAIN.com/?cb=$(date +%s)"

curl -s -o /dev/null -w "ChatGPT-User: %{http_code}\n" \
  -A "ChatGPT-User/1.0" "https://YOURDOMAIN.com/?cb=$(date +%s)"

curl -s -o /dev/null -w "ClaudeBot: %{http_code}\n" \
  -A "ClaudeBot/1.0" "https://YOURDOMAIN.com/?cb=$(date +%s)"

curl -s -o /dev/null -w "PerplexityBot: %{http_code}\n" \
  -A "PerplexityBot/1.0" "https://YOURDOMAIN.com/?cb=$(date +%s)"

Read the results plainly: 200 means this test request received a successful response; inspect the body to confirm it contains the page. 403 means this request was refused. Use server or firewall logs to identify the rule and verify actual crawler traffic. (The ?cb= bit busts caches so you’re seeing a live answer.) Test your homepage and a deep page or two; rules are sometimes scoped oddly. And run the same URLs with a normal browser user-agent as a control — if everything 403s, you have a different problem.

Fixing it in Cloudflare

If you saw 403s, inspect the Cloudflare security events for those requests before changing a rule. Review:

  • AI-bot policies. Find the controls under Security Settings and compare the search, training and agent policies with your intended access. Allow only the categories or crawlers you choose to admit.
  • Hunt for the legacy rule. If a “Block AI bots” managed rule or template exists — often scoped “Block on all pages” — it can override the newer panel. Set it to “Do not block” or remove it. This exact override was the root cause in our audit.
  • Check for scheduled behavior changes. Cloudflare has migrated these features more than once, and migration prompts sometimes default to re-blocking at a future date. If you’re offered a choice about what happens “when this feature is deprecated,” choose to keep AI crawlers allowed.
  • Leave your other protections alone. Allowing named AI crawlers does not mean turning off your WAF, bot-fight mode for malicious traffic, or rate limiting. This is a scalpel adjustment, not disarmament.

Then re-run the five commands and inspect the returned pages. In our client’s case, the test requests changed from 403 to 200 after the rule change. Confirm real crawler visits separately through logs and the provider’s verification method.

Should you maybe keep them blocked?

It’s a fair question with a real trade-off. If your business model is selling the content itself — journalism, stock imagery, paid research — blocking training crawlers like GPTBot is a defensible choice. But for a service business, your website exists to make you findable and chosen. The retrieval crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot) are how AI assistants read and cite you when a potential customer asks. Restricting retrieval crawlers can limit their access to your pages. Choose the policy that fits your publishing and discovery goals.

Access is step one, not the whole game

Once crawlers can read you, what they find has to be worth reading: entity schema that states plainly who and where you are, content that answers real questions, and consistent facts about your business across the web. That’s the larger discipline of generative engine optimization, and access is merely its front door. If you’d like the whole thing checked properly, an AI-visibility audit is part of our SEO & performance service — it takes us a day, and the crawler test above is literally its first step.

Frequently asked questions

Doesn't robots.txt control whether ChatGPT can read my site?

Only partly. Well-behaved crawlers honor robots.txt — but a firewall block happens before robots.txt matters. A site can allow every AI bot in robots.txt and still serve them all 403s at the edge. Always test the actual HTTP response.

Which AI crawlers are relevant to search and retrieval?

Check each provider’s current search and retrieval crawlers, which may differ from its training crawlers. Examples include OAI-SearchBot, ChatGPT-User, PerplexityBot and Claude-SearchBot. Use the provider’s documentation to choose and verify the policy.

Will allowing AI crawlers hurt my Google SEO?

No. AI-crawler policies and Google’s ranking systems are unrelated; Googlebot has its own rules. Allowing AI crawlers simply adds another set of readers to the audience your site already serves.

I'm not technical — can I still run the check?

If you can paste five commands into a terminal, yes — copy the block above and replace YOURDOMAIN. If you prefer help, ask your developer to run the checks or include them in an audit.

Have a website question?

We’ve built and maintained 450+ websites. Tell us what you’re working on and where you could use a hand.

Start a conversation