Dashboard

Cloudflare Now Blocks AI Crawlers by Default: What Changed

Cloudflare started splitting AI training crawlers from search crawlers by default on September 15, 2026. Here is exactly what changed and how to check your own site in minutes.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
16 September 20261 min read

Cloudflare Now Blocks AI Crawlers by Default: What Changed

On September 15, 2026, Cloudflare started blocking so-called "mixed-use" AI crawlers by default, bots that combine search indexing with AI training data collection, unless a site owner explicitly allows them. The practical effect: a site can now stay discoverable in Google, ChatGPT, and Perplexity search results while refusing to let that same crawler train a model on its content, something the two uses used to be bundled together.

What a "mixed-use" crawler actually is

Most AI companies run one crawler that serves two purposes at once: indexing pages for their search or answer product, and hoovering up the same pages as training data for their next model. Cloudflare calls this a mixed-use crawler, and until this change, a site owner had one lever: allow the bot and get both uses, or block it and lose both, including any chance of being cited in an AI-generated answer.

What changed on September 15

Cloudflare's new default blocks mixed-use crawlers unless a site explicitly opts in, and introduces a separate setting called Disallow AI Training, named for the Disallow: directive it writes into a site's robots.txt. Turning it on tells a compliant crawler: keep indexing this site for search and AI answers, but do not use it to train a model. Apple, Google, and Microsoft have committed to honoring that distinction, which is the part that makes the setting more than a polite request most crawlers ignore.

The defaults apply automatically to new Cloudflare customers, newly created sites on existing accounts, and all free-tier customers. Paid customers who already had crawler settings configured keep what they had; nothing changes under them silently.

How to check whether this affects your site

  • If your site sits behind Cloudflare (free or paid), log into the dashboard and check the AI Crawl Control section under Security, it lists which crawlers are currently allowed, blocked, and whether Disallow AI Training is on.

  • If you manage your own robots.txt directly, on Cloudflare or elsewhere, look for a Disallow: directive scoped to AI training user agents. Its absence does not mean you are exposed by default anymore on Cloudflare, but it does mean you have not made an explicit choice either way.

  • If you are not sure whether your host runs on Cloudflare, check the response headers of your homepage for a cf-ray header, or run your domain through a DNS lookup and look for Cloudflare nameservers.

None of this requires touching code. It is a dashboard toggle or a few lines in a text file that already governs how search engines see your site.

Where this is heading

Cloudflare has said it plans to extend the same control to AI-generated summaries specifically, letting a site owner decide how much of their content can appear inside an AI Overview or a chatbot's synthesized answer, separately from indexing and training, by early next year. That would turn one binary allow-or-block decision into three separate ones: be indexed, be trained on, be summarized. Worth revisiting once that ships, since it changes the calculus again.

Why this matters if you publish content at all

If you are running a blog, a docs site, or a product built with AI (including one built with AI), the crawler question sits upstream of a bigger one: how do you get found by an AI answer engine in the first place? Blocking training crawlers protects your content from being absorbed into a model with no attribution and no traffic back to you. It does nothing for your visibility if it also blocks the same bot from indexing you for search or citing you in an answer, which is exactly the tradeoff this change is meant to remove. See how to get cited in Google AI Overviews for the other half of that picture.

This is one entry in a fast-moving news cycle. For a general approach to keeping up without losing a day to it, see how to keep up with AI news. For the other big story this week, Google's new voice models, see what Gemini 3.8 Live changes.

Frequently asked questions

Does this affect regular search engine crawlers like Googlebot?

No. This targets crawlers Cloudflare classifies as AI-related, the ones that scrape for model training, agent browsing, or AI search products. Traditional web search indexing is not touched by the new default.

What happens if I do nothing?

If you are a free-tier or newly onboarded Cloudflare customer, mixed-use AI crawlers are now blocked by default until you choose otherwise. If you were already a paid customer with crawler rules set before September 15, your existing configuration was left alone.

Does turning on Disallow AI Training hurt my chances of being cited by ChatGPT or an AI Overview?

Not by design. The setting is built specifically to separate the two: a crawler can still index and cite your content in search and AI answers while being told not to use it for training, as long as that crawler honors the Disallow: directive, which Apple, Google, and Microsoft have committed to doing.

Do I need a developer to make this change?

No. If you use Cloudflare, it is a dashboard setting. If you manage robots.txt yourself, it is a text edit any site owner can make directly.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.