# No major AI crawler publicly meets all four of Cloudflare's conditions

Published: 2026-08-27T05:32:50.624Z · Source: Cloudflare Blog (https://blog.cloudflare.com/bot-preference-sync/)
Source date: 2026-08-27
Flags: wire analysis of published documentation — Cloudflare has published no compliance list, and a company could satisfy a condition privately
Entities: cloudflare, google, openai

When Cloudflare launched Bot Preference Sync on August 21 it set four conditions a crawler doing both search and training must meet to keep reaching sites that disallow training: honor the no-training preference, offer a way to opt out of AI summaries, report URL-level visibility into which pages were made available for training alongside search metrics, and publicly show that disallowing training does not hurt traditional search results. Cloudflare named no company that meets them.

Checking the published documentation of the three operators the rule binds — Google, OpenAI and Anthropic, each of which runs a separable search crawler and a training crawler — none satisfies all four.

Honor the no-training preference. Google and OpenAI both meet this and say so plainly. Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." OpenAI's crawler documentation states "a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot." Anthropic runs ClaudeBot and Claude-SearchBot as separate agents, but its published help page does not say that blocking the first preserves the second.

Offer a way to opt out of AI summaries. Google meets this, recently and incompletely: Search Console's Search generative AI control removes a site from AI Overviews, AI Mode and Discover's generative features, and is rolling out to "a subset of website owners." OpenAI offers no equivalent — its only lever is OAI-SearchBot, and opting out removes the site from ChatGPT's search answers altogether rather than from the summary alone. Anthropic describes no such control either: the levers on its content-removal page are noindex, robots.txt and password protection, each of which removes the page from Claude's search results rather than from the summary built on them.

Report URL-level training visibility alongside search metrics. Nobody does this. Google's Search Console reports how pages perform in AI features, which is appearance rather than training; Google-Extended is a robots.txt token with no reporting attached to it. OpenAI's documentation describes no publisher dashboard of any kind. Anthropic's describes none. This is the condition that decides the question, and it is unmet across the board.

Publicly show that disallowing training does not hurt search. Google and OpenAI both assert it. Google's is the sentence quoted above; OpenAI's documentation says the settings are independent and that a site can be in search while out of training. Anthropic's published pages carry no equivalent statement. Whether an assertion satisfies a condition Cloudflare worded as publicly demonstrate is unresolved, and Cloudflare has not said what would count — but Google and OpenAI have at least made the claim in public, which Anthropic has not.

The tally on published evidence: Google three of four, OpenAI two, Anthropic none documented. Which makes the rule, for now, a block rather than a standard — every dual-purpose crawler reaching a Cloudflare site that disallows training is doing so without publicly meeting the terms.

Two limits on this. Cloudflare has published no compliance list, so this is a reading of each company's own documentation against Cloudflare's stated conditions, not a verdict Cloudflare has issued. And absence from documentation is not proof of absence: an operator could satisfy a condition through a private arrangement with Cloudflare and never publish it. What can be said is what a publisher can check — and a publisher checking today finds no crawler that publicly clears the bar.

Why it matters: Cloudflare made continued access for search-and-training crawlers conditional on four public disclosures, and on the published evidence no major operator meets them — which decides whether the rule is a standard or a block.

## What this answers

**Do any AI crawlers meet Cloudflare's four conditions?**

Not on the public record. Checking Google, OpenAI and Anthropic against the four conditions on 2026-08-27, none satisfies all four: Google meets three, OpenAI two, Anthropic none it documents. Cloudflare has published no compliance list of its own.

**Which of Cloudflare's conditions do AI crawlers fail most?**

The third — URL-level visibility into which pages were made available for training, alongside search metrics. No major operator publishes it. Google's Search Console reports AI-feature performance but not training use, and OpenAI and Anthropic document no publisher reporting at all.

**Can a site block AI training without losing search traffic?**

With Google and OpenAI, yes, and both say so. Google states Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal", and OpenAI documents allowing OAI-SearchBot while disallowing GPTBot. Anthropic runs separate crawlers but does not document the guarantee.


Canonical: https://anythingengineoptimization.com/item/2026-08-27-no-major-ai-crawler-publicly-meets-all-four-of-cloudflare-s-conditions/
From Anything Engine Optimization (AEO Wire) — https://anythingengineoptimization.com/ · Standards: https://anythingengineoptimization.com/standards/
