Cloudflare's September 15 default blocks Googlebot on ad-bearing pages
On September 15 Cloudflare sets new defaults for the three crawler classifications it introduced in July. In its own words: “Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.”
The sentence that decides what this costs is further down the same post. “Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.” Those three do search and training behind a single user-agent, so a block aimed at training reaches them too — on the pages the default covers.
Put together: a site that takes the default has Google’s crawler blocked on every page that displays an ad.
Who this reaches. Cloudflare says the new defaults apply to “all new domains onboarding to Cloudflare.” An existing domain keeps whatever it has now. Owners who do not want the defaults “can easily mark this in their Security settings any time leading up to September 15.” The opt-out is a setting, not a support ticket — but it closes on the day.
Three classifications, defined by behavior rather than by company. Search is “any behavior that collects or indexes your content, so it can answer questions about it later.” Agent is “automated behavior that is acting, usually in real time, on a person’s behalf.” Training is “a crawler taking your content to train or fine-tune a model.” The categories are clean; the crawlers are not, which is the whole difficulty. A crawler that does two of the three is governed by both answers.
This sits against a condition Cloudflare wrote itself. When it launched Bot Preference Sync on August 21, Cloudflare set four conditions a crawler doing both search and training must meet to keep reaching sites that disallow training. The fourth is that the operator publicly demonstrate that disallowing training does not hurt traditional search results.
On a Cloudflare site that takes the September 15 default, disallowing training does hurt traditional search results — because the multi-purpose crawler carrying both behaviors is blocked. The condition governs what crawler operators must show; the default governs what site owners get. They are different actors and it is not a contradiction in terms. But a publisher reading both posts is being told that blocking training should be survivable, by the same company whose default makes it not.
What is unresolved. Cloudflare’s post scopes the change to new domains. Several third-party summaries state the defaults also reach all existing free customers, which would enlarge the affected population substantially. Cloudflare has not addressed the difference in the post, and this item follows Cloudflare’s own wording. A publisher on a free plan should check Security settings before the 15th rather than assume either reading.
Also unresolved is what “pages that display ads” resolves to in practice — whether it is determined per page at request time, per domain at onboarding, or by a signal the owner sets. Cloudflare’s post does not say, and the distinction decides whether a single ad unit on one template pulls a whole site’s Googlebot access with it.
For anyone weighing the decision, the wire’s reading of the four conditions found that no major crawler publicly meets all of them, and that the one every operator fails is URL-level reporting of what was used for training. That finding is unchanged. What is new is that the cost of acting on it is now measurable in the crawler most publishers cannot afford to lose.
Why it matters: Cloudflare's own conditions for crawlers, set in August, require an operator to publicly show that disallowing training does not hurt search — and the default Cloudflare turns on September 15 makes exactly that harm, by blocking the multi-purpose crawlers that do both.
The record: CloudflareGoogle
Via Cloudflare Blog ↗ · Cloudflare Blog ↗
Posted to the wire August 30, 2026.