HasData: Cloudflare's robots.txt change erased GPTBot bans at 118 sites
HasData’s AI Crawler Block Index found that 118 sites lost their Cloudflare-managed AI opt-out lines from robots.txt — a Content-Signal: ai-train=no tag, an EU copyright-directive rights reservation, and a GPTBot ban — when Cloudflare retired its “Managed Robots.txt” and “Block AI Bots” features on September 15, 2026, confirmed against archived Common Crawl copies of each file. Named among the affected: the Smithsonian, Ko-fi, Find a Grave, Vatican News, National Review, Bellingcat and Patreon. The replacement, Bot Preference Sync, requires the customer to review and confirm new rules before it writes them, so the file sits blank until they act — a step Cloudflare’s own announcements did not flag to existing customers. The same study, testing 10,894 domains in July and again in September, found a separate enforcement gap: 592 sites declared a GPTBot ban in robots.txt, and 234 of them (39.5%) still served GPTBot a live page when tested directly.
Why it matters: A robots.txt ban only holds if two separate things both work — the declaration and the enforcement behind it — and HasData's numbers show each failing independently, one through a product migration gap and one through sites that never blocked what they said they would.
The record: CloudflareOpenAI
Glossary: AI crawlers
Via PPC Land ↗ · HasData Blog ↗
Posted to the wire October 8, 2026. Edited by Joe Balewski.