Cloudflare has launched new controls that allow website owners to refuse AI training while remaining fully indexed in search results, ending a long-standing trade-off that forced publishers to choose between discoverability and control over their content. The changes introduce a “Disallow AI Training” setting and an “Accountable” designation for AI crawlers that meet specific transparency and opt-out criteria.
Mixed-use crawlers, which collect content for both search indexing and AI training, now account for 36.6% of verified crawler traffic on Cloudflare’s network, making them the single largest category. While fewer than 1% of site owners block search crawlers, 17% restrict AI training, according to Cloudflare. Until now, refusing training often meant losing search visibility from crawlers that lacked separate controls for the two functions.
“This is how we make the Internet better: preserving the openness that makes search valuable while giving the people and businesses behind the web meaningful control over how their work is used,” said Matthew Prince, co-founder and chief executive of Cloudflare. “We look forward to continued collaborative engagement with companies like Apple, Google, and Microsoft as we work together to build a healthy ecosystem.”
Cloudflare is replacing its single “Block AI Bots” switch with three independent controls for search, AI training and AI agents. New websites on Cloudflare will see recommended settings tailored to their business model: sites that carry advertising will have search crawling enabled, AI training disallowed, and AI agents blocked on ad-carrying pages; all other sites will allow search, AI training and AI agents by default.
The company is also replacing “Managed Robots.txt” with “Bot Preference Sync”, which lets site owners set crawling preferences once and applies them across all supported crawlers automatically.
To earn the “Accountable” designation, crawler operators must meet or commit to four requirements: providing a clear way to opt out of AI training via robots.txt or a comparable standard; allowing opt-out of AI-generated search summaries; giving URL-level visibility into how content is used for search and training; and publicly confirming that opting out of training will not affect traditional search rankings.
Apple, Google and Microsoft have demonstrated that they meet these criteria, with Google offering Google-Extended and Apple offering Applebot-Extended controls that enable sites to opt out of training without opting out of search. Other leading AI companies operate separate crawlers for search and training, allowing Cloudflare to block their training crawlers without affecting search.
Cloudflare says AI summaries are the next focus, with a goal to let site owners control how much of their content is included in AI-generated answers from a single dashboard by early next year. The company is also participating in open standards work at the Internet Engineering Task Force on “ai-prefs”, a specification to allow any website to express AI access preferences in a standardised, portable way.


