Web publishers using Cloudflare will soon be able to separately control an AI bot’s access based on what it does, allowing search indexing while blocking AI training or agent activity.
What’s new: Cloudflare – the content delivery network for nearly 20 percent of the internet – introduced new AI traffic controls which block training and agent bots by default while allowing search indexing crawlers. It also announced a monetization tool that would allow customers to charge specified visitors (whether humans or bots) for access to web pages, datasets, APIs, or Model Context Protocols.
How it works: Cloudflare customers can toggle crawlers’ access based on the bot’s use case, enabling or disabling access to crawlers for search, AI training, or AI agents. Crawlers serving multiple purposes will be subject to the most restrictive setting. For example, crawlers like Googlebot, Applebot, and Bingbot that both index content for search and collect training data for AI would be blocked if a publisher elects to block training crawlers. The controls will be applied beginning September 15.
- Cloudflare defines search bots as those that index content for search engines; publishers have historically welcomed these bots because they increase the likelihood that their sites’ content will appear in search results, which drives traffic. Agent bots are those acting on a user's behalf to perform tasks on a website (for example, fetching data or making purchases), and training bots collect data to train AI models.
- For each category, customers have three options: to block all pages, block only on pages with ads, or allow the bot. By default, pages displaying ads will block training and agent bots (or multi-purpose crawlers like those from Google, Apple, and Bing), but continue to allow search-only crawlers.
- Cloudflare also introduced BotBase, a public database of known bots and agents that classifies them by behavior and tracks how they use website content. Cloudflare’s “Verified” label indicates that a bot operator has been authenticated, is what it declares itself to be, obeys robots.txt instructions, and does not attempt to circumvent publisher restrictions. Publishers still control whether to allow that bot category, such as search or training, but this permits more granular restrictions, like blocking search, training, or agents from a particular company, while allowing others.
- Cloudflare's monetization system lets publishers charge AI agents for accessing online data on a per-request basis. When an AI agent requests a protected webpage or API, Cloudflare requires payment before forwarding the request to the publisher, handling payment verification at the network level. The tool has not yet been released, but customers can join the waitlist to use it.
Behind the news: AI bots account for a growing share of web traffic; Cloudflare says over half (57.5 percent) of all HTTP requests now come from automated systems rather than people. The value of search traffic has eroded enough that some publishers have even opted out of Google Search, in part because they do not want Google’s agents to bypass their display advertising or train on their content. Cloudflare's new framework attempts to preserve the benefits of search while giving publishers more control over other forms of AI crawling. It also rewards companies like OpenAI that separate their search and training crawlers, and punishes those, like Google, Apple, and Microsoft, that don’t.
Why it matters: Access to high-quality web data remains critical for developing and improving AI models, but publishers and partners like Cloudflare increasingly place limits on how that data can be collected. That friction has a technical and monetary cost. Such costs will likely have an outsized impact on smaller and newer AI developers, who lack the resources to secure agreements with publishers and face greater barriers to acquiring online knowledge that larger companies have already incorporated into their data sets.
We’re thinking: Cloudflare thinks a bargain can be struck between AI companies and online publishers, with the company playing the role of toll keeper. But it strikes at a basic principle held by most AI companies, which is that training on public web data should be fair use. It remains to be seen if Cloudflare has enough leverage to implement such a toll and whether its definition or most AI companies’ definition of “fairness” will prevail.