Correction 9/18/26 5:29 p.m.: This post originally implied publishers could not block Google from training on their content without being penalized in search, but Google-Extended allows for such control. The passage has been corrected.
Starting Sept. 15, Cloudflare will block AI training and agent crawlers by default on ad-supported pages for new domains and free-tier customers, while continuing to allow search crawlers unless site owners change their settings.
The change was outlined by Cloudflare executive Chema Alonso at Media Party Barcelona and detailed in The Media Stack’s account of the talk. Mixed-use crawlers will be treated according to their most restrictive function, meaning a bot that handles both search and AI training could be blocked if training access is denied.
Alonso said automated traffic has overtaken human traffic on Cloudflare’s network for the first time. Cloudflare says more than 20% of the web sits behind its infrastructure, giving it a broad view of how search crawlers, AI agents and training bots move across publisher sites.
The economic problem is familiar. Search engines traditionally indexed publisher pages and sent readers back, where advertising or subscriptions could generate revenue. AI systems can use the same reporting to answer a question without producing the click that generates revenue.
Cloudflare says 52% of crawler requests on its network are now tied to AI training, while more than a third come from mixed-use bots. The company wants crawlers classified as Search, Agent or Training so publishers can allow, block or charge each category separately. The controls build on Cloudflare’s earlier move to block AI training crawlers by default.
Publishers have long used robots.txt files to tell automated crawlers which parts of a site they may access. Separately, the Robots Exclusion Protocol standardizes how those instructions are written and interpreted. It governs crawler access rules, not payment terms or licensing.
Cloudflare is also changing how it thinks publishers should be paid. Its Pay Per Crawl model compensated a publisher when a bot fetched a page. The company now says payment should instead happen when publisher content appears in an AI answer.
Sponsored. AI Bot Access Check by oasy.ai: check any site's AI bot policy live: which AI crawlers and agents (GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended and more) its robots.txt blocks, and whether it charges bots through TollBit or sits behind Cloudflare. Free, no key: https://oasy.ai/api/aibots/r3r2a392?domain=nytimes.com
Two partners are already testing that approach. Ceramic.ai pays for non-paywalled news surfaced in its results, while You.com pays when premium paywalled content is used in a response.
Cloudflare says transparency is central to the model. Its argument is that clearer information about who is taking content and why can create scarcity, and that scarcity can give publishers more leverage to strike licensing agreements. Media Copilot has separately reported on publishers using blocking tools in AI licensing negotiations.
Cloudflare counts more than 50 publisher-AI agreements since 2023 but acknowledges that licensing is unlikely to replace all of the referral and advertising revenue publishers have lost.
For newsrooms, the immediate change is more control over which machines can access their work. Product and audience teams will have to decide whether to allow agent access, block training or seek payment for some uses.
Those decisions are unfolding alongside other compensation experiments. Separately, Google is testing payments based on publisher content value across some of its AI products.
Google also gives publishers a separate way to control whether their content is used for AI training. Since 2023, publishers have been able to use Google-Extended, a robots.txt control, to choose whether their content contributes to the training of future Gemini models. Google says using the control does not prevent a site from appearing in Search and is not used as a Search ranking signal, according to its public documentation.
The next test is whether other mixed-use crawlers follow Google’s lead and give publishers the same separation between search, AI training and other automated uses. If they do not, Cloudflare’s new defaults leave publishers with a simpler choice: permit the crawler as a whole or shut it out.







