The default bot settings Cloudflare announced on July 1 took effect on September 15. On pages that display ads, crawlers used for AI training and for AI agents are now blocked by default, while crawling for search indexing stays allowed. The change applies to newly added domains and to free-plan sites whose owners never touched the settings. If either describes a site you run, it is worth opening the dashboard to check.
Search Gets Through, Training and Agents Do Not
Cloudflare has split AI-related bot behavior into three categories: Search, Agent, and Training. Search covers collecting and indexing content so a system can answer questions about it later. Agent covers automated visits made on a person's behalf to get something done in the moment. Training covers crawling that takes content to train or fine-tune a model.
The new defaults are not applied uniformly across the three. Only on pages that display ads are Training and Agent blocked, while Search remains allowed. Cloudflare's reasoning is that the presence of ads signals the owner intended that page for human readers, whereas search is the behavior that most naturally sends visitors back.
Sites without ads are not affected by the default change. For a portfolio or lead-generation site, being cited in an AI answer can work as free marketing, so the opposite of the new default may be the better choice. Settings can be changed at any time from the security settings in the dashboard.
Who Is Affected: New Domains and Free Plans
The scope is deliberately narrow. It covers new customers and new sites added by existing customers. It also covers existing free-plan users who had not changed their dashboard settings by September 15.
In other words, paid customers who already configured their own preferences see nothing change automatically. Free-plan sites are the group that needs attention. Plenty of sites were put behind Cloudflare years ago for speed and DDoS protection and never revisited. Those sites just had a content policy applied on their owners' behalf.
Why Mixed Crawlers Were Singled Out
The clear target of this change is the mixed-use crawler. When an operator runs separately named bots for search, agent use, and training, a site owner can allow one and refuse another. When all three purposes are collapsed into a single bot, agreeing to be discoverable in search also means agreeing to be used for training.
Matthew Prince, co-founder and CEO of Cloudflare, said the company hopes the proposed default changes encourage mixed-use crawlers to separate search from agent use and training. Cloudflare also points out that the largest search engine currently has access to roughly twice as much information as leading AI companies, precisely because staying discoverable there is hard to separate from being used for AI.
Multi-purpose crawlers are evaluated against all of their behaviors, and the most restrictive applicable rule wins. As a result, on sites where the owner has chosen to block training, multi-purpose crawlers such as Googlebot, Applebot, and BingBot are blocked as well. Given the effect on discoverability, that is a point worth pausing over before changing settings.
The Numbers Behind It
Among the figures Cloudflare cites, the one site operators feel directly is bandwidth. According to the company's data, more than 50 percent of crawl traffic from AI crawlers is spent re-fetching pages that have not changed. Publishers waste bandwidth, AI companies waste compute, and nobody gets a better answer out of it. Because Cloudflare can see what has actually changed, it is testing signals with leading AI companies that indicate whether a page is worth re-fetching, and plans to make them broadly available later this year.
On the commercial side, Cloudflare counts more than 50 major content licensing agreements signed between publishers and AI platforms over the past year. The company has also evolved last year's Pay Per Crawl into Pay Per Use, so publishers are paid when their content actually creates value rather than when a bot merely fetches it. Ceramic.ai and You.com are named as the first partners.
A New Signal in robots.txt
Cloudflare has also begun testing a new robots.txt signal called use. It adds an optional fourth field to the three in the original Content Signals specification, with three values: immediate (interact, but store and reuse nothing), reference (index, excerpt, and link back, which is the default), and full (summarize and reproduce).
These are stated preferences rather than blocks. Sites that already have Cloudflare's managed robots.txt enabled get use=reference added automatically. Cloudflare has started tracking content use across bots as well, and says bots that abuse the signal will lose Verified status. As things stand, a bot that reproduces content in full cannot hold Verified status.
Summary
Cloudflare's new defaults took effect on September 15, blocking AI training and agent crawlers by default on pages that display ads. The change applies to new domains and to free-plan users who had not adjusted their settings. Search crawling remains allowed, but mixed-use crawlers that do not separate search from training are blocked across all ad-supported pages. More than 20 percent of web domains sit behind Cloudflare, which gives a default setting something close to the weight of an industry standard. Whether AI companies split their bots by purpose will shape what the next phase looks like.
