2026-10-11 16:37 UTC

Cloudflare claims its new Disallow AI Training setting preserves search access while excluding training use through operator commitments and selective crawler blocking, allowing publishers to refuse training without sacrificing search discoverability.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumai-crawlers training-data web-publishingCloudflareBryan BeckerAppleGoogleMicrosoft

What is this?

Cloudflare announced Disallow AI Training and an “Accountable” crawler designation on September 15, 2026, presenting them as a way for publishers to refuse training use without blocking search crawling. Its blog says the setting publishes no-training preferences in robots.txt, permits Accountable mixed-use crawlers such as Applebot, Bingbot, and Googlebot, and blocks other training crawlers, including training-only crawlers from operators that separate search from training. For permitted mixed-use crawlers, protection depends on operators honoring those preferences rather than preventing content access; the snippets establish Cloudflare’s claims and operator commitments, not independently verified exclusion from training or guaranteed search visibility. Earlier July coverage described mixed-use crawlers being blocked under training restrictions; the launch distinguishes the new Disallow AI Training option from Block, which still stops those crawlers entirely.

Why it matters to Scott

Cloudflare’s separation of discovery access from training permission converges with Scott’s bounded licensed-knowledge access goal and gives him a concrete policy option to evaluate for his Cloudflare-backed publishing estate. It does not establish architectural enforcement for permitted mixed-use crawlers: operator commitments remain on the cooperation side of his Manners vs Physics distinction, and the supplied radar hits do not show this launch already tracked.
dev:concept.anti-bulk-licensed-knowledge-slicingip:concept.manners-vs-physicswork:project.cloudflaredev:project.publishradar:concept.cloudflareradar:concept.training-dataradar:concept.ai-search
queries asked of Scott's wikis
  • publisher content sovereignty training consent
  • search discoverability versus AI training access
  • robots.txt policy signals versus technical enforcement
  • public knowledge publishing AI reuse licensing
  • Cloudflare hosted sites crawler access controls

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 650h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 14:00⭐ origin echo-reconstructedCloudflare announces Disallow AI Training, which publishes training preferences while allowing Accountable mixed-use crawlers for search and
Bryan Becker, Cloudflare on blog (echo) · attributed from hn.story.49721435
—
09-16 02:25first on hacker news · published · +36.4hStay discoverable in search while disallowing AI training
djfergus
—
09-17 18:14first on r/ClaudeAI · published · +76.2hMy agent is not a crawler, I sent it
MihaiDinculescu
—
09-16 02:25amplified on hacker news 👑hn.story.49721435
djfergus
peak 87 · 51 comments · 96% of case engagement
09-17 18:14amplified on r/ClaudeAIreddit.post.1wj1ple
MihaiDinculescu
peak 1 · 10 comments · 4% of case engagement
09-16 03:20our radar first saw it · +37.4hdiscovery anchor: hn.story.49721435—
pace: p71 vs 1032 stories at the 336h mark (now 650h old) — ahead of astra-video-to-3d-world-demo (1.0x), behind engrim-local-cli-memory (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnStay discoverable in search while disallowing AI training
Retrieved article excerpt

Open article · Retrieved 2026-09-16T03:21:47.349890+00:00

[Network Services](https://blog.cloudflare.com/tag/network-services/)[Product News](https://blog.cloudflare.com/tag/product-news/)[Security](https://blog.cloudflare.com/tag/security/)

[AI](https://blog.cloudflare.com/tag/ai/)[AI Bots](https://blog.cloudflare.com/tag/ai-bots/)[Bot Management](https://blog.cloudflare.com/tag/bot-management/)[Network Services](https://blog.cloudflare.com/tag/network-services/)[Product News](https://blog.cloudflare.com/tag/product-news/)[Security](https://blog.cloudflare.com/tag/security/)

September 15, 2026

# Have it both ways: stay discoverable in search while disallowing AI training

Bryan Becker

[Bryan Becker](https://blog.cloudflare.com/author/bryan-becker/)

12 minute read

COPY URL

Without proper controls, website owners have long faced a difficult tradeoff: allow your content to be used for AI training, or risk losing discoverability in search. That tradeoff exists because some of the largest organizations on the Internet use mixed-use crawlers: a single crawler serving both search and AI training. Refuse one, and you refuse the other.

Today, Cloudflare is announcing a new [Disallow AI Training](https://blog.cloudflare.com/bot-preference-sync/) setting that lets you easily stay indexed for search while refusing to let that same crawler train on your content. Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting.

Mixed-use crawlers were the hard part of the training question. AI Summaries are next. A site-wide yes or no is too blunt: how much of your content appears in a summary matters as much as whether it appears at all. An opt-out for AI summaries is already one of the requirements we've set for mixed-use crawler operators. By early next year, our goal is to let you control how much of your content is included — set once on Cloudflare, rather than with each operator separately.

## Why asking isn’t enough

Most site owners want to be found: by humans, agents, and (good) bots. But a significant portion of the open Internet is funded by advertising, subscriptions, or direct relationships with visitors, and those models only pay when someone actually arrives.

Almost every site owner considers Search beneficial: less than 1% of Cloudflare sites choose to block Search bots. Training, however, is a different story: 17% of sites choose to enable some mechanism to block training. This is exactly why we decided site owners needed more granular controls, rather than a one-size-fits-all “Block AI.”

A robots.txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it.

A network can solve it, however: we publish the preference, identify who is crawling, classify why they are crawling, and block the ones that ignore it – then report what each operator actually does on [Radar](https://radar.cloudflare.com/ai-insights#ai-bot-transparency).

But blocking removes a crawler. It doesn't change how crawlers behave. The better outcome is operators that don't make you choose at all. So since July, we've been talking to them directly. The response has been encouraging: almost all agreed that site owners should have control and transparency into how their content is used, and reassurance that their choices will be respected. To help site owners understand that, we created a designation: Accountable.

The Accountable designation recognizes both capabilities available today and concrete commitments to deliver them. To qualify, a bot operator must meet or commit to meeting the following requirements:

1. A mechanism for site owners to opt out of AI training, through robots.txt or a similar standard.
2. A mechanism for site owners to opt out of AI summaries set with the operator directly, and next year through Cloudflare (see section below for more detail).
3. URL-level visibility into which pages were made available for training, along with metrics showing how content appeared in search.
4. Assurance that opting out of AI training will not affect traditional search results.

Apple, Google, and Microsoft all demonstrate that they meet the qualifications to be Accountable. Each combines capabilities available today with time-bound commitments for those still in development. The details of each of these companies’ crawlers are shared below.

## New security setting options

Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior. Three behaviors are available as controls:

- **Search** - crawling to build a search index.
- **Training** - crawling to train or fine-tune a model.
- **Agent** - user-directed agents visiting a page on behalf of a human, such as chat fetch bots and browser-use agents.

A mixed-use crawler is a single crawler doing both Search and Training. Without controls, that combination creates the tradeoff described above: site owners cannot refuse one use without refusing the other.

To avoid blocking Accountable mixed-use crawlers — the ones that don't force that tradeoff on website owners — we are introducing a new setting: Disallow AI Training. Disallow AI Training is named for the Disallow: directive it publishes in your robots.txt.

### “Block” setting now means something different

Block and “Block on pages with ads” previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability. Now that we have the new Disallow AI Training setting, Block and “Block on pages with ads” apply to *all* training crawlers, including mixed-use crawlers.

Training, Search, and Agent controls are applied at the domain level. With the addition of Disallow AI Training, the available settings are:

1. **Allow**: All crawlers are allowed, unless blocked by another setting or a WAF rule.
2. **Disallow AI Training**: Bot Preference Sync publishes the applicable no-training preference in robots.txt. Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search. Disallow AI Training is only available as a setting for Training, not Search or Agent.
3. **Block on pages with ads**: Crawlers, including mixed-use crawlers, are blocked only on pages detected to be serving an ad.
4. **Block**: All crawlers, including mixed-use crawlers, are blocked.

Disallow AI Training works by publishing a preference in robots.txt. An ads-only preference cannot be expressed that way: Cloudflare can detect which pages serve ads, but that list is too large and changes too frequently to enumerate in robots.txt. That's why there's no Disallow AI Training on pages with ads.

Agents do not create the same search-discoverability tradeoff as mixed-use crawlers, and the Internet does not yet have a well-established directive for expressing Disallow preferences to agents. For now, we’re not including a Disallow setting for Agents. As standards such as [ai-prefs](https://datatracker.ietf.org/wg/aipref/documents/) mature, we will revisit this approach.

## What changes on September 15?

We are making the following changes to Bot Management and AI Crawl Control:

1. Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training. To stop training and *keep* search, use Disallow AI Training.
2. “Block AI Bots” will be deprecated in favor of the more granular Search, Training, and Agent controls.
3. Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Customers who enabled Managed Robots.txt will migrate to the new system.
4. Disallow AI Training will become part of the recommended configuration for certain new domains.
5. Existing customers will have their preferences migrated to the new controls as described below.

### What you need to do

Nothing, in almost every case. Your current settings carry over on their own.

If you want mixed-use crawlers gone entirely, you now have to say so. Select Block. It will stop Applebot, Bingbot, and Googlebot from reaching your site — search included.

#### Existing domains that never used the Search/Training/Agent controls

Site owners that never configured the more granular controls will be migrated to the new settings based on their legacy Block AI Bots setting:

| (Legacy) “Block AI” setting | (New) Search setting | (New) Training setting | (New) Agent setting |
| --- | --- | --- | --- |
| Disabled (unselected) | Allow | Allow | Allow |
| Block | Allow | Disallow AI Training | Block on pages with ads |
| Block on pages with ads | Allow | Disallow AI Training | Block on pages with ads |

#### Existing domains that previously configured the Search/Training/Agent controls

For domains that previously configured the granular controls, we will preserve the practical effect of their selections under the new definitions. Previous Training selections of Block or Block on pages with ads will migrate to Disallow AI Training.

| Control | Legacy setting | New setting |
| --- | --- | --- |
| Search | Allow | Allow |
| Block | Block |
| Block on pages with ads | Block on pages with ads |
| Training | Allow | Allow |
| Block | **Disallow AI Training** |
| Block on pages with ads | **Disallow AI Training** |
| Agent | Allow | Allow |
| Block | Block |
| Block on pages with ads | Block on pages with ads |

### Recommendations for new domains

Beginning September 15, customers onboarding a new domain will be offered one of two preset configurations, depending on whether the site earns money from advertising. Ad revenue depends on a human actually seeing the page. Training replaces that visit with an answer; agents fetch the page with nobody there to see the ads. So the presets for ad-supported sites are more restrictive. You can change any of these settings during onboarding, or at any time afterward.

| Setting | Site does not monetize using ads | Site is monetized using ads |
| --- | --- | --- |
| Preference Sync | Enabled | Enabled |
| Search | Allow | Allow |
| Training | Allow | Disallow AI Training |
| Agent | Allow | Block on pages with ads |

*Recommended settings for new domains*

BLOG-3499 2.png

Screenshot of onboarding flow for a new domain, showing the recommended settings when “I monetize pages that serve ads” is selected.

## What does this mean for specific mixed-use crawlers?

Applebot, Bingbot, and Googlebot are Accountable. Apple, Google, and Microsoft are committed to the same principles of publisher choice and transparency. Under Disallow AI Training they can keep crawling your site for search. Selecting Block stops them entirely.

We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.

### Applebot

Applebot allows site owners to opt out of training by adding a Disallow rule to robots.txt for “Applebot-Extended”. Site owners can also currently express preferences for AI Summaries via their nosnippet [directive](https://support.apple.com/en-us/119829#:~:text=nosnippet%3A%20Applebot,products%20and%20services.) in the page HTML. Content can also be labeled as [paywalled content](https://support.apple.com/en-us/119829#:~:text=Marking%20paywalled%20content,the%20next%20section.) to exclude it from generative output. Applebot does not yet provide a tool for URL-level inspection. However, we have met with their team, and they have shared details of their in-progress solution for next year. Apple has also stated that disallowing training [does not impact search ranking](https://support.apple.com/en-us/119829#:~:text=Applebot%2DExtended%20and%20controlling%20data%20usage).

### Googlebot

Googlebot allows site owners to opt out of training by adding a Disallow rule 
djfergus8746
🟧 echo.blog ⭐Cloudflare announces Disallow AI Training, which publishes training preferences while allowing Accountable mixed-use crawlers for search andBryan Becker, Cloudflare——
🟠 redditMy agent is not a crawler, I sent it
ClaudeAI
MihaiDinculescu010

Interpretation history

Decision trace