Anthropic has resumed charging for safeguard-blocked requests in low-false-positive categories (biology, distillation attacks, frontier LLM development) as a stated defense layer against coordinated attacks, making blocked calls a real line item in agent API economics and testing whether its <0.1% false-positive tuning holds under billing pressure.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: highanthropic inference-economics agentic-securityAnthropic
Surfaced 2026-09-29T16:56:19Z β "Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit β The velocity spike was the launch-day Reddit thread's bump peaking and digesting β the thread grew to ~62 pts/23 comments and is already easing (65β62, comments flat), i.e. intra-thread amplification of FP reports the case already holds, not a third community or outlet. No Anthropic response, no CVP extension, no pricing disclosure; the case's meaning shifts from 'unfolding FP controversy' to 'established policy with a dormant two-community FP grievance awaiting Anthropic's next move', so heat cools to low despite the 78th percentile reading β that percentile reflects a small cooling cohort, not live momentum.
What is this?
Anthropic announced on September 24, 2026 (ClaudeDevs post, echoed in its platform release notes and refusal/fallback docs) that it has resumed charging for API requests its safeguards block before Claude produces any output β limited to three categories it tunes to a <0.1% false-positive rate: biology, distillation attacks, and frontier LLM development (documented as the `bio`, `frontier_llm`, and `reasoning_extraction` stop_details categories; other pre-output refusals stay unbilled but count against rate limits). The company frames billing as one layer of defense against coordinated attacks, citing a September 2026 threat-intelligence report on large-scale illicit distillation campaigns (the supplied snippet names Alibaba's, peaking near 3M exchanges/day from 3,500+ fraudulent accounts), and reports 99.7% of Claude Code/claude.ai/Cowork accounts triggered none of the newly billable blocks in testing. Community counter-evidence on the false-positive claim β benign prompts tripping bio/cyber flags β remains first-hand but anecdotal (HN, r/ClaudeAI launch-day thread), and the supplied coverage does not confirm the case's specific attribution to DeepSeek/Moonshot/MiniMax, any pricing figures, or whether CVP-gated cyber blocks fall inside billed categories.
Why it matters to Scott
Anthropic has independently arrived at the full-cost-per-transaction accounting Scott argues in his Agent Token Manifesto (dev:project.llmreport) β failure paths are now literally on the invoice β and the new cross-community false-positive reports plus the context-carried [cyber]-trigger finding are a dated receipt for Guardrail Illusion: a probabilistic safeguard promoted to a billing boundary, failing in exactly the unreliable-permission-boundary way that concept names. It also bears on his own stack operationally: every project routes through the LiteLLM proxy, so stop_details refusal categories now carry monetary cost, and a billed block plus blind retry in an autonomous harness becomes a burn loop until stop_details-aware resumption/autonomy-budget handling exists.
ip:concept.guardrail-illusiondev:project.llmreportdev:technology.litellmdev:project.askradar:fable-5-safeguard-fallbacksradar:concept.model-safetyradar:concept.inference-economicsradar:concept.anthropicradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
- AI unit economics full cost per transaction failure paths
- agent retry policy on blocked refusals autonomy budget
- guardrail false positives safety classifier reliability cost
- distillation attacks defense economics making probing expensive
- LiteLLM harness handling stop_details refusal categories
- guardrail illusion layered safeguards model-level vs around-model
Measured heat
now 0 pts/hpeak 36 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 434h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p72 vs 1032 stories at the 336h mark (now 434h old) β ahead of nemotron-3-diarization-release (1.0x), behind cheatbench-reward-gaming-benchmark (1.0x)
Evidence (7) β β canonical anchor
| source | object | author | score | comments |
| π§ hn | Anthropic resumes charging for requests blocked by safeguardsRetrieved article excerptOpen article Β· Retrieved 2026-09-24T20:37:54.155775+00:00 [@ClaudeDevs](https://x.com/ClaudeDevs)
[ClaudeDevs](https://x.com/ClaudeDevs)
[Anthropic](https://twitter.com/AnthropicAI)
[@ClaudeDevs](https://x.com/ClaudeDevs)
Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false positive rates: biology, distillation attacks, and frontier LLM development. We've seen some coordinated attacks on our systems in recent weeks, and this is one layer of defense.
In recent testing, 99.7% of accounts using Claude Code, Claudeβ.ai, or Cowork did not hit any of these newly "billable blocks." The classifiers behind the blocks weβre resuming charging for today are tuned to have a <0.1% false positive rate. We know that's not 0%, and we're going to keep improving them so they interrupt your work less often. If you think a request has been blocked incorrectly, please report it with /feedback in Claude Code. [platform.claude.com/docs/en/build-β¦](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#how-refusals-are-billed)
[Refusals and fallback](https://t.co/uX6JhvN6He)[From platform.claude.com](https://t.co/uX6JhvN6He)
[5:11 PM Β· Sep 24, 2026](https://x.com/ClaudeDevs/status/2103170368794185758)Β·[310.1K
Views](https://x.com/ClaudeDevs/status/2103170368794185758)
[225](https://x.com/ClaudeDevs/status/2103170368794185758)
66
1.5K
352 | jeudesprits | 5 | 2 |
| π§ echo.x β | "Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit | Anthropic (@ClaudeDevs) | β | β |
| π reddit | Insane Claude Safeguards since Sonnet 5.5 release ClaudeAI | ofhgtl | 75 | 30 |
| π reddit | Opus 5.5 keeps flagging my non-security Claude Code sessions as [cyber]. Anyone else? ClaudeAI | jonntanny | 9 | 6 |
| π reddit | False flagging βreasoning_extractionβ is real problem ClaudeAI | marcandreewolf | 18 | 9 |
| π§ hn | AI safeguards are slowing developers down | dollar | 2 | 0 |
| π reddit | Do flagged chats really fallback to Opus 4.8? Mine routed to Opus 4.5. ClaudeAI Retrieved article excerptOpen article Β· Retrieved 2026-10-02T04:27:35.534824+00:00 # Prove your humanity
Weβre committed to safety and security. But not for bots. Complete the challenge below and let us know youβre
a real person.
[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)
[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us) | cross_peach | 1 | 3 |
Interpretation history
2026-10-02T05:08:33Z
The Oct 2 addition (Fable-5 flagged for kitten discussion, fallback routing to Opus 4.5 instead of the documented 4.8) repeats the established benign-false-flag pattern from a 1-point thread β no watch condition fired. The periphery that justified medium heat on Oct 1 has stopped expanding: thread sizes decay 74β15β10β1, velocity is ~0.3 pts/h vs ~39 peak at age ~207h, and no new community, outlet, or Anthropic response has appeared, so the case demotes from accelerating to corroborated and cools to low β a documented-but-unanswered billing dispute now waiting on Anthropic's next move rather than spreading.
2026-10-02T04:54:04Z
evidence attached: reddit.post.1wvk6w3 β Repeated flags on kitten discussion are direct evidence on whether the case's <0.1% false-positive tuning holds under the live billing regime.
2026-10-01T03:24:01Z
The FP grievance crossed from intra-community anecdote into press (VentureBeat, cross-vendor OpenAI+Anthropic) and acquired its first quantified financial harm ($15 paid credits plus weekly-limit burn on a benign legislative-analysis task, with cache-write charges on flagged sessions) β both were pre-declared material watch conditions. The case's meaning shifts from 'established policy with a dormant two-community FP complaint' to 'a billing policy whose <0.1% FP tuning claim is now publicly, measurably contested across three platforms', with a dedicated r/ClaudeAI megathread institutionalizing the grievance. Acceleration is in the counter-evidence periphery, not velocity: ~0.7 pts/h vs ~29 peak, but the periphery is expanding while each addition is thin β count the periphery, so heat holds at medium rather than cooling.
2026-10-01T02:29:13Z
evidence attached: hn.story.49916420 β VentureBeat coverage of developers at both OpenAI and Anthropic losing time to false safeguard flags is independent cross-vendor spread of the same false-positive episode the billing case is testing.
2026-10-01T02:29:13Z
evidence attached: reddit.post.1wulda1 β First-hand report of a false-positive reasoning_extraction block burning weekly limits and $15 of paid credits is direct evidence on whether the <0.1% FP tuning holds under billing pressure.
2026-09-30T09:44:13Z
The new r/ClaudeAI thread repeats the benign [cyber]-false-flag pattern on Opus 5.5 (fallback to 4.8) from a distinct user, showing the FP reports persist past launch day β marginally weakening the pure launch-confound caveat β but it is intra-community repetition of evidence already held, not a third independent line; no watch condition fired (no Anthropic response, no CVP extension, no pricing). The case's meaning is unchanged: established billing policy with a dormant two-community FP grievance awaiting Anthropic's next move. Heat stays low: ~1 pt/h at age ~164h vs ~23 peak β the 86th peer percentile reads a uniformly cooled cohort, not momentum, and the periphery is not expanding.
2026-09-30T09:24:16Z
evidence attached: reddit.post.1wtzkvn β Benign Claude Code sessions being flagged [cyber] with fallback to Opus 4.8 is user-visible evidence on the safeguard false-positive rate the case's <0.1% tuning claim hangs on.
2026-09-29T16:47:52Z
grounded: converges/high β Anthropic has independently arrived at the full-cost-per-transaction accounting Scott argues in his Agent Token Manifesto (dev:project.llmreport) β failure path
2026-09-29T16:38:48Z
relevance=high case never alerted; deterministic escalation to deliver
2026-09-28T22:18:33Z
The watch condition set at last review fired: billing-era false-positive reports now span a second independent community β r/ClaudeAI launch-day [cyber] blocks on Sonnet 5.5, compounded by a CVP verification gap on the 5.5 models β moving the counter-evidence on the <0.1% FP claim from lone anecdote to a cross-community pattern, though model-launch confounds and unconfirmed category billing keep it short of proof. Attention ticked up with the substance (86th percentile, steady momentum), so heat rises to medium to catch whether Anthropic tunes, acknowledges, or extends CVP.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wsoi7f β Launch-day reports of spurious [cyber] safeguard blocks plus the CVP coverage gap are false-positive evidence bearing on whether the <0.1% tuning claim holds.
2026-09-26T00:43:37Z
The case graduates from announced-policy seed to corroborated: the official docs now specify how refusals are billed, and the first billing-era false-positive reports (benign 'Use Unicode graphemes' flagged as biology, layout gibberish flagged unsafe) put concrete though anecdotal counter-evidence on the <0.1% FP claim inside billed categories β while engagement itself has gone quiet (0 pts/h, 13th percentile), so heat drops as substance firms up.
2026-09-24T21:12:08Z
grounded: converges/high β Anthropic has independently arrived at the full-cost-per-transaction accounting Scott already argues in AI Unit Economics and the Agent Token Manifesto/LLM Repo
2026-09-24T21:04:34Z
case created β Fresh first-party billing-policy act from Anthropic's own dev channel that no open anthropic case covers; materially changes API cost expectations for agent builders.
Decision trace
- 10-08 10:38review_screenjev screen: no material development (noul=0.11)
- 10-04 03:55review_screenThe diff swaps one anecdote for another: the deleted comment was a repeat of the already-documented 'drop to Opus 4.8' workaround (still corroborated by the launch thread), so its removal do
- 10-02 15:08repriceThe Oct 2 addition (Fable-5 flagged for kitten discussion, fallback routing to Opus 4.5 instead of the documented 4.8) repeats the established benign-false-flag pattern from a 1-point thread β no watc
- 10-02 14:54attachRepeated flags on kitten discussion are direct evidence on whether the case's <0.1% false-positive tuning holds under the live billing regime.
- 10-01 13:24repriceThe FP grievance crossed from intra-community anecdote into press (VentureBeat, cross-vendor OpenAI+Anthropic) and acquired its first quantified financial harm ($15 paid credits plus weekly-limit burn
- 10-01 12:29attachVentureBeat coverage of developers at both OpenAI and Anthropic losing time to false safeguard flags is independent cross-vendor spread of the same false-positive episode the billing case is testing.
- 10-01 12:29attachFirst-hand report of a false-positive reasoning_extraction block burning weekly limits and $15 of paid credits is direct evidence on whether the <0.1% FP tuning holds under billing pressure.
- 10-01 12:26propose_attachVentureBeat coverage of developers at both OpenAI and Anthropic losing time to false safeguard flags is independent cross-vendor spread of the same false-positive episode the billing case is testing.
- 10-01 12:26propose_attachFirst-hand report of a false-positive reasoning_extraction block burning weekly limits and $15 of paid credits is direct evidence on whether the <0.1% FP tuning holds under billing pressure.
- 09-30 19:44repriceThe new r/ClaudeAI thread repeats the benign [cyber]-false-flag pattern on Opus 5.5 (fallback to 4.8) from a distinct user, showing the FP reports persist past launch day β marginally weakening the pu
- 09-30 19:24attachBenign Claude Code sessions being flagged [cyber] with fallback to Opus 4.8 is user-visible evidence on the safeguard false-positive rate the case's <0.1% tuning claim hangs on.
- 09-30 19:23propose_attachBenign Claude Code sessions being flagged [cyber] with fallback to Opus 4.8 is user-visible evidence on the safeguard false-positive rate the case's <0.1% tuning claim hangs on.
- 09-30 02:56push"Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit β The velocity spike was the launch-day Reddit th
- 09-30 02:47repriceThe velocity spike was the launch-day Reddit thread's bump peaking and digesting β the thread grew to ~62 pts/23 comments and is already easing (65β62, comments flat), i.e. intra-thread amplifica
- 09-30 02:47groundAnthropic has independently arrived at the full-cost-per-transaction accounting Scott argues in his Agent Token Manifesto (dev:project.llmreport) β failure paths are now literally on the invoice β and
- 09-30 02:38alert_held"Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit β The velocity spike was the launch-day Reddit th
- 09-30 02:38alert_route"Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit β The velocity spike was the launch-day Reddit th
- 09-29 12:22sensor_dirtycomment_update
- 09-29 09:21sensor_dirtyvelocity_spike
- 09-29 08:18repriceThe watch condition set at last review fired: billing-era false-positive reports now span a second independent community β r/ClaudeAI launch-day [cyber] blocks on Sonnet 5.5, compounded by a CVP verif
- 09-29 07:36attachLaunch-day reports of spurious [cyber] safeguard blocks plus the CVP coverage gap are false-positive evidence bearing on whether the <0.1% tuning claim holds.
- 09-29 07:32propose_attachLaunch-day reports of spurious [cyber] safeguard blocks plus the CVP coverage gap are false-positive evidence bearing on whether the <0.1% tuning claim holds.
- 09-26 10:43repriceThe case graduates from announced-policy seed to corroborated: the official docs now specify how refusals are billed, and the first billing-era false-positive reports (benign 'Use Unicode graphem
- 09-26 10:42review_screenNew story adds a first-party Anthropic docs link on refusals plus concrete first-hand false-positive reports (benign 'Use Unicode graphemes' flagged as biology, layout gibberish flagged unsa
- 09-26 10:41review_screenjev screen borderline (noul=0.62) β luna review
- 09-25 07:12groundAnthropic has independently arrived at the full-cost-per-transaction accounting Scott already argues in AI Unit Economics and the Agent Token Manifesto/LLM Report work: safeguard blocks are now a pric
- 09-25 07:04createFresh first-party billing-policy act from Anthropic's own dev channel that no open anthropic case covers; materially changes API cost expectations for agent builders.