2026-10-11 16:38 UTC

Anthropic has resumed charging for safeguard-blocked requests in low-false-positive categories (biology, distillation attacks, frontier LLM development) as a stated defense layer against coordinated attacks, making blocked calls a real line item in agent API economics and testing whether its <0.1% false-positive tuning holds under billing pressure.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highanthropic inference-economics agentic-securityAnthropic
Surfaced 2026-09-29T16:56:19Z β€” "Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit β€” The velocity spike was the launch-day Reddit thread's bump peaking and digesting β€” the thread grew to ~62 pts/23 comments and is already easing (65β†’62, comments flat), i.e. intra-thread amplification of FP reports the case already holds, not a third community or outlet. No Anthropic response, no CVP extension, no pricing disclosure; the case's meaning shifts from 'unfolding FP controversy' to 'established policy with a dormant two-community FP grievance awaiting Anthropic's next move', so heat cools to low despite the 78th percentile reading β€” that percentile reflects a small cooling cohort, not live momentum.

What is this?

Anthropic announced on September 24, 2026 (ClaudeDevs post, echoed in its platform release notes and refusal/fallback docs) that it has resumed charging for API requests its safeguards block before Claude produces any output β€” limited to three categories it tunes to a <0.1% false-positive rate: biology, distillation attacks, and frontier LLM development (documented as the `bio`, `frontier_llm`, and `reasoning_extraction` stop_details categories; other pre-output refusals stay unbilled but count against rate limits). The company frames billing as one layer of defense against coordinated attacks, citing a September 2026 threat-intelligence report on large-scale illicit distillation campaigns (the supplied snippet names Alibaba's, peaking near 3M exchanges/day from 3,500+ fraudulent accounts), and reports 99.7% of Claude Code/claude.ai/Cowork accounts triggered none of the newly billable blocks in testing. Community counter-evidence on the false-positive claim β€” benign prompts tripping bio/cyber flags β€” remains first-hand but anecdotal (HN, r/ClaudeAI launch-day thread), and the supplied coverage does not confirm the case's specific attribution to DeepSeek/Moonshot/MiniMax, any pricing figures, or whether CVP-gated cyber blocks fall inside billed categories.

Why it matters to Scott

Anthropic has independently arrived at the full-cost-per-transaction accounting Scott argues in his Agent Token Manifesto (dev:project.llmreport) β€” failure paths are now literally on the invoice β€” and the new cross-community false-positive reports plus the context-carried [cyber]-trigger finding are a dated receipt for Guardrail Illusion: a probabilistic safeguard promoted to a billing boundary, failing in exactly the unreliable-permission-boundary way that concept names. It also bears on his own stack operationally: every project routes through the LiteLLM proxy, so stop_details refusal categories now carry monetary cost, and a billed block plus blind retry in an autonomous harness becomes a burn loop until stop_details-aware resumption/autonomy-budget handling exists.
ip:concept.guardrail-illusiondev:project.llmreportdev:technology.litellmdev:project.askradar:fable-5-safeguard-fallbacksradar:concept.model-safetyradar:concept.inference-economicsradar:concept.anthropicradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
  • AI unit economics full cost per transaction failure paths
  • agent retry policy on blocked refusals autonomy budget
  • guardrail false positives safety classifier reliability cost
  • distillation attacks defense economics making probing expensive
  • LiteLLM harness handling stop_details refusal categories
  • guardrail illusion layered safeguards model-level vs around-model

Measured heat

now 0 pts/hpeak 36 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 434h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-23 14:00⭐ origin echo-reconstructed"Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false posit
Anthropic (@ClaudeDevs) on x (echo) Β· attributed from hn.story.49835073
β€”
09-24 18:42first on hacker news Β· published Β· +28.7hAnthropic resumes charging for requests blocked by safeguards
jeudesprits
β€”
09-28 19:36first on r/ClaudeAI Β· published Β· +125.6hInsane Claude Safeguards since Sonnet 5.5 release
ofhgtl
β€”
09-24 18:42amplified on hacker newshn.story.49835073
jeudesprits
peak 5 Β· 2 comments Β· 7% of case engagement
09-28 19:36amplified on r/ClaudeAI πŸ‘‘reddit.post.1wsoi7f
ofhgtl
peak 75 Β· 30 comments Β· 62% of case engagement
09-30 08:33amplified on r/ClaudeAIreddit.post.1wtzkvn
jonntanny
peak 10 Β· 6 comments Β· 10% of case engagement
10-01 00:27amplified on r/ClaudeAIreddit.post.1wulda1
marcandreewolf
peak 18 Β· 9 comments Β· 16% of case engagement
10-01 00:59amplified on hacker newshn.story.49916420
dollar
peak 2 Β· 0 comments Β· 2% of case engagement
10-02 04:08amplified on r/ClaudeAIreddit.post.1wvk6w3
cross_peach
peak 2 Β· 3 comments Β· 3% of case engagement
09-24 20:21our radar first saw it Β· +30.4hdiscovery anchor: hn.story.49835073β€”
09-29 16:47reached heat=high Β· +146.8h Β· via ledgerβ€”β€”
pace: p72 vs 1032 stories at the 336h mark (now 434h old) β€” ahead of nemotron-3-diarization-release (1.0x), behind cheatbench-reward-gaming-benchmark (1.0x)

Evidence (7) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAnthropic resumes charging for requests blocked by safeguards
Retrieved article excerpt

Open article Β· Retrieved 2026-09-24T20:37:54.155775+00:00

[@ClaudeDevs](https://x.com/ClaudeDevs)

[ClaudeDevs](https://x.com/ClaudeDevs)

[Anthropic](https://twitter.com/AnthropicAI)

[@ClaudeDevs](https://x.com/ClaudeDevs)

Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false positive rates: biology, distillation attacks, and frontier LLM development. We've seen some coordinated attacks on our systems in recent weeks, and this is one layer of defense.
In recent testing, 99.7% of accounts using Claude Code, Claude​.ai, or Cowork did not hit any of these newly "billable blocks." The classifiers behind the blocks we’re resuming charging for today are tuned to have a <0.1% false positive rate. We know that's not 0%, and we're going to keep improving them so they interrupt your work less often. If you think a request has been blocked incorrectly, please report it with /feedback in Claude Code. [platform.claude.com/docs/en/build-…](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#how-refusals-are-billed)

[Refusals and fallback](https://t.co/uX6JhvN6He)[From platform.claude.com](https://t.co/uX6JhvN6He)

[5:11 PM Β· Sep 24, 2026](https://x.com/ClaudeDevs/status/2103170368794185758)Β·[310.1K

Views](https://x.com/ClaudeDevs/status/2103170368794185758)

[225](https://x.com/ClaudeDevs/status/2103170368794185758)

66

1.5K

352
jeudesprits52
🟧 echo.x ⭐"Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false positAnthropic (@ClaudeDevs)β€”β€”
🟠 redditInsane Claude Safeguards since Sonnet 5.5 release
ClaudeAI
ofhgtl7530
🟠 redditOpus 5.5 keeps flagging my non-security Claude Code sessions as [cyber]. Anyone else?
ClaudeAI
jonntanny96
🟠 redditFalse flagging β€œreasoning_extraction” is real problem
ClaudeAI
marcandreewolf189
🟧 hnAI safeguards are slowing developers downdollar20
🟠 redditDo flagged chats really fallback to Opus 4.8? Mine routed to Opus 4.5.
ClaudeAI
Retrieved article excerpt

Open article Β· Retrieved 2026-10-02T04:27:35.534824+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
cross_peach13

Interpretation history

Decision trace