2026-10-11 16:38 UTC

Reddit user cross_peach claims Anthropic's stated fallback for repeatedly flagged Fable-5 conversations routed him to Opus 4.5 instead of the documented Opus 4.8, after flags fired on trivially benign content; wider reports, replication, or Anthropic acknowledgment resolves whether the published fallback contract is broken.

state: seedheat: lowuncertainty: mediumconvergesscott: lowanthropic safeguard-fallback model-routingAnthropic

What is this?

Claude Fable 5 is Anthropic's Mythos-class flagship (2026), shipped with safety classifiers over cybersecurity, bio/chem, and distillation content that β€” per Anthropic's documentation and its RSP v3.0 commitments β€” route flagged requests to Claude Opus 4.8 ('our next-most-capable model') with user notification rather than refusing outright. That the classifiers fire on benign content is well corroborated across the supplied sources β€” a GitHub issue on legitimate anti-fraud work, Reddit posts on defensive security, press coverage of blocked biology questions β€” and Anthropic concedes launch tuning is deliberately conservative (<5% of sessions; Karpathy called it 'a little too trigger happy'). However, no supplied source corroborates the case's central claim: every source that names a fallback target names Opus 4.8 (Opus 4.5 appears only as its predecessor in the safety line, never as a documented fallback target), and there is no replication or Anthropic acknowledgment of a 4.5 routing β€” the claim rests on a single Reddit post seen here only by title. The documented contract also varies by surface (automatic switching on Claude app surfaces including Claude Code; opt-in configuration on the API), which any 'broken contract' claim would need to pin down.

Why it matters to Scott

The alleged wrong-target fallback (Opus 4.5 where Opus 4.8 is documented) is uncorroborated as supplied β€” a single Reddit post, no replication or acknowledgment, every other source naming Opus 4.8 β€” so it cannot yet be priced as a contradiction; but if it holds it lands squarely on the failure mode Scott's canon is built around: provider-side safeguard routing that cannot be trusted from documentation and must be verified from API model-field and billing metadata (his 12-factor silent-swap detection; owned-output model levels' 'absent, never substituted, exactly one explicit marked fallback'). It extends the radar's Fable-5 safeguard surface with a new falsifiable edge β€” predecessor-downgrade, the opposite direction of the silent-upgrade sibling episode β€” and replication or Anthropic acknowledgment would turn his detection patterns and LiteLLM verification discipline into a dated-receipts publishing opportunity.
ip:framework.12-factor-agents-frameworkdev:concept.owned-output-model-levelsdev:technology.litellmdev:concept.provider-bound-reasoning-continuityradar:fable-5-safeguard-fallbacksradar:anthropic-fable55-silent-routingradar:anthropic-blocked-request-billingradar:concept.anthropicradar:concept.model-routingradar:concept.model-identityradar:concept.model-safety
queries asked of Scott's wikis
  • agent harness handling of model fallback and refusal stop reasons
  • detecting silent model swaps from API model field and billing rate metadata
  • safety classifier false positives on benign professional content
  • critique of Anthropic over-refusal and paternalistic safety posture
  • open-weight local models as hedge against provider-side safeguard routing
  • sticky session downgrade as an AI product reliability pattern

Measured heat

now 0 pts/hpeak 45 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 3002h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

06-08 14:00⭐ origin echo-reconstructedAnthropic's "Claude Fable 5 and Claude Mythos 5" post introduced the fallback policy and assurance the Reddit post disputes: "When Fable's c
Anthropic on blog (echo) Β· attributed from reddit.post.1wvk6w3
β€”
10-02 04:08first on r/ClaudeAI Β· published Β· +2774.1hDo flagged chats really fallback to Opus 4.8? Mine routed to Opus 4.5.
cross_peach
β€”
10-02 04:08amplified on r/ClaudeAIreddit.post.1wvk6w3
cross_peach
peak 2 Β· 3 comments Β· 6% of case engagement
10-04 18:02amplified on r/ClaudeAI πŸ‘‘reddit.post.1wxm5bh
MysteriousAvocado580
peak 37 Β· 12 comments Β· 57% of case engagement
10-06 08:04amplified on r/ClaudeAIreddit.post.1wywx78
True_Profile3695
peak 5 Β· 10 comments Β· 18% of case engagement
10-07 09:20amplified on r/ClaudeAIreddit.post.1wzryhk
dthrdr
peak 9 Β· 7 comments Β· 19% of case engagement
10-02 04:20our radar first saw it Β· +2774.3hdiscovery anchor: reddit.post.1wvk6w3β€”

Evidence (5) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDo flagged chats really fallback to Opus 4.8? Mine routed to Opus 4.5.
ClaudeAI
Retrieved article excerpt

Open article Β· Retrieved 2026-10-02T04:27:35.534824+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
cross_peach13
🟧 echo.blog ⭐Anthropic's "Claude Fable 5 and Claude Mythos 5" post introduced the fallback policy and assurance the Reddit post disputes: "When Fable's cAnthropicβ€”β€”
🟠 redditIn June Anthropic apologized for silently rerouting flagged Fable 5 requests. The Sept 8 NSA/CISA/FBI advisory tells labs to downgrade suspected distillers without telling them. Does the June promise still hold?
ClaudeAI
MysteriousAvocado5803712
🟠 redditAm I banned from Opus 5.5 or something?
ClaudeAI
True_Profile3695510
🟠 redditNew safeguards are way too restrictive, it won't even download a firmware update from a public OTA server
ClaudeAI
dthrdr97

Interpretation history

Decision trace