Reddit user cross_peach claims Anthropic's stated fallback for repeatedly flagged Fable-5 conversations routed him to Opus 4.5 instead of the documented Opus 4.8, after flags fired on trivially benign content; wider reports, replication, or Anthropic acknowledgment resolves whether the published fallback contract is broken.
state: seedheat: lowuncertainty: mediumconvergesscott: lowanthropic safeguard-fallback model-routingAnthropic
What is this?
Claude Fable 5 is Anthropic's Mythos-class flagship (2026), shipped with safety classifiers over cybersecurity, bio/chem, and distillation content that β per Anthropic's documentation and its RSP v3.0 commitments β route flagged requests to Claude Opus 4.8 ('our next-most-capable model') with user notification rather than refusing outright. That the classifiers fire on benign content is well corroborated across the supplied sources β a GitHub issue on legitimate anti-fraud work, Reddit posts on defensive security, press coverage of blocked biology questions β and Anthropic concedes launch tuning is deliberately conservative (<5% of sessions; Karpathy called it 'a little too trigger happy'). However, no supplied source corroborates the case's central claim: every source that names a fallback target names Opus 4.8 (Opus 4.5 appears only as its predecessor in the safety line, never as a documented fallback target), and there is no replication or Anthropic acknowledgment of a 4.5 routing β the claim rests on a single Reddit post seen here only by title. The documented contract also varies by surface (automatic switching on Claude app surfaces including Claude Code; opt-in configuration on the API), which any 'broken contract' claim would need to pin down.
Why it matters to Scott
The alleged wrong-target fallback (Opus 4.5 where Opus 4.8 is documented) is uncorroborated as supplied β a single Reddit post, no replication or acknowledgment, every other source naming Opus 4.8 β so it cannot yet be priced as a contradiction; but if it holds it lands squarely on the failure mode Scott's canon is built around: provider-side safeguard routing that cannot be trusted from documentation and must be verified from API model-field and billing metadata (his 12-factor silent-swap detection; owned-output model levels' 'absent, never substituted, exactly one explicit marked fallback'). It extends the radar's Fable-5 safeguard surface with a new falsifiable edge β predecessor-downgrade, the opposite direction of the silent-upgrade sibling episode β and replication or Anthropic acknowledgment would turn his detection patterns and LiteLLM verification discipline into a dated-receipts publishing opportunity.
ip:framework.12-factor-agents-frameworkdev:concept.owned-output-model-levelsdev:technology.litellmdev:concept.provider-bound-reasoning-continuityradar:fable-5-safeguard-fallbacksradar:anthropic-fable55-silent-routingradar:anthropic-blocked-request-billingradar:concept.anthropicradar:concept.model-routingradar:concept.model-identityradar:concept.model-safety
queries asked of Scott's wikis
- agent harness handling of model fallback and refusal stop reasons
- detecting silent model swaps from API model field and billing rate metadata
- safety classifier false positives on benign professional content
- critique of Anthropic over-refusal and paternalistic safety posture
- open-weight local models as hedge against provider-side safeguard routing
- sticky session downgrade as an AI product reliability pattern
Measured heat
now 0 pts/hpeak 45 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 3002h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
Evidence (5) β β canonical anchor
Interpretation history
2026-10-07T10:36:10Z
The dthrdr thread's key testimony is that the mechanism itself has changed shape β on Fable 5.1, safeguards now visibly notify-and-continue rather than shunting to another model β which dissolves the case's remaining testable edge: the never-replicated 4.5 misroute describes behavior of a superseded mechanism, while the corroborated story remains the already-known over-triggering. The watch shifts to whether the notify-and-continue mechanism gets documented or silently diverges from the published switch-to-4.8 contract.
2026-10-07T10:25:58Z
evidence attached: reddit.post.1wzryhk β Independent corroboration of safeguards firing on benign content, plus a reported mechanism change to notify-and-continue instead of shunting to another model that re-judging the fallback contract must account for.
2026-10-06T10:21:39Z
The wider report the seed named as resolution-relevant has arrived β and it describes the contract working as documented (visible 'Cyber' flag, visible switch to Opus 4.8), leaving the 4.5 misroute a non-replicating singleton riding on already-established classifier over-triggering. The case cools from a live test of Anthropic's June promise to a watch for whether any wrong-target fallback ever produces a metadata-grade receipt.
2026-10-06T09:27:25Z
evidence attached: reddit.post.1wywx78 β Second independent user reporting benign content triggering 'Cyber' flags and forced model fallback β the wider-report evidence the seed case says would resolve it.
2026-10-04T18:29:45Z
The 4.5-routing claim itself gains nothing β still one reporter, zero replication β but the attached thread pins the exact contract under test (June visible-fallback promise, Fortune-quoted) and supplies a plausible violation mechanism (Sept 8 NSA/CISA/FBI advisory reportedly urging silent downgrades of suspected distillers), reframing the case from a lone anomaly into a live test of whether Anthropic's promised transparency survives government pressure.
2026-10-04T18:25:58Z
evidence attached: reddit.post.1wxm5bh β Documents the June visible-fallback promise Anthropic made and the Sept 8 NSA/CISA advisory pressure to silently downgrade suspected distillers β core evidence on whether the published fallback contract holds; spread across ClaudeAI adds weight.
2026-10-02T04:59:23Z
origin walked (opencode/cheap-glm, conf 0.9): anchor reddit.post.1wvk6w3 -> echo.blog.147a989f3e by Anthropic
2026-10-02T04:54:04Z
grounded: converges/medium β The alleged wrong-target fallback (Opus 4.5 where Opus 4.8 is documented) is uncorroborated as supplied β a single Reddit post, no replication or acknowledgment
2026-10-02T04:45:35Z
case created β A checkable contradiction of Anthropic's stated fallback mechanism β a distinct claim from the silent-upgrade and billing episodes sharing the same safeguard surface.
Decision trace
- 10-07 21:36repriceThe dthrdr thread's key testimony is that the mechanism itself has changed shape β on Fable 5.1, safeguards now visibly notify-and-continue rather than shunting to another model β which dissolves
- 10-07 21:25attachIndependent corroboration of safeguards firing on benign content, plus a reported mechanism change to notify-and-continue instead of shunting to another model that re-judging the fallback contract mus
- 10-07 21:25propose_attachIndependent corroboration of safeguards firing on benign content, plus a reported mechanism change to notify-and-continue instead of shunting to another model that re-judging the fallback contract mus
- 10-06 21:21repriceThe wider report the seed named as resolution-relevant has arrived β and it describes the contract working as documented (visible 'Cyber' flag, visible switch to Opus 4.8), leaving the 4.5 m
- 10-06 20:27attachSecond independent user reporting benign content triggering 'Cyber' flags and forced model fallback β the wider-report evidence the seed case says would resolve it.
- 10-06 20:23propose_attachSecond independent user reporting benign content triggering 'Cyber' flags and forced model fallback β the wider-report evidence the seed case says would resolve it.
- 10-05 21:47review_screenjev screen: no material development (noul=0.13)
- 10-05 10:21sensor_dirtycomment_update
- 10-05 05:29repriceThe 4.5-routing claim itself gains nothing β still one reporter, zero replication β but the attached thread pins the exact contract under test (June visible-fallback promise, Fortune-quoted) and suppl
- 10-05 05:25attachDocuments the June visible-fallback promise Anthropic made and the Sept 8 NSA/CISA advisory pressure to silently downgrade suspected distillers β core evidence on whether the published fallback contra
- 10-05 05:23propose_attachDocuments the June visible-fallback promise Anthropic made and the Sept 8 NSA/CISA advisory pressure to silently downgrade suspected distillers β core evidence on whether the published fallback contra
- 10-02 14:59promote_anchororigin walk conf 0.9
- 10-02 14:54groundThe alleged wrong-target fallback (Opus 4.5 where Opus 4.8 is documented) is uncorroborated as supplied β a single Reddit post, no replication or acknowledgment, every other source naming Opus 4.8 β s
- 10-02 14:45createA checkable contradiction of Anthropic's stated fallback mechanism β a distinct claim from the silent-upgrade and billing episodes sharing the same safeguard surface.