Reddit user National_Wolverine_7 reports Claude Code's auto mode intermittently hard-stops even trivial edits behind its safety check and stays stuck for days โ an Anthropic acknowledgment or fix, or wider user reports of the same fail-closed gate, would establish auto mode's server-side safety layer as a recurring blocker of routine agent work rather than a one-off glitch.
state: resolvedheat: lowuncertainty: mediumconvergesscott: mediumagent-harnesses claude-code agent-safetyAnthropic
What is this?
Claude Code's auto mode is Anthropic's permission system for its coding agent: a safety classifier adjudicates each shell command and tool call in context instead of asking per action, and it became the default mode for Pro/Max/Team plans on August 14, 2026 โ Anthropic cites an internal study with a ~89% dangerous-command catch rate vs ~13.6% for humans. Release v2.1.278 (Sept 19, 2026) moved the classifier server-side and unbilled, but only for Enterprise/API/Bedrock/Vertex/Foundry accounts; Pro/Max/Team keep a local classifier, and the documented server-path failure mode is an LLM gateway stripping the safeguards fields, which drops to the old billed local check. The case's specific claims โ one user stuck for days as the gate refuses trivial edits, another seeing it approve a destructive repo-discard while denying a harmless read โ appear nowhere in the supplied coverage, and there is no Anthropic acknowledgment; an independent stress-test (the AmPermBench paper on alphaXiv) actually emphasizes the opposite failure direction, an ~81% end-to-end false-negative rate on ambiguous scenarios with in-project file edits passing ungated. So the web establishes the gate's design, defaults, and documented failure modes but nothing corroborating these misfires โ and a 'shared server-side regression' hypothesis is awkward, since the server-side path never reached consumer plans.
Why it matters to Scott
Converges: a probabilistic classifier now sits as the default permission boundary on Scott's primary coding agent, and these reports document both failure directions his guardrail-illusion and risk-based-triage pages predict โ over-gating trivial reversible edits for days, and approving an irreversible repo-discard while denying a harmless read as 'Irreversible Local Destruction', an irreversibility-gradient inversion in the wild that a deterministic boundary (his DAI argument) would not make. Medium not high: two low-engagement single-platform posts with no vendor acknowledgment, and the grounding shows consumer plans still run a local classifier โ so the episodes share classifier-as-gate, not yet a confirmed server-side regression; that makes this dated receipts plus an operational hazard to his unattended Claude Code runs rather than something that changes what he builds.
ip:concept.guardrail-illusionip:concept.risk-based-triageip:concept.irreversibility-gradientip:framework.decision-authority-infrastructureip:framework.long-running-agentsdev:project.askradar:claude-code-server-side-write-classifierradar:claude-code-auto-mode-defaultradar:fable-5-safeguard-fallbacksradar:concept.agent-safety
queries asked of Scott's wikis
- "guardrail illusion" classifier-based safety critique
- risk-based permission triage agent tool calls
- fail-closed vs fail-open agent permission gates
- unattended headless agent run reliability
- router proxy passthrough of unknown request fields safeguards
- claude code auto mode sandbox versus classifier judgment
Measured heat
now 0 pts/hpeak 4 pts/hcomments 1/hpeers p64momentum: steady1 platformsage 140h
points/hour across evidence ยท reading as of 2026-10-07 04:22:34.109904+11:00 ยท deterministic, not a model opinion
How the heat travelled
Evidence (3) โ โญ canonical anchor
Interpretation history
2026-10-06T18:49:40Z
The third independent report โ auto mode newly refusing a task routine for six months, change dated to 'the past week or so' โ completes the case's own resolution condition ('wider user reports of the same fail-closed gate'): three users, three distinct signatures (fail-closed stall, irreversibility inversion, new refusals of routine work) within roughly a week, establishing the gate's misbehavior as recurring rather than a one-off. Resolving absorbed even though the evidence never crossed into cross-platform corroboration โ the hypothesis's stated bar was 'wider reports', not vendor confirmation, and any Anthropic acknowledgment or patch is a new episode under the already-hot claude-code topic rather than a continuation of this one.
2026-10-06T16:42:15Z
evidence attached: reddit.post.1wz49z1 โ A second user report of auto mode newly refusing routine tasks โ exactly the wider-report evidence the case names as resolving.
2026-10-06T15:36:47Z
grounded: converges/medium โ Converges: a probabilistic classifier now sits as the default permission boundary on Scott's primary coding agent, and these reports document both failure direc
2026-10-06T15:25:59Z
A second independent user report recasts the case from a single stall anecdote into a small cluster of adjudication failures with two distinct signatures: over-gating of trivial reversible edits (stuck for days) and outright misattribution โ the destructive repo-discard command passed while a harmless file read was denied as 'Irreversible Local Destruction'. That widening moves it past seed, but both lines are same-subreddit, micro-engagement, unconfirmed-by-vendor testimony, so it stops short of corroborated.
2026-10-06T13:34:27Z
evidence attached: reddit.post.1wz1prh โ Another documented misfire of the same auto-mode safety layer, this time denying the wrong command, widening the case's failure characterization from stalls to misattribution.
2026-10-01T00:40:55Z
grounded: converges/high โ Third Reddit episode in the radar's server-side auto-mode-classifier failure series (after mazarax's write-block disclosure and wacoder's no-verdict outage), an
2026-10-01T00:31:23Z
case created โ A concrete, bounded harness-failure claim โ a non-overridable safety gate blocking routine auto-mode edits โ matches the queue's established single-report seed pattern for Claude Code bugs and resolves on Anthropic's response or corroboration.
Decision trace
- 10-07 05:49resolveThe third independent report โ auto mode newly refusing a task routine for six months, change dated to 'the past week or so' โ completes the case's own resolution condition ('wider
- 10-07 03:42attachA second user report of auto mode newly refusing routine tasks โ exactly the wider-report evidence the case names as resolving.
- 10-07 03:34propose_attachA second user report of auto mode newly refusing routine tasks โ exactly the wider-report evidence the case names as resolving.
- 10-07 02:36repriceA second independent user report recasts the case from a single stall anecdote into a small cluster of adjudication failures with two distinct signatures: over-gating of trivial reversible edits (stuc
- 10-07 02:36groundConverges: a probabilistic classifier now sits as the default permission boundary on Scott's primary coding agent, and these reports document both failure directions his guardrail-illusion and ri
- 10-07 00:34attachAnother documented misfire of the same auto-mode safety layer, this time denying the wrong command, widening the case's failure characterization from stalls to misattribution.
- 10-07 00:26propose_attachAnother documented misfire of the same auto-mode safety layer, this time denying the wrong command, widening the case's failure characterization from stalls to misattribution.
- 10-01 10:40groundThird Reddit episode in the radar's server-side auto-mode-classifier failure series (after mazarax's write-block disclosure and wacoder's no-verdict outage), and it converges with his g
- 10-01 10:31createA concrete, bounded harness-failure claim โ a non-overridable safety gate blocking routine auto-mode edits โ matches the queue's established single-report seed pattern for Claude Code bugs and re