Claude Code users including wacoder report a same-day wave of server-side auto-mode classifier failures that block Bash tool calls with 'no verdict' errors; Anthropic's acknowledgment or a fix β or users adopting the CLAUDE_CODE_AUTO_MODE_SERVER=0 bypass β will establish remote safety-verdict dependence as a recognized agent-harness failure mode.
state: resolvedheat: lowuncertainty: lowconvergesscott: highagent-containment agent-harnesses claude-code safety-classifiersAnthropic
What is this?
Claude Code's auto mode gates state-changing tool calls (Bash, Write, WebFetch) behind a safety classifier; as of v2.1.278 (Sept 19, 2026) that classifier runs server-side for Enterprise and platform/API accounts (Bedrock, Vertex, Foundry), with CLAUDE_CODE_AUTO_MODE_SERVER=0 as a documented opt-out back to the older billed local classifier. Multiple GitHub issues spanning JuneβSeptember 2026 (#74351, #74248, #78576, #85850) document recurring classifier-availability failures during which the harness fails safe β refusing the gated tool call outright, with no client-side retry and no fallback to interactive approval β at times pushing users into bypass-permissions mode to keep working. The specific same-day wave, the user 'wacoder', and the exact 'no verdict' error string in the hypothesis are not corroborated by the supplied snippets; what is well-corroborated is that hard-blocked tool calls pending a remote safety verdict are a real, ongoing, multi-month failure pattern, with no fix acknowledged in the supplied material beyond gateway-fallback notes.
Why it matters to Scott
Anthropic made a probability classifier the enforcement layer for consequential tool calls, and when it's unavailable the harness hard-blocks with no retry and no degraded human-approval tier β the absence-freezes-the-system anti-pattern his owned-output-model-levels and risk-based-triage pages warn against, and a receipt-backed live demonstration of guardrail-illusion: users fleeing to bypass-permissions mode show fail-closed gating without a middle tier drives people to remove safety entirely (the flip side of his deliberately fail-open guarded-agent-inbox). It bears directly on tooling he operates daily (Claude Code, including the hourly/weekly router agents), and the corroborated multi-month issue trail plus the documented CLAUDE_CODE_AUTO_MODE_SERVER=0 flag gives dated receipts for a 'remote verdict dependency' failure-mode piece β though the grounding flags the specific same-day 'no verdict' wave as uncorroborated, so the argument should rest on the pattern, not the spike.
ip:concept.guardrail-illusionip:concept.risk-based-triagedev:concept.owned-output-model-levelsdev:concept.guarded-agent-inboxdev:technology.claude-codedev:project.routerradar:agent-chaperone-jev-tool-screeningradar:agent-trace-tamperingradar:ac2-agent-security-protocol
queries asked of Scott's wikis
- agent harness remote dependency fail-safe tool gating
- fail-open vs fail-open guardrail agent containment sandbox
- local model classifier offline tool approval coding agent
- agent scaffold retry backoff graceful degradation provider outage
- Claude Code permission modes bypass hooks workflow
- safety classifier false positive friction agent reliability cost
Measured heat
no measured readings yet β the hourly heat pass fills this in
How the heat travelled
Evidence (2) β β canonical anchor
Interpretation history
2026-09-29T15:13:14Z
The velocity-spike trigger was a small-denominator artifact (+1 upvote against a 0.167/h peer baseline); measured activity is now flat zero at 31.6h age, past the early peak. The Sep 28 wave itself corroborated across two lines (Reddit thread plus three same-day GitHub issues #97812/#97854/#97855 and the Rupali59 primary report), but the episode subsided on its own within about a day with no Anthropic acknowledgment, no fix, and no adoption wave beyond the already-documented bypass flag β so the hypothesis's establishing event never landed and the live-outage window is closed. Durable value (the multi-month fail-closed classifier trail as receipts) is already grounded and preserved in the assessment.
2026-09-28T16:54:02Z
origin walked (opencode/cheap-glm, conf 0.75): anchor reddit.post.1wseypz -> echo.github.0b7d108cbd by Rupali59
2026-09-28T16:17:55Z
grounded: converges/high β Anthropic made a probability classifier the enforcement layer for consequential tool calls, and when it's unavailable the harness hard-blocks with no retry and
2026-09-28T16:07:08Z
case created β Live multi-user safety-classifier outage with a documented bypass flag is a bounded, actively moving harness-failure episode absent from the open set.
Decision trace
- 09-30 01:13resolveThe velocity-spike trigger was a small-denominator artifact (+1 upvote against a 0.167/h peer baseline); measured activity is now flat zero at 31.6h age, past the early peak. The Sep 28 wave itself co
- 09-29 04:22sensor_dirtyvelocity_spike
- 09-29 02:54promote_anchororigin walk conf 0.75
- 09-29 02:17groundAnthropic made a probability classifier the enforcement layer for consequential tool calls, and when it's unavailable the harness hard-blocks with no retry and no degraded human-approval tier β t
- 09-29 02:07createLive multi-user safety-classifier outage with a documented bypass flag is a bounded, actively moving harness-failure episode absent from the open set.