Claude Code's own verbatim error text, reported by Reddit user mazarax, discloses that local Write actions are gated by a server-side Anthropic auto-mode safety classifier whose failures block writes โ if confirmed as standing architecture, Claude Code's local writes depend on remote classifier availability and every write is observable to Anthropic.
state: resolvedheat: lowuncertainty: lowconvergesscott: highagent-harnesses agentic-security coding-agentsAnthropic
What is this?
Claude Code's Auto Mode โ launched as a research preview in March 2026 and made the default permission model for Pro/Max/Team sessions on August 14, 2026 โ replaces per-action human approval with a separate classifier model (Sonnet 5 by default, falling back to the session model) that evaluates each proposed tool call before execution. The supplied snippets confirm the gate is fail-closed and remote: Anthropic's docs describe a specific error when the classifier cannot determine action safety, and a wire-captured GitHub issue (#78372) shows the classifier call โ a ~107KB policy prompt plus the proposed action โ routes through ANTHROPIC_BASE_URL (undocumented), so under Auto Mode the proposed action does cross the network to whoever serves the classifier endpoint (Anthropic by default; the user's gateway if a custom base URL is set). The specific mazarax Reddit post anchoring this case is not among the supplied snippets, but the architecture it describes is independently corroborated by first-party docs, the outage error text, and the wire capture. Security context is already hot: Rehberger disclosed a 60โ80%-success RCE chain against Auto Mode on August 26, 2026, and CSA frames these as harness-level designs that undermine otherwise-reasonable safety controls.
Why it matters to Scott
Anthropic's Auto Mode puts a probabilistic LLM classifier in the enforcement path for every Claude Code write โ exactly the Guardrail Illusion / Compliance Cosplay pattern (probabilistic behaviour sold as a permission boundary), now with dated receipts: fail-closed outage errors, a wire-captured ~107KB classifier call routed via ANTHROPIC_BASE_URL, and Rehberger's 60โ80% RCE success against the same gate. The fail-closed remote gate also makes unattended-agent writes hostage to vendor availability and ships every proposed write to Anthropic โ bearing on his long-running-agent reliability work, his SiloOS answer to the same gating problem, and his LiteLLM/base-URL routing practice (a concrete check: which permission model and endpoint his own Claude Code sessions actually run under).
ip:concept.guardrail-illusionip:source.compliance-cosplayip:framework.decision-authority-infrastructureip:concept.manners-vs-physicsip:framework.long-running-agentsip:framework.sovereign-software-assurancedev:technology.claude-codedev:project.silo-osdev:technology.litellmradar:claude-code-denied-read-secret-bypassradar:claude-code-remote-attribution-injectionradar:mcp-tool-sequence-guardrail-bypassradar:hosted-open-model-endpoint-failuresradar:runbook-mcp-fail-closed-workflows
queries asked of Scott's wikis
- agent harness tool-call safety gating fail-closed architecture
- coding agent permission classifier prompt injection defense
- local-first privacy vendor observability of agent file writes
- Anthropic-compatible gateway local open model llama.cpp
- unattended long-running agent remote service availability dependency
- Claude Code settings workflows agent-maintained file writes
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
| 09-25 04:16 | โญ origin directly observed | Claude Phoning Home before write? mazarax on r/ClaudeAI | โ |
| 09-25 04:16 | amplified on r/ClaudeAI ๐ | reddit.post.1wpmq81 mazarax | peak 1 ยท 5 comments ยท 99% of case engagement |
| 09-25 04:20 | our radar first saw it ยท +0.1h | discovery anchor: reddit.post.1wpmq81 | โ |
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-26T01:43:43Z
The new comments confirm the mechanism with independent specificity (auto mode routes every non-read-only action through a server-side classifier) but narrow the claim: gating applies only under auto mode โ the default permission model since Aug 14, 2026 โ and is opt-out-able via settings.json allow rules or a permission-mode switch, correcting 'standing architecture gating every write' to 'default-mode writes'. With first-party docs, the wire-captured classifier call, and the verbatim error text all aligned, the checkable fact is established and this episode has no open question left to track.
2026-09-25T04:32:00Z
grounded: converges/high โ Anthropic's Auto Mode puts a probabilistic LLM classifier in the enforcement path for every Claude Code write โ exactly the Guardrail Illusion / Compliance Cosp
2026-09-25T04:23:22Z
case created โ The product's own error message is first-party-adjacent evidence of a consequential, easily checkable harness-architecture fact (remote gating of local writes) at the intersection of two hot topics.
Decision trace
- 09-26 11:43resolveThe new comments confirm the mechanism with independent specificity (auto mode routes every non-read-only action through a server-side classifier) but narrow the claim: gating applies only under auto
- 09-26 11:41review_screenTwo new comments (iniyanai, kevin_g_g) independently confirm and specify the mechanism: auto mode routes non-read-only actions (writes, shell commands) through a server-side classifier, explaining the
- 09-26 11:41review_screenjev screen borderline (noul=0.62) โ luna review
- 09-25 18:21sensor_dirtycomment_update
- 09-25 14:32groundAnthropic's Auto Mode puts a probabilistic LLM classifier in the enforcement path for every Claude Code write โ exactly the Guardrail Illusion / Compliance Cosplay pattern (probabilistic behaviou
- 09-25 14:23createThe product's own error message is first-party-adjacent evidence of a consequential, easily checkable harness-architecture fact (remote gating of local writes) at the intersection of two hot topi