2026-10-11 17:09 UTC

Claude Code's own verbatim error text, reported by Reddit user mazarax, discloses that local Write actions are gated by a server-side Anthropic auto-mode safety classifier whose failures block writes โ€” if confirmed as standing architecture, Claude Code's local writes depend on remote classifier availability and every write is observable to Anthropic.

state: resolvedheat: lowuncertainty: lowconvergesscott: highagent-harnesses agentic-security coding-agentsAnthropic

What is this?

Claude Code's Auto Mode โ€” launched as a research preview in March 2026 and made the default permission model for Pro/Max/Team sessions on August 14, 2026 โ€” replaces per-action human approval with a separate classifier model (Sonnet 5 by default, falling back to the session model) that evaluates each proposed tool call before execution. The supplied snippets confirm the gate is fail-closed and remote: Anthropic's docs describe a specific error when the classifier cannot determine action safety, and a wire-captured GitHub issue (#78372) shows the classifier call โ€” a ~107KB policy prompt plus the proposed action โ€” routes through ANTHROPIC_BASE_URL (undocumented), so under Auto Mode the proposed action does cross the network to whoever serves the classifier endpoint (Anthropic by default; the user's gateway if a custom base URL is set). The specific mazarax Reddit post anchoring this case is not among the supplied snippets, but the architecture it describes is independently corroborated by first-party docs, the outage error text, and the wire capture. Security context is already hot: Rehberger disclosed a 60โ€“80%-success RCE chain against Auto Mode on August 26, 2026, and CSA frames these as harness-level designs that undermine otherwise-reasonable safety controls.

Why it matters to Scott

Anthropic's Auto Mode puts a probabilistic LLM classifier in the enforcement path for every Claude Code write โ€” exactly the Guardrail Illusion / Compliance Cosplay pattern (probabilistic behaviour sold as a permission boundary), now with dated receipts: fail-closed outage errors, a wire-captured ~107KB classifier call routed via ANTHROPIC_BASE_URL, and Rehberger's 60โ€“80% RCE success against the same gate. The fail-closed remote gate also makes unattended-agent writes hostage to vendor availability and ships every proposed write to Anthropic โ€” bearing on his long-running-agent reliability work, his SiloOS answer to the same gating problem, and his LiteLLM/base-URL routing practice (a concrete check: which permission model and endpoint his own Claude Code sessions actually run under).
ip:concept.guardrail-illusionip:source.compliance-cosplayip:framework.decision-authority-infrastructureip:concept.manners-vs-physicsip:framework.long-running-agentsip:framework.sovereign-software-assurancedev:technology.claude-codedev:project.silo-osdev:technology.litellmradar:claude-code-denied-read-secret-bypassradar:claude-code-remote-attribution-injectionradar:mcp-tool-sequence-guardrail-bypassradar:hosted-open-model-endpoint-failuresradar:runbook-mcp-fail-closed-workflows
queries asked of Scott's wikis
  • agent harness tool-call safety gating fail-closed architecture
  • coding agent permission classifier prompt injection defense
  • local-first privacy vendor observability of agent file writes
  • Anthropic-compatible gateway local open model llama.cpp
  • unattended long-running agent remote service availability dependency
  • Claude Code settings workflows agent-maintained file writes

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-25 04:16โญ origin directly observedClaude Phoning Home before write?
mazarax on r/ClaudeAI
โ€”
09-25 04:16amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wpmq81
mazarax
peak 1 ยท 5 comments ยท 99% of case engagement
09-25 04:20our radar first saw it ยท +0.1hdiscovery anchor: reddit.post.1wpmq81โ€”

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญClaude Phoning Home before write?
ClaudeAI
mazarax15

Interpretation history

Decision trace