2026-10-11 16:38 UTC

Independent replications will determine whether conventional approve-or-deny prompts cause users to approve a substantial share of dangerous coding-agent commands and whether stronger permission controls materially reduce those misses.

state: significantheat: lowuncertainty: mediumconvergesscott: highcoding-agent-security human-in-the-loop agent-permissionsWirbelwindScaleX
Surfaced 2026-08-28T15:40:11Z — priced heat=high at reprice: A headline now claims Claude Code Opus 5 Auto Mode can be broken, creating a potentially consequential post-rollout test of automated permission handling, but no mechanism, affected configuration, reproduction, or impact is supplied. The AST-based command reviewer is another concrete control implementation without comparative safety results.

What is this?

The case tracks the practitioner and vendor debate over whether approve-or-deny prompts are a workable security boundary for coding agents, and whether stronger permission controls actually do better. The miss-rate magnitudes rest on two non-independent sources — Scale X's self-described permission game (~40,000 runs, ~409,000 decisions, roughly one in three dangerous commands approved; Scale X's own blog concedes several prompts were legitimately ambiguous, e.g. 45.9% approving `cat ~/.zshrc`) and Anthropic's vendor-reported human-vs-classifier figures used to justify making Claude Code Auto Mode the default on Aug 14 — while the supplied web material adds broad qualitative agreement from security firms (NCC Group, Zenity, Telerik, Fiddler) that per-action prompts are insufficient on their own, favoring sandboxing, task-scoped allowlists, and graded checks, but no independent replication of the numbers. The newest turn is a hook author's (Wirbelwind's) public self-correction documenting that Claude Code PreToolUse hooks fail open by contract — only exit code 2 blocks, so crashes, timeouts, and malformed output drop the verdict and the command proceeds — upgrading the dominant DIY control class's failure mode from anecdote to documented platform semantics. One mild counter-current in the supplied material: security vendor Manifold still markets per-action human approval ('Ask') as a product, unchanged by the fatigue consensus.

Why it matters to Scott

Converges at the doctrine's sharpest edge: Wirbelwind's public self-correction that Claude Code PreToolUse hooks fail open by contract (only exit code 2 blocks; crashes, timeouts and malformed verdicts drop it and the command runs) is a dated receipt for Architecture Not Vibes' 'a gate is defined by its failure path, not its happy path' claim, and extends Guardrail Illusion into deterministic-tool territory — mechanically enforced code with fail-open semantics is still a probability barrier, not a permission boundary. It directly bears on Ask, whose complex-work approval is still behavioural: if Scott mechanicalizes it, this documents that explicit fail-closed handling (deny on timeout/malformed/crash) is the required primitive rather than an option, and sharpens the contrast with his own guarded-agent-inbox, where failing open is a deliberate low-stakes choice for mail screening that would be fatal for command gating.
ip:framework.architecture-not-vibesip:concept.guardrail-illusionip:framework.decision-authority-infrastructureip:framework.siloosdev:project.askdev:concept.padded-cell-agent-architecturedev:concept.guarded-agent-inboxradar:concept.deterministic-guardrailsradar:concept.policy-enforcementradar:concept.human-in-the-loopradar:concept.agent-authorizationradar:concept.agent-containmentradar:runbook-mcp-fail-closed-workflowsradar:bulwark-agent-security-gateway
queries asked of Scott's wikis
  • fail-closed vs fail-open enforcement in agent hooks
  • human approval as a security boundary approval fatigue
  • scoped capabilities least-privilege agent permissions SiloOS
  • architecture not vibes mechanical vs behavioral enforcement
  • sandboxing and containment for coding agent harnesses
  • allowlist deny-list bypass deterministic command gating

Measured heat

now 0 pts/hpeak 11 pts/hcomments 0/hpeers p23momentum: steady3 platformsage 1588h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-06 12:22 (minted)⭐ origin echo-reconstructedAnalysis of 409,000 approve-or-deny decisions from roughly 40,000 runs of a coding-agent permission game found that participants missed abou
Wirbelwind on blog (echo) · attributed from reddit.post.1vh1y03, hn.story.49195468 · published time unknown
—
08-06 11:47first on r/ClaudeAI · published · lag ?Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays
Wirbelwind
—
08-06 11:58first on hacker news · published · lag ?Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Wirbelwind
—
08-18 02:11first on r/OpenAI · published · lag ?The AI review that said "looks good, solid implementation" and approved a bug that took down production three days later
ClickOk5811
—
08-18 16:32first on r/singularity · published · lag ?And Samsung has started using Anthropic’s Claude Code for chip design, reportedly compressing a month of work into two days, but...
mvandemar
—
08-18 19:50first on r/LocalLLaMA · published · lag ?I built a tool to constrain model access locally
erratic_parser
—
08-06 11:47amplified on r/ClaudeAIreddit.post.1vh1y03
Wirbelwind
peak 111 · 10 comments · 2% of case engagement
08-06 11:58amplified on hacker newshn.story.49195468
Wirbelwind
peak 336 · 243 comments · 16% of case engagement
08-06 19:07amplified on hacker newshn.story.49200925
sbulaev
peak 2 · 0 comments · 0% of case engagement
08-07 16:19amplified on r/ClaudeAIreddit.post.1vi54ks
BugeacAlexandru
peak 1 · 4 comments · 0% of case engagement
08-07 19:11amplified on hacker newshn.story.49214994
tosh
peak 19 · 22 comments · 1% of case engagement
08-08 11:35amplified on hacker newshn.story.49220827
maxloh
peak 18 · 4 comments · 1% of case engagement
92 more amplifiers in ainews.case_chain
08-06 12:20our radar first saw it · lag ?discovery anchor: reddit.post.1vh1y03—
08-28 15:40reached heat=high · lag ? · via ledger——

Evidence (99) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditHumans missed 1 in 3 threats approving AI agent commands across 40,000 plays
ClaudeAI
Wirbelwind11110
🟧 hnHumans missed 1 in 3 threats approving AI agent commands across 40k game runsWirbelwind336243
🟧 echo.blog ⭐Analysis of 409,000 approve-or-deny decisions from roughly 40,000 runs of a coding-agent permission game found that participants missed abouWirbelwind——
🟧 hnHumans in the loop miss a third of dangerous AI coding agent requestssbulaev20
🟠 redditI open-sourced the Claude Code setup that helps me work faster and more safely
ClaudeAI
BugeacAlexandru04
🟧 hnClaude Code: Starting August 14, auto mode will be the default permission modetosh1922
🟧 hnAuto Mode will be the default in Claude Code – because humans can't be trustedmaxloh184
🟠 redditAnthropic Flips Claude Code to Auto Mode by Default Aug 14, after finding AI blocks 80%+ dangerous queries while humans only 14%
ClaudeAI
Justgototheeffinmoon1261189
🟧 hnShow HN: RunOnMine – local policy and approvals for AI access to your machineademisler20
🟠 reddit'Dangerously skip permissions'
ClaudeAI
forsaken3400031
🟧 hnShow HN: dbward – Approval workflows for production DBs, built for AI agentsmetapox20
🟧 hnA least-privilege linter for Claude Code agent policiesjdubray10
🟠 redditDid you seen?
ClaudeAI
dradn4ts11
🟧 hnPrompt Injection Experiments with Opus-5 in Claude Code – Auto-Mode Editionveganmosfet31
🟠 redditWebfetch risks
ClaudeAI
Sagnet015
🟠 redditAuto mode goes default tomorrow (Aug 14) — anyone actually stress-tested it yet?
ClaudeAI
mike_aitrends013
🟧 hnShow HN: A UI to see all the assumptions your coding agents are makingnikisweeting21
🟧 hnShow HN: OpenCode Auto Permissions – automatic review of permission promptshuey7720
🟧 hnAgent-kit – a size-based mandatory review chain for Claude Codedanielhorilla10
🟠 redditKind of burnt out reviewing every single line claude code writes
ClaudeAI
Emergency_Mobile70159076
🟧 hnThe road to Seahaven, or how I run my agent harnesses without permission promptspmw20
🟠 redditI built a spend kill-switch for the cloud commands Claude Code runs — the aws ec2 run-instances, not the tokens
ClaudeAI
Substantial-Fuel-519114
🟧 hnShow HN: A pre-execution guard that stops AI agents running destructive commandsandevandith21
🟠 redditThe AI review that said "looks good, solid implementation" and approved a bug that took down production three days later
OpenAI
ClickOk581105
🟧 hnShow HN: Cermet – allow github.push where owner = "suarezc" and name = "cermet"notshore11
🟠 redditAnd Samsung has started using Anthropic’s Claude Code for chip design, reportedly compressing a month of work into two days, but...
singularity
mvandemar50369
🟠 redditI built a tool to constrain model access locally
LocalLLaMA
erratic_parser08
🟧 hnWhen deny doesn't win: least-privilege permissions for Claude Codesixthsense10
🟧 hnA merge gate for coding agents cannot be a booleanfskroes_is_me20
🟧 hnShow HN: Do-over, undo for AI agent shell commandsCayden2721
🟧 hnAI agent suggested installing a malware package. Engineer almost took its advicesbulaev30
🟠 redditYour `deny Bash(aws:*)` rule is either too strict or too loose. There's a third option.
ClaudeAI
OneOfBenders02
🟠 redditAnyone else experiencing wild levels of overreach?
ClaudeAI
memetican450
🟠 redditThe danger of auto approval and default being on?
ClaudeAI
bregottextrasaltat19
🟠 redditWhy did you delete the database?
ClaudeAI
helu_ca06
🟠 redditWelp thats just great !
ClaudeAI
toxic_prince21374100
🟠 redditBuilding a personal AI agent with an approval gate. I don't really know AI security, what am I getting wrong?
ClaudeAI
UniversityFuzzy620917
🟠 redditClaude wiped my friend's production .env (definitely not mine). I wrote him a guide.
ClaudeAI
FeralFancyBop26
🟧 hnShow HN: AgentMachinist, issue to reviewed PR with a SHA-bound spec approvalvscarpenter10
🟠 redditRead only website access
ClaudeAI
misanthropic78928
🟠 redditHow do I use claude code as a mid level developer?
ClaudeAI
Sashinkis16
🟠 redditWhat belongs in the binary vs. what can only live in the prompt
ClaudeAI
Dynmiwang11
🟠 redditAnthropic published an AI-native SDLC playbook. The interesting part isn't the six stages, it's what replaces line-by-line review
ClaudeAI
Forward_Mind68869250
🟠 redditFirst time getting hosed by Claude Code
ClaudeAI
Z7N6Qo139
🟧 hnLocus – A small, deterministic safety barrier for AI coding agentsahmadshadi200410
🟧 hnAuto Review for shell commands with AST parsing and a subagentyags50
🟧 hnBreaking Claude Code Opus 5 Auto ModeBerislavLopac10
🟠 redditClaude committing code without asking
ClaudeAI
Ffilib015
🟧 hn80% Prompt Injection Success Rate Against Claude Auto Modegr_norm30
🟧 hnBreaking Claude Code Opus 5 Auto Modejaksa72
🟧 hnBreaking Claude Code Opus 5 Auto ModeRecursing396119
🟧 hnAn AI coding agent silently erased 92% of AI nodes in n8n's most-cited datasetahunia10
🟠 redditI got tired of babysitting Claude Code, so I built an extension to auto-accept commands (with safety guardrails)
ClaudeAI
Rait7020
🟧 hnWhen Claude Code went rogue, years of Bengaluru heritage work disappeareddsr12179
🟠 redditThe part of Claude Code that concerns me most is how easy it is to approve work I only half understand
ClaudeAI
Pretend_Sell65924384
🟠 redditCodex deleted test cases that uncovered a bug to “Prioritize getting the test cases to pass”…
OpenAI
PuzzleheadedAnt95036424
🟧 hnWhat happens when your AI agent edits its own tests to pass?itsub_sa20
🟠 redditClaude code almost wiped out a week's worth of my work!
ClaudeAI
OxyMC016
🟠 redditCodex got me a birthday present. I paid for it.
OpenAI
MantheaLabs013
🟧 hnAsk HN: How do you gate an autonomous coding agent's shell access?alanfuNZ23
🟧 hnExecution gating and micro-rollbacks for AI agentsitsub_sa10
🟠 redditThe "I don't know, Claude wrote this" pandemic
ClaudeAI
fagnerbrack06
🟧 hnWhy AI coding agents fake completion, and how to build a bijective validatortanjianbo10
🟧 hn98% of What Claude Code Does, I Never See. The 0.27% It Stops for Is the Pointedf1310
🟧 hnShow HN: ActraDeck – Put risky coding-agent actions back in front of a humanTKMD10
🟧 hnShow HN: DashClaw – policy and approval layer for unattended coding agentspracticalsystem10
🟠 redditComprehensive Claude Code Permission Guard (settings.json)
ClaudeAI
Puzzled-Ad-685423
🟠 redditI built a tiny CLI that finds Claude Code jobs in CI that will hang (or silently fail) on permission prompts
ClaudeAI
Obluness31
🟧 hnDeterministic Rule based auto-approver for Claude/Codexaditaymatt11
🟠 redditRules for what reviewing the AI misses, after months of building a real codebase with Claude
ClaudeAI
FirstOrDefault22
🟧 hnShow HN: Opair, a coding harness that eschews autonomyphilbo20
🟧 hnShow HN: GuardRail, shell guards that stop Claude Code before it pushes to mainpromptandbuild50
🟧 hnI Let Claude Merge 38 PRs. Here's What Brokecraigphares22
🟠 redditI turned off auto mode's classifier, ran 255 evasion cases against my "snapshot before any system change" hook, and 100 got through. Here's what a PreToolUse hook has to handle when it's the only safety layer
ClaudeAI
No_Diet_114508
🟠 redditClaude Code deleted 2,000+ files from my Dropbox (including my dissertation). Dropbox the day. Back up your stuff :)
ClaudeAI
DrEvaWolf050
🟧 hnShow HN: I bypassed my Claude Code deny-list 8 ways; only an allow-list heldglitchbound10
🟠 redditSix code reviews said "request changes". My board recorded six approvals. The model was right every time.
ClaudeAI
Julien_Builds05
🟧 hnShow HN: Vigilator – human-in-the-loop layer for AI agentsnoahhafn30
🟠 redditMy agent deleted the file that was stopping it from merging its own PRs
ClaudeAI
echowrecked04
🟧 hnMiSeGuard – Deterministic runtime safety layer for AI coding agentsmisetech20
🟧 hnShow HN: Review agent changes locally before pushingGrinningFool20
🟧 hnSub-ms deterministic parsing vs. LLM-based policy for agent safety?misetro20
🟧 hnShow HN: GuardRail – 13 guards that stop Claude Code before it pushes to mainpromptandbuild11
🟠 redditI built a small tool that checks what Claude Code actually changed vs what I asked, before I commit
ClaudeAI
sandeepwastaken06
🟧 hnShow HN: Pi-jev-auto-mode – a probability model gates Pi's shell commandsjomatsu10
🟠 redditClaude Code's overreach is getting a bit severe
ClaudeAI
therealwench16365
🟠 redditClaude destroyed my entire project and home directory while adding a simple delete feature
ClaudeAI
FeatureCurrent9416017
🟠 redditClaude Code Safety Classifier Not Overridable
ClaudeAI
ReckonCrew03
🟧 hnTell HN: Claude Code just accepted and signed a contract for me. Without askingfranze5198
🟠 redditHow do you run scheduled Claude Code jobs that need browser automation, without granting blanket tool permissions?
ClaudeAI
ayushgupta0610033
🟠 redditOpus 5.5 implemented a feature in a single prompt, but also introduced a major security vulnerability.
ClaudeAI
LoverOfCoding04
🟧 hnAI: Who Still Uses Permissions?meredithbloom21
🟠 redditBe honest… how many permission prompts does your coding agent get before you stop reading them?
ClaudeAI
iamjessew356
🟧 hnShow HN: Guardrails for Claude Code: blocks rm -RF, reports what loadedahmadalstaty10
🟧 hnWhen Every AI Agent Action Is Authorized–and the Overall Decision Is Still Wrongtasinsight10
🟧 hnShow HN: AgentMachinist: make your coding agent show its workvscarpenter20
🟠 redditCorrection to my hook post from last week: it fails open. The fix, and 3 more ways a hook lets rm -rf through.
ClaudeAI
Individual-Shower973010
🟠 redditWhy does claude asks for my permission to do anything. STOP
ClaudeAI
Boom511107
🟠 redditHow much context should an approval prompt give before an AI agent takes action?
OpenAI
Sumsub_Insights15

Interpretation history

Decision trace