2026-10-11 18:04 UTC

Independent testing will determine whether long policy documents such as Handbook.md fail to reliably constrain agent behavior, driving adoption of shorter or tool-enforced controls.

state: resolvedheat: lowuncertainty: highnovelscott: noneagent-safety agent-harnesses policy-enforcement

What is this?

Handbook.md is presented as a test of whether agents can follow lengthy company-policy handbooks while completing real tasks, but the supplied results do not include the underlying paper’s methods, authors, or findings. A secondary benchmark reports that shorter agent instruction files outperform longer ones and notes that some coding-agent systems cap or truncate instruction chains. Separate governance sources argue that documented policies are insufficient without independently testable, runtime or tool-level enforcement, though the snippets do not yet establish that this shift was caused by Handbook.md itself.

Why it matters to Scott

No intersection found: there are no Scott wiki or radar hits establishing that this claim bears on a position, project, or tracked development.
queries asked of Scott's wikis
  • long context instructions versus agent compliance
  • CLAUDE.md AGENTS.md instruction-file design
  • tool-enforced policy versus prompt-based rules
  • agent harness runtime guardrails and permissions
  • state-action graphs for agent control
  • governance tests for coding agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (10) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnHandbook.md shows that long policy documents do not reliably govern agentsspIrr325210
🟧 echo.blog ⭐The original report says HANDBOOK.md tests whether agents can follow long company handbooks across real tasks. It reports: “No frontier modeSurge AI——
🟠 redditI replaced our agent's CLAUDE.md with a POMDP-style state-action graph, +16 to +20pts task success
ClaudeAI
Mysterious-Try-196611
🟠 reddit100 days in: where Claude Code beat everything else I tried, and where I stopped using it.
ClaudeAI
vibecodejoe07
🟠 redditClaude writes code before we fully agree on a plan
ClaudeAI
Square_Reason_6490323
🟠 redditFollow Dr. DYK, because bloat is the sickness and simple is the trick
ClaudeAI
PilgrimOfHaqq06
🟧 hnI Stop LLMs Drifting in Production Codebasesmainsong30
🟠 redditClaude wiped every prod env-var on my Render service and my local .env too 💀
ClaudeAI
Purple-Release5132941
🟠 redditCLAUDE.md for Opus 5 based on Anthropic's official platform docs to fix verbosity and more.
ClaudeAI
Puzzled-Ad-685445463
🟠 redditTIL why my agent.md file was making my Claude Code sessions worse, not better
ClaudeAI
No-Respect-104032

Interpretation history

Decision trace