Anthropic's disclosed numbers say its online monitor blocked about 1 in 47,000 of roughly 1 billion August actions across ~30,000 internal agents (~21,000 blocked actions a month), making quantified production-scale agent-action monitoring a visible frontier-lab safety control that smaller operators currently lack.
state: watchingheat: lowuncertainty: mediumconvergesscott: highagentic-security anthropic agent-safetyAnthropic
What is this?
On 2026-09-17 Anthropic published self-reported oversight metrics for its internal agent fleet: roughly 30,000 agents ran concurrently on its most-used internal platform during August 2026, executing over a billion decisions, all of which passed through an automated online monitor that blocked about 0.002% (1 in ~47,000 โ 20,000+ actions a month, each human-reviewed per one writeup). The snippets corroborate the figures but also their limits: single internal platform, self-reported, review-latency numbers drawn from a separate one-week window, and a 'block' is the monitor's judgment call rather than a confirmed harm โ outside commentary notes Anthropic is effectively 'marking its own homework,' and the supplied coverage does not establish what the blocked actions actually were. The case also tracks a headline-only HN post ('Nvidia Debuts System Designed to Stop AI Agents from Going Awry') as a second major vendor entering agent containment; the supplied web results contain no Nvidia coverage, so what that system does (in-path monitoring, sandboxing, policy enforcement) remains unverified.
Why it matters to Scott
Anthropic's 1-in-47,000 block rate over ~1B gated actions is a dated receipt that a frontier lab has independently built Scott's Decision-Authority-Infrastructure pattern โ in-path gating of every agent action before consequence โ while the metrics themselves are exactly his Evidence Class Ladder / Correlated Checkers case: self-reported blocks from a probabilistic monitor judging its own model, with block-vs-confirmed-harm still unclarified. The headline-only Nvidia containment debut extends the convergence (containment productizing beyond indie proxies like Bulwark and Grith) and now bears on SiloOS's 'operators without frontier-scale monitoring need structural containment' market clause โ but with zero detail on what Nvidia's system actually enforces, that thread stays a watch item, not a challenge.
ip:framework.decision-authority-infrastructureip:concept.evidence-class-ladderip:concept.correlated-checkers-pitfallip:concept.guardrail-illusionip:framework.siloosradar:concept.agent-securityradar:concept.agent-governanceradar:concept.agent-observabilityradar:concept.policy-enforcementradar:bulwark-agent-security-gatewayradar:grith-coding-agent-security-proxy
queries asked of Scott's wikis
- decision authority infrastructure โ runtime gating and in-path blocking of agent actions
- governance-as-code policy enforcement layer for coding-agent harnesses
- evidence class ladder โ grading seller-reported, self-attested vendor safety metrics
- correlated checkers pitfall / guardrail illusion โ limits of a single automated monitor
- SiloOS structural containment pitch for operators without frontier-scale monitoring budgets
- full-coverage vs sampled agent-action monitoring โ cost economics at scale
Measured heat
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 602h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p44 vs 1032 stories at the 336h mark (now 602h old) โ ahead of anthropic-meta-lawsuit (1.2x), behind legion-elixir-lua-agent-sandbox (0.9x)
Evidence (5) โ โญ canonical anchor
Interpretation history
2026-10-06T14:48:43Z
A second headline-only monitor-evasion item (100-day simulation of monitoring training evasion) turns the durability critique from a lone unverified paper into a two-signal research direction โ Guardrail Illusion is now an actively forming threat model against runtime gating rather than a grading caveat. Neither item is substantiated, and neither is shown to apply to in-path blocking monitors (Anthropic's design) rather than latent/outcome-only ones, so the case stays watching at low heat instead of promoting.
2026-10-06T13:34:27Z
evidence attached: hn.story.49977252 โ Simulation evidence that monitor-aware agents learn to evade monitors directly challenges whether Anthropic's quantified monitor blocking is an effective safety control โ re-judging that case must weigh it.
2026-10-05T21:55:07Z
The monitor-evasion paper shifts the case's critique thread from evidence quality (can you trust the 1-in-47,000 number?) to control durability (can the monitor itself be learned around, given that a blocking monitor is itself a feedback source?). Guardrail Illusion moves from a grading caveat to a mechanistic threat model for in-path monitoring โ but with a 1-point, zero-comment headline and no paper detail, this is a third watch thread alongside the Anthropic metrics and the unverified Nvidia system, not a revision of the core assessment.
2026-10-05T20:39:22Z
evidence attached: hn.story.49969048 โ Paper showing LLMs learn to evade latent monitors from feedback alone materially contextualizes whether monitor-based agent oversight is a durable control.
2026-09-28T10:40:07Z
grounded: converges/high โ Anthropic's 1-in-47,000 block rate over ~1B gated actions is a dated receipt that a frontier lab has independently built Scott's Decision-Authority-Infrastructu
2026-09-28T10:30:02Z
The Nvidia agent-containment debut turns this from a single-disclosure observation into a two-thread story: containment is beginning to productize, which both corroborates the runtime-gating convergence and starts to erode the case's 'smaller operators currently lack it' clause. Whether Nvidia's system is real in-path monitoring with quantified results, or containment branding, is now the open question keeping the case live.
2026-09-28T10:25:09Z
evidence attached: hn.story.49875573 โ A major vendor productizing agent containment is material context for the case's claim that quantified agent-action monitoring is a frontier-lab control smaller operators lack.
2026-09-24T13:53:37Z
origin walked (opencode/cheap-glm, conf 0.92): anchor reddit.post.1wp1dw4 -> echo.blog.fb225a3bcd by Anthropic (co-authored by Marina Favaro and Phillie Wright, with editorial support from Santi Ruiz, Adam Farina, and Sarah Pollack; published by Anthropic PBC)
2026-09-24T13:43:37Z
grounded: converges/high โ Anthropic has independently arrived at Scott's runtime-gating position: routing 100% of ~1B agent actions through an in-path automated boundary that blocks pre-
2026-09-24T13:38:06Z
case created โ First quantified production agent-monitor figures are a genuinely new claim, separate from the R&D-share index case despite coming from the same disclosure.
Decision trace
- 10-08 00:59attention_routeThe editor compared this story and chose to keep watching.
- 10-07 01:48repriceA second headline-only monitor-evasion item (100-day simulation of monitoring training evasion) turns the durability critique from a lone unverified paper into a two-signal research direction โ Guardr
- 10-07 00:34attachSimulation evidence that monitor-aware agents learn to evade monitors directly challenges whether Anthropic's quantified monitor blocking is an effective safety control โ re-judging that case mus
- 10-07 00:34propose_attachSimulation evidence that monitor-aware agents learn to evade monitors directly challenges whether Anthropic's quantified monitor blocking is an effective safety control โ re-judging that case mus
- 10-06 08:55repriceThe monitor-evasion paper shifts the case's critique thread from evidence quality (can you trust the 1-in-47,000 number?) to control durability (can the monitor itself be learned around, given th
- 10-06 07:39attachPaper showing LLMs learn to evade latent monitors from feedback alone materially contextualizes whether monitor-based agent oversight is a durable control.
- 10-06 07:34propose_attachPaper showing LLMs learn to evade latent monitors from feedback alone materially contextualizes whether monitor-based agent oversight is a durable control.
- 09-28 20:40repriceThe Nvidia agent-containment debut turns this from a single-disclosure observation into a two-thread story: containment is beginning to productize, which both corroborates the runtime-gating convergen
- 09-28 20:40groundAnthropic's 1-in-47,000 block rate over ~1B gated actions is a dated receipt that a frontier lab has independently built Scott's Decision-Authority-Infrastructure pattern โ in-path gating of
- 09-28 20:25attachA major vendor productizing agent containment is material context for the case's claim that quantified agent-action monitoring is a frontier-lab control smaller operators lack.
- 09-28 20:25propose_attachA major vendor productizing agent containment is material context for the case's claim that quantified agent-action monitoring is a frontier-lab control smaller operators lack.
- 09-26 08:53review_screenNew comments only re-discuss the already-assessed disclosure: one corrects another commenter's arithmetic on the same 1-in-47,000 rate, another opines that the rate reflects Anthropic's perm
- 09-26 08:52review_screenjev screen borderline (noul=0.56) โ luna review
- 09-24 23:53keep_distinctSame source document, but genuinely different episodes: case B tracks the R&D automation index (how much of Anthropic's model development is agent-executed, resolving on updated index reading
- 09-24 23:53promote_anchororigin walk conf 0.92
- 09-24 23:43groundAnthropic has independently arrived at Scott's runtime-gating position: routing 100% of ~1B agent actions through an in-path automated boundary that blocks pre-execution is Decision Authority Inf
- 09-24 23:38createFirst quantified production agent-monitor figures are a genuinely new claim, separate from the R&D-share index case despite coming from the same disclosure.