2026-10-11 16:38 UTC

Anthropic's disclosed numbers say its online monitor blocked about 1 in 47,000 of roughly 1 billion August actions across ~30,000 internal agents (~21,000 blocked actions a month), making quantified production-scale agent-action monitoring a visible frontier-lab safety control that smaller operators currently lack.

state: watchingheat: lowuncertainty: mediumconvergesscott: highagentic-security anthropic agent-safetyAnthropic

What is this?

On 2026-09-17 Anthropic published self-reported oversight metrics for its internal agent fleet: roughly 30,000 agents ran concurrently on its most-used internal platform during August 2026, executing over a billion decisions, all of which passed through an automated online monitor that blocked about 0.002% (1 in ~47,000 โ€” 20,000+ actions a month, each human-reviewed per one writeup). The snippets corroborate the figures but also their limits: single internal platform, self-reported, review-latency numbers drawn from a separate one-week window, and a 'block' is the monitor's judgment call rather than a confirmed harm โ€” outside commentary notes Anthropic is effectively 'marking its own homework,' and the supplied coverage does not establish what the blocked actions actually were. The case also tracks a headline-only HN post ('Nvidia Debuts System Designed to Stop AI Agents from Going Awry') as a second major vendor entering agent containment; the supplied web results contain no Nvidia coverage, so what that system does (in-path monitoring, sandboxing, policy enforcement) remains unverified.

Why it matters to Scott

Anthropic's 1-in-47,000 block rate over ~1B gated actions is a dated receipt that a frontier lab has independently built Scott's Decision-Authority-Infrastructure pattern โ€” in-path gating of every agent action before consequence โ€” while the metrics themselves are exactly his Evidence Class Ladder / Correlated Checkers case: self-reported blocks from a probabilistic monitor judging its own model, with block-vs-confirmed-harm still unclarified. The headline-only Nvidia containment debut extends the convergence (containment productizing beyond indie proxies like Bulwark and Grith) and now bears on SiloOS's 'operators without frontier-scale monitoring need structural containment' market clause โ€” but with zero detail on what Nvidia's system actually enforces, that thread stays a watch item, not a challenge.
ip:framework.decision-authority-infrastructureip:concept.evidence-class-ladderip:concept.correlated-checkers-pitfallip:concept.guardrail-illusionip:framework.siloosradar:concept.agent-securityradar:concept.agent-governanceradar:concept.agent-observabilityradar:concept.policy-enforcementradar:bulwark-agent-security-gatewayradar:grith-coding-agent-security-proxy
queries asked of Scott's wikis
  • decision authority infrastructure โ€” runtime gating and in-path blocking of agent actions
  • governance-as-code policy enforcement layer for coding-agent harnesses
  • evidence class ladder โ€” grading seller-reported, self-attested vendor safety metrics
  • correlated checkers pitfall / guardrail illusion โ€” limits of a single automated monitor
  • SiloOS structural containment pitch for operators without frontier-scale monitoring budgets
  • full-coverage vs sampled agent-action monitoring โ€” cost economics at scale

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 602h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-16 14:00โญ origin echo-reconstructedThe Reddit title paraphrases Anthropic's oversight-of-AI-agents section: "We've built a system that lets us oversee and intervene in actions
Anthropic (co-authored by Marina Favaro and Phillie Wright, with editorial support from Santi Ruiz, Adam Farina, and Sarah Pollack; published by Anthropic PBC) on blog (echo) ยท attributed from reddit.post.1wp1dw4
โ€”
09-24 13:17first on r/ClaudeAI ยท published ยท +191.3hAnthropic's monitor blocks about 1 in 47,000 actions from its internal agents. That still works out to roughly 21,000 a month
InterviewAsleep639
โ€”
09-28 09:38first on hacker news ยท published ยท +283.6hNvidia Debuts System Designed to Stop AI Agents from Going Awry
monkeydust
โ€”
09-24 13:17amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wp1dw4
InterviewAsleep639
peak 1 ยท 4 comments ยท 35% of case engagement
09-28 09:38amplified on hacker newshn.story.49875573
monkeydust
peak 2 ยท 0 comments ยท 25% of case engagement
10-05 19:04amplified on hacker newshn.story.49969048
Betelbuddy
peak 2 ยท 0 comments ยท 25% of case engagement
10-06 12:06amplified on hacker newshn.story.49977252
creator77
peak 1 ยท 0 comments ยท 14% of case engagement
09-17 21:22our radar first saw it ยท +31.4hdiscovery anchor: reddit.post.1wp1dw4โ€”
pace: p44 vs 1032 stories at the 336h mark (now 602h old) โ€” ahead of anthropic-meta-lawsuit (1.2x), behind legion-elixir-lua-agent-sandbox (0.9x)

Evidence (5) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditAnthropic's monitor blocks about 1 in 47,000 actions from its internal agents. That still works out to roughly 21,000 a month
ClaudeAI
InterviewAsleep63914
๐ŸŸง echo.blog โญThe Reddit title paraphrases Anthropic's oversight-of-AI-agents section: "We've built a system that lets us oversee and intervene in actionsAnthropic (co-authored by Marina Favaro and Phillie Wright, with editorial support from Santi Ruiz, Adam Farina, and Sarah Pollack; published by Anthropic PBC)โ€”โ€”
๐ŸŸง hnNvidia Debuts System Designed to Stop AI Agents from Going Awrymonkeydust20
๐ŸŸง hnLLMs Learn to Evade Latent Monitors from Prior Feedback AloneBetelbuddy20
๐ŸŸง hnMonitoring an AI agent trains it to evade the monitor (100-day simulation)creator7710

Interpretation history

Decision trace