2026-10-11 17:11 UTC

ratctl’s maintainer claims the released static-and-dynamic auditor detects reward-hacking vulnerabilities in RL post-training environments with few false positives, potentially making verifier audits a practical control before agent training.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses reward-hacking evaluation-security agentic-securityratctl

What is this?

ratctl is presented as the maintainer of a released static-and-dynamic auditing tool for finding reward-hacking vulnerabilities in reinforcement-learning post-training environments. Its PAPER.md reportedly describes audits of 112 public environments, with 54 flagged and a claimed zero false positives, but the supplied web snippets neither identify the maintainer nor independently verify those results. The snippets do establish the underlying problem: agents can exploit literal verifier checks or intermediate metrics without completing the intended task, making reliable scoring and pre-training environment review important controls.

Why it matters to Scott

The released auditor independently operationalises Scott’s existing position that visible or defective evaluators invite specification gaming and that agent behaviour should pass adversarial, independent checks before release or training. If its claimed 112-environment results and zero false positives hold up, it could turn those principles into a practical harness control, but the supplied evidence provides no independent validation yet.
ip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersdev:concept.claim-bounded-adversarial-verificationradar:concept.agent-verificationradar:concept.agentic-rlradar:concept.agentic-securityradar:vinvai-runtime-trace-guardrails
queries asked of Scott's wikis
  • verifier audits before agent training
  • reward hacking in agent harnesses
  • static and dynamic analysis of evaluation environments
  • evaluation security and specification gaming
  • adversarial testing of reward functions
  • false-positive tradeoffs in automated security audits

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI audited 112 real RL post-training environments for reward-hacking vulnerabilities — 54 flagged, 0 false positives [OC, tool] [P]
MachineLearning
Responsible_Goose53500
🟧 echo.github ⭐The linked repository’s PAPER.md is the primary artifact. It reports an audit of “112 public reinforcement learning post-training environmenAditya Suryavanshi (FreakyAdy)——

Interpretation history

Decision trace