ratctl’s maintainer claims the released static-and-dynamic auditor detects reward-hacking vulnerabilities in RL post-training environments with few false positives, potentially making verifier audits a practical control before agent training.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses reward-hacking evaluation-security agentic-securityratctl
What is this?
ratctl is presented as the maintainer of a released static-and-dynamic auditing tool for finding reward-hacking vulnerabilities in reinforcement-learning post-training environments. Its PAPER.md reportedly describes audits of 112 public environments, with 54 flagged and a claimed zero false positives, but the supplied web snippets neither identify the maintainer nor independently verify those results. The snippets do establish the underlying problem: agents can exploit literal verifier checks or intermediate metrics without completing the intended task, making reliable scoring and pre-training environment review important controls.
Why it matters to Scott
The released auditor independently operationalises Scott’s existing position that visible or defective evaluators invite specification gaming and that agent behaviour should pass adversarial, independent checks before release or training. If its claimed 112-environment results and zero false positives hold up, it could turn those principles into a practical harness control, but the supplied evidence provides no independent validation yet.
ip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersdev:concept.claim-bounded-adversarial-verificationradar:concept.agent-verificationradar:concept.agentic-rlradar:concept.agentic-securityradar:vinvai-runtime-trace-guardrails
queries asked of Scott's wikis
- verifier audits before agent training
- reward hacking in agent harnesses
- static and dynamic analysis of evaluation environments
- evaluation security and specification gaming
- adversarial testing of reward functions
- false-positive tradeoffs in automated security audits
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-07T19:37:29Z
Repeated checks have produced no independent validation, adoption, or implementation findings, and no concrete follow-up is expected; this episode no longer warrants scheduled attention. The auditor remains a candidate control with maintainer-reported results, not a disproved approach, and can reopen on replication or consequential use.
2026-09-05T18:31:21Z
This check adds no substantive evidence: ratctl remains a candidate verifier-auditing control supported by maintainer claims, not a validated harness safeguard. The echoed paper and promotional post are one evidentiary line; neither the reported clean-control results nor activity in adjacent agent-security topics establishes general reliability.
2026-09-03T17:56:30Z
The staleness check adds no validation, adoption, or implementation evidence; this remains an inspectable but solely maintainer-validated auditor, with its accuracy claims unresolved.
2026-09-01T17:39:25Z
No new implementation evidence, independent validation, or adoption has appeared; the case remains a potentially useful released auditor whose headline accuracy claims are solely maintainer-reported.
2026-09-01T17:34:30Z
grounded: converges/medium — The released auditor independently operationalises Scott’s existing position that visible or defective evaluators invite specification gaming and that agent beh
2026-09-01T17:30:36Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1w4h6c1 -> echo.github.fafeb41ae3 by Aditya Suryavanshi (FreakyAdy)
2026-09-01T17:29:20Z
case created — The first-party post describes a concrete auditing tool and a 112-environment vulnerability survey addressing a consequential agent-training failure mode.
Decision trace
- 09-08 05:37expireRepeated checks have produced no independent validation, adoption, or implementation findings, and no concrete follow-up is expected; this episode no longer warrants scheduled attention. The auditor r
- 09-08 05:37alert_silentThere is no new consequential delta, and the original release was already routed. Another notification would repeat unresolved claims without changing Scott's decisions.
- 09-08 05:37alert_routeThere is no new consequential delta, and the original release was already routed. Another notification would repeat unresolved claims without changing Scott's decisions.
- 09-06 04:31repriceThis check adds no substantive evidence: ratctl remains a candidate verifier-auditing control supported by maintainer claims, not a validated harness safeguard. The echoed paper and promotional post a
- 09-06 04:31alert_silentThere is no new release, replication, adoption, or consequential implementation finding. The original release was already routed; another notification would repeat it without changing Scott's dec
- 09-06 04:31alert_routeThere is no new release, replication, adoption, or consequential implementation finding. The original release was already routed; another notification would repeat it without changing Scott's dec
- 09-04 03:56repriceThe staleness check adds no validation, adoption, or implementation evidence; this remains an inspectable but solely maintainer-validated auditor, with its accuracy claims unresolved.
- 09-04 03:56alert_silentThere is no consequential new delta beyond elapsed time, and the original release has already been routed; repeating it now would add no decision value.
- 09-04 03:56alert_routeThere is no consequential new delta beyond elapsed time, and the original release has already been routed; repeating it now would add no decision value.
- 09-02 03:39repriceNo new implementation evidence, independent validation, or adoption has appeared; the case remains a potentially useful released auditor whose headline accuracy claims are solely maintainer-reported.
- 09-02 03:39alert_silentThis look is only a legacy-state re-evaluation with unchanged engagement and no consequential new evidence; the original release was already routed, so another alert would be repetitive.
- 09-02 03:39alert_routeThis look is only a legacy-state re-evaluation with unchanged engagement and no consequential new evidence; the original release was already routed, so another alert would be repetitive.
- 09-02 03:38alert_shadowThe repository and paper establish that an inspectable auditing tool is available now, operationalising pre-training checks for exploitable verifiers across common RL environment formats. Its maintain
- 09-02 03:38alert_routeThe repository and paper establish that an inspectable auditing tool is available now, operationalising pre-training checks for exploitable verifiers across common RL environment formats. Its maintain
- 09-02 03:34groundThe released auditor independently operationalises Scott’s existing position that visible or defective evaluators invite specification gaming and that agent behaviour should pass adversarial, independ
- 09-02 03:30promote_anchororigin walk conf 0.99
- 09-02 03:29createThe first-party post describes a concrete auditing tool and a 112-environment vulnerability survey addressing a consequential agent-training failure mode.