2026-10-11 16:36 UTC

evaluation-security

band: coolmomentum: stable score: 0.031
temperature history

Episodes (2)

HarnessOpt-Bench’s authors claim their held-out benchmark can measure whether frontier models improve other agents’ harnesses without exploiting test data or grading signals, enabling safer evaluation of recursive agent optimization.
expiredconvergesscott: high
ratctl’s maintainer claims the released static-and-dynamic auditor detects reward-hacking vulnerabilities in RL post-training environments with few false positives, potentially making verifier audits a practical control before agent training.
expiredconvergesscott: medium

Trajectory notes