2026-10-11 17:09 UTC

benchmarks

band: coolmomentum: stable score: 0.171
temperature history

Episodes (6)

CVE-Bench's publishers present a benchmark for evaluating AI agents' ability to exploit web vulnerabilities, potentially giving builders a task-specific measure of offensive agent capability.
expiredconvergesscott: medium
Independent use will determine whether ExtractBench provides reproducible schema-extraction evaluations that reveal meaningful reliability differences among models and extraction systems.
expiredknownscott: medium
John Sous and coauthors claim expert repairs and regrading reveal near-saturation of retained physics benchmark questions by frontier models, undermining low leaderboard scores as evidence of weak closed-form physics capability.
watchingconvergesscott: medium
Raycaster's released Biopharma Bench V0.1 is headlined as showing open-weight DeepSeek agents beating OpenAI's GPT-6 Sol on autonomous drug-development tasks β€” though its retrieved clinical-hold task page shows GPT-6 Astra as the only passing model with DeepSeek V4.1 Flash failing β€” so the full leaderboard either establishes a real open-vs-frontier agent narrowing in a specialized domain or exposes headline overreach.
seedconvergesscott: medium
SecondState's ex-EY team claims its released FAB benchmark β€” 50 tasks, 160 documents and 231 grading criteria in a synthetic data room, with published traces β€” shows frontier agents finding relevant financial facts but failing to carry them through to complete, reliable due-diligence analysis; adoption by evaluators and labs, or expansion to more companies and models, would make FAB the reference benchmark for long-horizon financial agent work.
seedconvergesscott: medium
Gamow Labs claims its released LabBench β€” 20 held-out wet-lab decision tasks built from real drug-discovery and genomics records β€” shows five frontier agents pass only ~40% of decision criteria, fail every criterion on which experiment should come first, and recover on failed decisions only when a one-sentence attention redirect is appended, and adoption by AI-for-science evaluators would establish it as the reference benchmark for whether agents can decide the next wet-lab experiment.
seedconvergesscott: medium

Trajectory notes