2026-10-11 16:36 UTC

coding-benchmarks

band: coolmomentum: stable score: 0.11
temperature history

Episodes (2)

Autoprompt's publisher claims its released coding skill raised DeepSeek V4 Flash 0731's Terminal-Bench 2.1 success rate from 67.42% to 82.02% in OpenCode, potentially reducing coding-task failures at the expense of longer runs and higher token costs.
seedconvergesscott: medium
Researchers from Meta Superintelligence Labs, Stanford, Harvard, and UW (SWE-bench lineage, led by Kilian Lieret and Ofir Press) released SWE-sweep โ€” 100 repos, 4.1k bugs, where agents must find and fix unreported bugs with no hints โ€” measuring proactive bug discovery at a stark 4.7% best (Sol 5.6 xhigh); leaderboard movement past that level or external adoption as a tracked agent-coding capability establishes proactive bug discovery as a benchmarked frontier, while stagnation marks the current gap as durable.
watchingconvergesscott: medium