2026-10-11 16:36 UTC

benchmarking

band: coolmomentum: stable score: 0.157
temperature history

Episodes (4)

Krishna Balasubramanian, Sasha Podkopaev, and Shiva Kasiviswanathan claim their unsupervised Ising-model aggregator improves ten-judge panel accuracy by 9โ€“14% over weighted voting on three tasks, enabling better automated evaluation without human reference labels for training.
seedconvergesscott: medium
Specific Labs claims its newly released Real-SWE benchmark finds tested model-and-harness combinations resolve at most 38.8% of private enterprise tasks, exposing a company-context and cross-service reliability gap relevant to production coding-agent deployment.
seedconvergesscott: medium
DeepSWE-mini creator asankhs claims the released 16-instance subset preserves the full DeepSWE leaderboard's relative model rankings, potentially reducing the cost of routine local coding-agent evaluation without reproducing absolute scores.
seedconvergesscott: medium
Beatnothing creator kaustubhspatil claims none of 25 tested equity-signal strategies demonstrates statistically significant net Sharpe improvement over an equal-weight passive baseline under its cost and leakage controls, challenging apparent model gains that omit trading costs or universe-selection effects.
seedconvergesscott: low