2026-10-11 17:10 UTC

llm-benchmarks

band: coolmomentum: stable score: 0.254
temperature history

Episodes (3)

Independent use will determine whether the released ctx-cliff benchmark reproducibly identifies context-length, VRAM-fit, and serving-configuration failure boundaries in local LLM deployments.
expiredknownscott: medium
Benchmark author mauricekleine's Nonobench v1.2 reports that among 43 tested LLMs no open-weight model solves the new 20ร—20 Hard mode (0/10) while GPT-6 Astra posts the first perfect 15ร—15 run (30/30) and Opus 5.5 scores 8/10; if the open-weight shutout holds as open-weight releases are added, it stands as a measured open-frontier gap on long-constrained verifiable reasoning, and any open model clearing 20ร—20 refutes it.
watchingconvergesscott: medium
Open Codenames benchmark gains traction as a cited reference for evaluating LLM reasoning and communication in multi-agent settings.
seednovelscott: low