Benchmark author mauricekleine's Nonobench v1.2 reports that among 43 tested LLMs no open-weight model solves the new 20Γ20 Hard mode (0/10) while GPT-6 Astra posts the first perfect 15Γ15 run (30/30) and Opus 5.5 scores 8/10; if the open-weight shutout holds as open-weight releases are added, it stands as a measured open-frontier gap on long-constrained verifiable reasoning, and any open model clearing 20Γ20 refutes it.
state: watchingheat: lowuncertainty: mediumconvergesscott: mediumllm-benchmarks open-weight-modelsmauricekleine
What is this?
Nonobench is a reproducible community benchmark run by independent developer Maurice Kleine that scores LLMs on nonogram puzzles β one attempt per puzzle, no tools, answers machine-checked against the clues β with a public leaderboard at v1.2 covering 43 models. Its headline claims: no open-weight model has solved any 20Γ20 Hard-mode puzzle (0/10; best open entry DeepSeek V4 Pro sits at 83.3% overall), while OpenAI's GPT-6 Astra posted the first perfect 15Γ15 run and Anthropic's Opus 5.5 leads Hard at 8/10 β a shutout that held in a 2026-09-30 follow-up in which 11 of 15 models (including Astra, at 5/10) solved zero 20Γ20s. The supplied web results corroborate the landscape but not the benchmark: the model names and release dates (GPT-6 Astra 2026-09-03, GPT-6.1 Sol 2026-09-29, Opus/Sonnet 5.5 in late September) match release trackers, and independent 2026 reporting (Lumiere, NerdLevelTech, Artificial Analysis figures) measures a persistent open-vs-closed frontier gap that is widest on long-horizon and written-from-scratch benchmarks, with harness effects flagged as a known confound. But zero coverage of Nonobench or Kleine surfaces anywhere, so the specific 0/10 shutout remains a single-source claim β falsified if any open model clears 20Γ20 raw, while a harnessed-only clear would instead refute the 'open models can't' reading.
Why it matters to Scott
The Hard-mode follow-up sharpens this from an open-gap data point into an instrument for Scott's benchmark-unit argument: the case's own raw-vs-harnessed falsification split operationalizes his model-plus-harness unit claim (with Nonobench's tool-free, one-attempt design as the wrong-unit counterfactual), and Opus taking 8/10 at 'high' while Sol (max) and Astra (xhigh) trail at 7 and 5 is measured evidence that reasoning-effort burn doesn't convert on constraint puzzles. Stays medium β still single-source and quiet, consequential only on a trigger (any open 20Γ20 clear, and especially a harnessed-only one, which would be dated receipts for the harness-unit position over the 'open models can't' reading).
ip:concept.model-plus-harness-benchmark-unitip:concept.benchmarking-the-wrong-unitip:concept.inference-time-scalingip:concept.capability-symmetryradar:concept.open-modelsradar:concept.open-weight-modelsradar:concept.benchmark-integrityradar:concept.llm-evaluationradar:pcss-zebra-puzzle-reasoning-transferradar:minizinc-mcp-solver-toolsradar:ship-harness-benchradar:claude-sonnet-55-token-economicsradar:qwen38-27b-reasoning-effort
queries asked of Scott's wikis
- open-weight vs frontier capability gap position
- harness or scaffold as part of the model eval
- reproducible verifiable LLM benchmark design
- benchmark contamination fresh held-out evals
- LLM constraint satisfaction puzzle reasoning limits
- reasoning token cost per solved task economics
Measured heat
now 0 pts/hpeak 11 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 386h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p65 vs 1032 stories at the 336h mark (now 386h old) β ahead of gfx906-expert-pool-admission-fix (1.0x), behind ordewell-editable-multi-runner-plans (1.0x)
Evidence (4) β β canonical anchor
Interpretation history
2026-10-04T08:24:24Z
Open-sourcing converts the single-source watch into an inspectable one: the method (Standard 30 puzzles 5Γ5β15Γ15 from the CC BY 4.0 dataset; Hard ten unique-solution 20Γ20s, five not line-logic-solvable) is now fully replicable, so an outside run is the reachable second evidence line rather than a hypothetical. Roster 43β49 with headline results unchanged extends the open-weight 20Γ20 shutout under more models. Still zero outside coverage or independent replication and no periphery expansion β the new post itself scored 1 β so it stays a quiet falsifiable watch, not corroborated.
2026-10-04T08:23:29Z
evidence attached: reddit.post.1wxa2bs β Same benchmark now public and open source with roster expanded 43β49 and headline results unchanged β adoption/scrutiny substrate the case's resolution depends on.
2026-09-30T14:46:34Z
grounded: converges/medium β The Hard-mode follow-up sharpens this from an open-gap data point into an instrument for Scott's benchmark-unit argument: the case's own raw-vs-harnessed falsif
2026-09-30T14:37:34Z
Same-author Hard-mode follow-up extends the series: 15 models on the 20Γ20s, 11 solve zero, Opus 5.5 still #1 (8/10), new GPT-6.1 Sol 7/10, and Astra's perfect 15Γ15 doesn't transfer (5/10 Hard) β the shutout survived its first extension to the newest frontier models, but there is still no open model close, no second evidence line, and no periphery expansion. Moved seedβwatching because the benchmark is now demonstrably an actively iterating series (January post β v1.2 β Hard-mode update within days), which is what makes the falsification watch actionable; evidence itself remains single-source.
2026-09-30T14:25:24Z
evidence attached: reddit.post.1wu5w1c β Independent update from the same benchmark author's Nonobench Hard-mode runs (Opus 5.5 8/10, Sol 7/10, Astra 5/10) plus Sonnet 5.5 long-thinking cost parity β direct evidence for the open-weight gap case and its token-economics angle.
2026-09-28T01:42:23Z
Engagement-only update: two comment churns and a transient 4.5x velocity spike that has already decayed (now 0.67 pts/h, cooling). No new model runs, no independent corroboration, and the 20x20 open-weight shutout stands unchanged β the case's meaning is unchanged; it remains a quiet falsifiable watch with no periphery expansion.
2026-09-27T11:40:52Z
origin walked (opencode/cheap-glm, conf 0.78): anchor reddit.post.1wrh7b2 -> echo.other.a4b7d4b3b2 by Maurice Kleine
2026-09-27T11:33:54Z
grounded: converges/medium β Independently measured (though single-source β grounding found no coverage of Nonobench or mauricekleine, so the 0/10 is a claim to watch, not a fact) confirmat
2026-09-27T11:25:26Z
case created β Second release of a reproducible community benchmark with a specific falsifiable open-vs-frontier result and an active maintainer iterating per model additions; no open case tracks Nonobench.
Decision trace
- 10-07 06:29review_screenjev screen: no material development (noul=0.07)
- 10-07 06:24sensor_dirtycomment_update
- 10-04 19:24repriceOpen-sourcing converts the single-source watch into an inspectable one: the method (Standard 30 puzzles 5Γ5β15Γ15 from the CC BY 4.0 dataset; Hard ten unique-solution 20Γ20s, five not line-logic-solva
- 10-04 19:23attachSame benchmark now public and open source with roster expanded 43β49 and headline results unchanged β adoption/scrutiny substrate the case's resolution depends on.
- 10-04 19:23propose_attachSame benchmark now public and open source with roster expanded 43β49 and headline results unchanged β adoption/scrutiny substrate the case's resolution depends on.
- 10-01 00:46repriceSame-author Hard-mode follow-up extends the series: 15 models on the 20Γ20s, 11 solve zero, Opus 5.5 still #1 (8/10), new GPT-6.1 Sol 7/10, and Astra's perfect 15Γ15 doesn't transfer (5/10 H
- 10-01 00:46groundThe Hard-mode follow-up sharpens this from an open-gap data point into an instrument for Scott's benchmark-unit argument: the case's own raw-vs-harnessed falsification split operationalizes
- 10-01 00:25attachIndependent update from the same benchmark author's Nonobench Hard-mode runs (Opus 5.5 8/10, Sol 7/10, Astra 5/10) plus Sonnet 5.5 long-thinking cost parity β direct evidence for the open-weight
- 10-01 00:23propose_attachIndependent update from the same benchmark author's Nonobench Hard-mode runs (Opus 5.5 8/10, Sol 7/10, Astra 5/10) plus Sonnet 5.5 long-thinking cost parity β direct evidence for the open-weight
- 09-28 11:42repriceEngagement-only update: two comment churns and a transient 4.5x velocity spike that has already decayed (now 0.67 pts/h, cooling). No new model runs, no independent corroboration, and the 20x20 open-w
- 09-28 07:20sensor_dirtycomment_update
- 09-28 01:21sensor_dirtyvelocity_spike
- 09-27 22:20sensor_dirtycomment_update
- 09-27 21:40promote_anchororigin walk conf 0.78
- 09-27 21:33groundIndependently measured (though single-source β grounding found no coverage of Nonobench or mauricekleine, so the 0/10 is a claim to watch, not a fact) confirmation of the Model Barbell's open-fro
- 09-27 21:25createSecond release of a reproducible community benchmark with a specific falsifiable open-vs-frontier result and an active maintainer iterating per model additions; no open case tracks Nonobench.