2026-10-11 16:38 UTC

Benchmark author mauricekleine's Nonobench v1.2 reports that among 43 tested LLMs no open-weight model solves the new 20Γ—20 Hard mode (0/10) while GPT-6 Astra posts the first perfect 15Γ—15 run (30/30) and Opus 5.5 scores 8/10; if the open-weight shutout holds as open-weight releases are added, it stands as a measured open-frontier gap on long-constrained verifiable reasoning, and any open model clearing 20Γ—20 refutes it.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumllm-benchmarks open-weight-modelsmauricekleine

What is this?

Nonobench is a reproducible community benchmark run by independent developer Maurice Kleine that scores LLMs on nonogram puzzles β€” one attempt per puzzle, no tools, answers machine-checked against the clues β€” with a public leaderboard at v1.2 covering 43 models. Its headline claims: no open-weight model has solved any 20Γ—20 Hard-mode puzzle (0/10; best open entry DeepSeek V4 Pro sits at 83.3% overall), while OpenAI's GPT-6 Astra posted the first perfect 15Γ—15 run and Anthropic's Opus 5.5 leads Hard at 8/10 β€” a shutout that held in a 2026-09-30 follow-up in which 11 of 15 models (including Astra, at 5/10) solved zero 20Γ—20s. The supplied web results corroborate the landscape but not the benchmark: the model names and release dates (GPT-6 Astra 2026-09-03, GPT-6.1 Sol 2026-09-29, Opus/Sonnet 5.5 in late September) match release trackers, and independent 2026 reporting (Lumiere, NerdLevelTech, Artificial Analysis figures) measures a persistent open-vs-closed frontier gap that is widest on long-horizon and written-from-scratch benchmarks, with harness effects flagged as a known confound. But zero coverage of Nonobench or Kleine surfaces anywhere, so the specific 0/10 shutout remains a single-source claim β€” falsified if any open model clears 20Γ—20 raw, while a harnessed-only clear would instead refute the 'open models can't' reading.

Why it matters to Scott

The Hard-mode follow-up sharpens this from an open-gap data point into an instrument for Scott's benchmark-unit argument: the case's own raw-vs-harnessed falsification split operationalizes his model-plus-harness unit claim (with Nonobench's tool-free, one-attempt design as the wrong-unit counterfactual), and Opus taking 8/10 at 'high' while Sol (max) and Astra (xhigh) trail at 7 and 5 is measured evidence that reasoning-effort burn doesn't convert on constraint puzzles. Stays medium β€” still single-source and quiet, consequential only on a trigger (any open 20Γ—20 clear, and especially a harnessed-only one, which would be dated receipts for the harness-unit position over the 'open models can't' reading).
ip:concept.model-plus-harness-benchmark-unitip:concept.benchmarking-the-wrong-unitip:concept.inference-time-scalingip:concept.capability-symmetryradar:concept.open-modelsradar:concept.open-weight-modelsradar:concept.benchmark-integrityradar:concept.llm-evaluationradar:pcss-zebra-puzzle-reasoning-transferradar:minizinc-mcp-solver-toolsradar:ship-harness-benchradar:claude-sonnet-55-token-economicsradar:qwen38-27b-reasoning-effort
queries asked of Scott's wikis
  • open-weight vs frontier capability gap position
  • harness or scaffold as part of the model eval
  • reproducible verifiable LLM benchmark design
  • benchmark contamination fresh held-out evals
  • LLM constraint satisfaction puzzle reasoning limits
  • reasoning token cost per solved task economics

Measured heat

now 0 pts/hpeak 11 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 386h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-25 14:00⭐ origin echo-reconstructedNonobench is Kleine's own LLM nonogram benchmark leaderboard; the v1.2 results it publishes are the source of every claim in the Reddit titl
Maurice Kleine on other (echo) Β· attributed from reddit.post.1wrh7b2
β€”
09-27 10:51first on r/LocalLLaMA Β· published Β· +44.9hNonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20Γ—20 Hard mode
mauricekleine
β€”
09-30 14:07first on r/ClaudeAI Β· published Β· +120.1hClaude Opus 5.5 leads my nonogram benchmark's Hard mode: 8 of 10 random 20x20 puzzles, while 11 of 15 models solve not even one
mauricekleine
β€”
10-04 07:57first on r/MachineLearning Β· published Β· +210.0hNonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
mauricekleine
β€”
09-27 10:51amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wrh7b2
mauricekleine
peak 28 Β· 26 comments Β· 66% of case engagement
09-30 14:07amplified on r/ClaudeAIreddit.post.1wu5w1c
mauricekleine
peak 9 Β· 3 comments Β· 15% of case engagement
10-04 07:57amplified on r/MachineLearningreddit.post.1wxa2bs
mauricekleine
peak 8 Β· 8 comments Β· 20% of case engagement
09-27 11:20our radar first saw it Β· +45.3hdiscovery anchor: reddit.post.1wrh7b2β€”
pace: p65 vs 1032 stories at the 336h mark (now 386h old) β€” ahead of gfx906-expert-pool-admission-fix (1.0x), behind ordewell-editable-multi-runner-plans (1.0x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20Γ—20 Hard mode
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-27T11:24:15.544855+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
mauricekleine2826
🟧 echo.other ⭐Nonobench is Kleine's own LLM nonogram benchmark leaderboard; the v1.2 results it publishes are the source of every claim in the Reddit titlMaurice Kleineβ€”β€”
🟠 redditClaude Opus 5.5 leads my nonogram benchmark's Hard mode: 8 of 10 random 20x20 puzzles, while 11 of 15 models solve not even one
ClaudeAI
mauricekleine93
🟠 redditNonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
MachineLearning
mauricekleine88

Interpretation history

Decision trace