2026-10-11 16:37 UTC

Dunnolab claims NetHackers' released registry, held-out evaluation and shared elite bots provide a reproducible substrate for humans and coding agents to cumulatively improve modern NetHack bots toward the first verified 3.6.6 ascension.

state: watchingheat: lowuncertainty: highknownscott: lowagent-harnesses agent-evaluation long-horizon-agentsDunnolab

What is this?

NetHackers is presented as an open effort and leaderboard for improving inspectable, symbolic NetHack bots under a common held-out evaluation, with Dunnolab identified by the case as the party behind it. The case further claims a released registry and shared elite bots intended to support cumulative human-and-coding-agent progress toward a verified NetHack 3.6.6 ascension. The supplied web results do not directly cover NetHackers, so the platform details, release status, participation, and claimed ascension target remain uncorroborated here.

Why it matters to Scott

Scott’s Long-Running Agents, Hidden Gates, Replay-Driven Design Evolution, and Trace-backed Agent Comparison pages already carry the core position: durable agent work should run in reproducible harnesses with retained evidence and held-out checks. NetHackers is presently only an uncorroborated game-domain instance of that position; without demonstrated participation, useful traces, or frontier progress, it does not yet change what Scott would build or argue.
ip:framework.long-running-agentsip:framework.hidden-gates-frameworkip:framework.replay-driven-design-evolutiondev:concept.trace-backed-agent-comparisonradar:longhorizon-harness-validationradar:harnessopt-agent-harness-optimization-benchmarkradar:spire-agent-long-horizon-gameplayradar:dspy-factorio-rlm-gepa-agents
queries asked of Scott's wikis
  • reproducible harnesses for long-horizon agents
  • held-out evaluation and benchmark integrity
  • inspectable symbolic agents versus LLM agents
  • human and coding-agent cumulative improvement loops
  • shared agent traces artifacts and elite baselines
  • game environments for long-horizon agent evaluation

Measured heat

now 0 pts/hpeak 6 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 434h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-23 14:00⭐ origin echo-reconstructedNetHackers launches an open effort and leaderboard for improving inspectable symbolic bots under common held-out evaluation, with registered
Dunnolab on blog (echo) · attributed from hn.story.49823869
—
09-23 23:04first on hacker news · published · +9.1hShow HN: NetHackers
vokneruk
—
09-23 23:04amplified on hacker news 👑hn.story.49823869
vokneruk
peak 1 · 2 comments · 38% of case engagement
09-26 05:18amplified on hacker newshn.story.49853442
EvgeniyZh
peak 1 · 0 comments · 13% of case engagement
09-29 17:47amplified on hacker newshn.story.49897392
kenforthewin
peak 2 · 0 comments · 25% of case engagement
10-04 06:10amplified on hacker newshn.story.49951137
bigmadshoe
peak 2 · 0 comments · 25% of case engagement
09-24 00:21our radar first saw it · +10.4hdiscovery anchor: hn.story.49823869—
pace: p42 vs 1032 stories at the 336h mark (now 434h old) — ahead of agent-memory-add-search-evaluation (1.2x), behind agenticos-self-hosted-governance (0.9x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: NetHackers
Retrieved article excerpt

Open article · Retrieved 2026-09-24T00:31:41.120221+00:00

## @ What it is? solving = AGI?



**NetHackers is an open effort to build the first program that can
reliably win NetHack 3.6.6.** Write the bot by hand, evolve it with coding agents, improve the thing
that improves it, or recurse until something interesting happens. Whatever path you take, the evaluation
is the same: the symbolic bot that comes out of the loop. The ultimate goal is an
**ascension**; until then, the board measures how far each bot gets.

The challenge

**NetHack** (1987) is among the oldest **unsolved** challenges in games.
Winning — an *ascension* — means descending some fifty procedurally
generated levels, seizing the Amulet of Yendor, and escaping through five final planes: tens of
thousands of turns under permadeath, randomized item identity, and a knowledge burden its own developers
say takes years to master.

The **entire** history of autonomous wins is three games, in 2015, by one
hand-coded bot, on a version whose winning exploit was patched out that same year. On modern
NetHack, **no program has ever won.**

THE STATE OF THE DUNGEON — every line of attack, ascensions on NLE's NetHack 3.6.6

| Line of attack | Best known result | Asc. |
| --- | --- | --- |
| Humans, all logged games | — | 0.4% |
| Humans, experts | 15.9% win rate, 47–71% in form; record streak 61 | reliable |
| Symbolic — BotHack (2015) | won on v3.4.3 via a pudding-farming exploit, removed in 3.6 | 3\* |
| Symbolic — AutoAscend (2021) | **7.8%** progression — the floor this board measures; Medusa in 0 of 109,545 games | 0 |
| Learned — imitation (2021–) | scaling flattens **below** the very bot it imitates | 0 |
| Learned — RL (2026) | **11.78%** progression score-trained, **16.98%** depth-trained — first to beat AutoAscend since 2021 | 0 |
| LLM direct play (2026) | **13.2%** progression in BALROG | 0 |

Every machine line: **zero ascensions on NetHack 3.6.6**, the NLE research
standard. Progression is one 0–1 metric throughout, but each line is
measured on its own seeds and episode counts, so read it as indicative, not head-to-head.
\*BotHack's three wins were on the older 3.4.3, via an exploit the developers then removed.

One number did just move. In September 2026 an RL agent beat
AutoAscend for the first time since 2021 — roughly doubling its challenge median, with
depth-trained policies reaching the Castle, deeper than AutoAscend has ever gone. On
progression that is 16.98% against AutoAscend's 7.8%.

Why now

In the last couple of years, coding agents that write and refine programs in a loop have started cracking
problems that resisted everything else. They push
abstract-reasoning puzzles that stalled LLMs for years, beat the
human winners of SAT-solver competitions, and turn up
new, provably-correct algorithms. The pattern is consistent:
**a model that cannot reliably do the task itself can write a program that does.**

NetHack is where that pattern has not yet held. The researchers behind the
NetHack Learning Environment put it forward as
a grand challenge for AI, and it is still unsolved years
later; NLE co-author Tim Rocktäschel marks the anniversary each year with "AI still
can't learn to play NetHack." It is the hardest game in the BALROG
suite, and the one frontier models are worst at by a wide margin. Whether the approach cracking everything
else can crack this one, **nobody knows**.

That is what NetHackers is for: an open attempt to find out **together**. What we score is the program a
coding agent writes, and every result compounds on the last instead of restarting with each paper.

How it works

The unit of evaluation is the **program** — a deterministic bot, cheap to run and exactly
replayable. Objectives grid over the **73 starting identities** and the milestone ladder, so
specialists and generalists all have somewhere to land. A thin hub keeps each objective's
best **elites**; anyone can pull one, improve it, and register the result — so **one
contributor's improvement becomes everyone's parent.** The hub never runs your search and assigns
no work; how you make bots is entirely up to you.

---



## ! Why you should care

There are at least three ways to fall into this dungeon.

If you cannot leave a loop alone

If Recursive,
Ricursive,
Discovery Loop,
AIDE², and the
Darwin Gödel Machine all appeared in your timeline before
breakfast; if every benchmark bump is *“it’s happening”* and every plateau means
*“add another outer loop”*; if your honest answer to “what improves the
improver?” is “another improver” — welcome. Build the seed, mutate the harness,
fork the fork. The intelligence explosion can start with not dying to a grid bug.

If you do AI research

Winning NetHack is the headline; the research problem is **generalization**. A bot must turn wiki
knowledge into action, decompose a tens-of-thousands-of-steps objective, discover and compose reusable
skills, and recover when unfamiliar seeds or stochastic events break its plan.
Held-out evaluation tests whether those skills transfer rather
than whether one trajectory was memorized — and the result is a symbolic program you can inspect.

These are not game-only problems. ASPIRE applies a similar loop
to robotics, repairing code-as-policy programs after failed rollouts and saving skills for new tasks;
Code as Policies composes perception, control, and tools into
executable robot behavior. NetHack is a cheap, fast arena for studying long-horizon planning, skill
composition, tool creation, and generalization to unfamiliar seeds — without a robot lab.
[» how held-out evaluation works](https://github.com/dunnolab/nethackers/blob/main/docs/verification.md)

If you just think it's cool

You don't need to be good at NetHack, or an ML researcher. A coding agent and a laptop will do.
Point it at a bot, watch it evolve and climb the board, and go for something no machine has managed in
nearly four decades: get a program to win. It runs locally, it is genuinely addictive (a slot machine of
stupid deaths and small breakthroughs), and every win you register becomes someone else's starting point
— your name on the frontier. [» start solving](https://github.com/dunnolab/nethackers)

---



## < The Frontier

Mean progression (0–100%) across all 73 identities, and how far each sits above the
**AutoAscend floor**. Click any identity for its directly comparable leaderboard, sources, episodes, dates, and current holder.

dungeons:

Public Dungeons (15)
Private Dungeons (15)
?

progression:

0% → 100%
 AutoAscend floor
 Δ above floor · below
 ★ ascended ≥ 87.5%

---



## & Hackers moving the frontier

Recognition is attached to concrete records, roles, and source commits — not collapsed into one global score.

Frontier Keepers

dungeons:

Public Dungeons (15)
Private Dungeons (15)
?

Current identity leaders above AutoAscend. Click a row for the hacker's complete contribution history.

Loading frontier keepers…

Greatest Breakthroughs

dungeons:

Public Dungeons (15)
Private Dungeons (15)
?

Loading breakthroughs…

---



## \* The Oracle

NetHack's Oracle sells prophecy for gold; ours is free — and it's
**you**. Two questions on **whether**, and **how**, a program finally wins. Cast your prophecy to see where the
crowd — and **your own tribe** — lands.

Q1. Which approach writes the first ascending program?

Q2. When does the first ascension on held-out seeds happen?

your tribe:

nethack:

Consult the Oracle
answer Q1 & Q2 to cast · one prophecy per browser

segment:

Everyone
By role
By NetHack XP

[ change my prophecy ]

---



## % FAQ

How is this different from the other NetHack competitions?

Same dungeon, three different jobs.

THREE NETHACK CONTESTS — WHAT EACH ONE ACTUALLY BUILDS

| Competition | Your job | Evaluation |
| --- | --- | --- |
| **NetHackers** | Build an autonomous symbolic bot that plays NLE's NetHack 3.6.6. | An ongoing, open leaderboard for progression and ascension across fixed identities, with every bot available as a starting point for the next improvement. |
| [NeurIPS NetHack Challenge 2021](https://nethackchallenge.com/) | Build an agent to play the full game through NLE 3.6.6. | A fixed 2021 event: ascensions first, then median in-game score over random characters. |
| [Mazes of Menace](https://mazesofmenace.ai/) | Port NetHack 5.0 from C and Lua to readable ES6 JavaScript. | Bit-exact screen and PRNG parity on public and held-out sessions, followed by a generalization phase. |

How does this compare to ARC-AGI?

They ask opposite questions about the same gap.

ARC-AGI limits what the solver is told: infer a transformation
from a few examples, then apply it to a new input. It measures how efficiently a system adapts when task
evidence is deliberately scarce.

NetHackers makes the inverse choice. NetHack's rules are public — you can read the implementation,
consult decades of wiki knowledge, study every bot that came before, and run the game as often as you like.
Knowledge from earlier attempts is something to keep and exploit, not something to withhold. The question is
how much competence a process can build out of all that.

Having the rules still leaves the decisions. A running program sees only what it has explored. A simulator
does not hand you a tractable strategy — you still have to choose what to consider and what to ignore.
And choices interact across tens of thousands of turns, where progress now can consume what you needed
later.

The two are **complementary**. A system can recognize a new pattern from three examples and still fail
to build and maintain a large, reliable controller; a strong NetHack specialist would say little about
adapting to unrelated tasks.

Why AutoAscend?

Because it is a strong starting point, not a finished answer.
AutoAscend won the 2021 challenge and already contains serious
symbolic machinery for exploration, combat, inventory, altars, and Sokoban. It also contains plenty of
strange, sometimes plainly dumb behavior: every role begins by farming dungeon level 1 until experience
level 8; Monks are hard-coded never to choose a melee weapon or body-armor suit; autopickup is replaced
by a brittle hand-written item-priority system; bag use is disabled; shopping is unimplemented; and the
post-Mines plan still turns into TODOs. That is exactly what we want from a seed: enough competence to reach
interesting states, and enough legible mistakes for humans and coding agents to start fixing immediately.

Why not BOINC or an “@home” project?

Because volunteer computing solves the opposite problem. BOINC ships
code the project wrote and signed to machines it does not trust, then validates results by having two of them
agree. Both halves invert here. The untrusted thing is the **code** — a stranger's bot, or a coding
agent running with its permission prompts switched off — and a volunteer's side of that bargain is that
they never have to read it, because the project signed it and stands behind it. We cannot make that promise
about code we neither wrote nor reviewed, and routing it in as an input file to a signed wrapper only means
the signature stops covering the part that actually runs. Agreement would not buy us much either: a bot
overfitted to the public seeds is not a disagreement between hosts, it is a number every replica reproduces
perfectly. The only check that catches it is re-running on dungeons the author has never seen — and a
volunteer machine cannot hold a secret. Cycles were never the scarce thing anyway; a full private pass is
roughly a thousand episodes, hours on one box. What is scarce is good programs, and the agent tokens to find
them. So the hub hands out no work at all: you search on your own machine, and we keep the link.
vokneruk12
🟧 echo.blog ⭐NetHackers launches an open effort and leaderboard for improving inspectable symbolic bots under common held-out evaluation, with registeredDunnolab——
🟧 hnAn LLM Beat NetHackEvgeniyZh10
🟧 hnAn LLM Beat NetHackkenforthewin20
🟧 hnAn LLM Beat NetHackbigmadshoe20

Interpretation history

Decision trace