GPT-6 Astra is OpenAI's frontier model released September 3, 2026, trained on 100,000+ GPUs with prior OpenAI models supervising training, and shipped with an updated Codex harness (1.9x faster computer-use task completion, persistent notes across context windows). The web results strongly document a viral Minecraft episode: a Vals AI 141-hour livestreamed run in which Astra outpaced all prior AI systems, then lost everything to a Creeper destroying its chest and bed and spent hours 'farming potatoes' with visible behavioral overcorrection. However, the snippets say little about the specific claim in this case — a 'MineTrials' benchmark by creator 'mxls' reporting fixed-seed one-hour evaluations across 13 model–harness combinations where Astra's worst run beat every rival's best run. That MineTrials report is not directly corroborated by the supplied results; only the tangentially related Vals AI long-horizon run and Astra's strong long-horizon/computer-use numbers (e.g., OSWorld 2.0 at 72.6%) appear.
MineTrials independently arrives at Scott's Model-Plus-Harness Benchmark Unit — evaluating 13 model–harness combinations rather than raw weights — and its headline finding (Astra+Codex's worst one-hour run beating every rival's best) is a direct empirical test of his Breaking the 1hr Barrier thesis that the one-hour limit is architectural, not temporal. The attribution question it raises — was it the model or Codex's new persistent notes across context windows — cuts straight at his harness-vs-model attribution and external-state claims, and touches his own remote-exec/trace-backed comparison work; the dated-receipts publishing angle is that a consequential outside benchmark adopted his methodology and produced a result he'd want to interpret.
ip:concept.model-plus-harness-benchmark-unitip:source.breaking-the-1hr-barrierip:framework.long-running-agentsdev:project.remote-exec
queries asked of Scott's wikis
- harness-vs-model attribution: how much agent benchmark performance comes from the harness (compaction, adapter, reasoning-state preservation) rather than the model
- long-horizon agent evaluation methodology: fixed-seed, time-boxed, open-world tasks like Minecraft vs terminal/web benchmarks
- agent memory and compaction: does persistent note-taking across context windows change sustained-task performance
- self-hosted vs frontier model gaps on long-horizon interactive tasks
- evaluation infrastructure or benchmark projects Scott has built or written about for coding agents
- cost-of-evaluation: frontier agent runs costing $18k–$26k+ and what that means for independent benchmarking
now 0 pts/hpeak 185 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 458h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
| source | object | author | score | comments |
| 🟧 hn | MineTrials: How far can AI agents get in an hour of Minecraft?Retrieved article excerptOpen article · Retrieved 2026-09-23T14:28:01.139171+00:00 September 2026
# MineTrials
How far can today's agents get in an hour of Minecraft?
Best
Mean
[Best-run cumulative Minecraft advancements for 13 model–harness combinations over one hour. GPT-6 Astra with Codex reaches 24; icons mark all 24 advancements. Other setups peak at 18.](https://massiminoe.github.io/minetrials/assets/progress.png)
Figure 1 Best run for each model–harness combination in a fixed-seed survival world. Ties use the earliest final advancement. Icons follow Astra’s selected run. [View full size](https://massiminoe.github.io/minetrials/assets/progress.png)
## Introduction
MineTrials explores how effectively today’s models can play Minecraft. Specifically, how many achievements can they collect within a one-hour time limit?
The answer I arrive at is… quite a few! And Astra is really good at Minecraft.
All tested models were capable of playing with some reasonable capacity. Many were able to get diamond gear, and a few even reached the Nether. It almost astounds me that they can play at all, being such square pegs for this task.
I’m also fascinated by the importance of time in this test: both the world still running in real time and the constraint of a time limit. There’s something a bit more tangible about it to me. I like to think that [METR’s work on measuring AI ability to complete long tasks](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/) was so captivating for similar reasons.
Another aspect being evaluated here is the harness. MineTrials is really a meta-harness: we plug into Codex, Claude Code, Cursor, or OpenCode, connecting them to a custom MCP server. I took this direction after initial efforts to build my own were so easily outperformed by out-of-the-box Claude Code. I wonder if there is a lesson here about building harnesses…
## Astra’s best run
You can watch Astra’s best attempt below, including its occasional commentary in chat. For more details, including full traces, see the [dataset on Hugging Face](https://huggingface.co/datasets/mxls/MineTrials).
Astra’s selected best run, with commentary captions.
On why Astra seemed to perform so well, some thoughts: it was certainly more reliable and consistent than the other models. **Even its worst run was better than the best of every other setup** (20 versus 18).
On vibes, I observed it to be more adaptive and flexible. All models made a lot of errors and encountered unexpected situations throughout their runs. Astra seemed the most robust. For instance, in this run, after dying in pursuit of blaze rods, it readily abandoned that objective and took up the more peaceful pastime of fishing.
I also hypothesise that its compaction was more effective than what I believe occurred in Claude Code, which seemed to keep accumulating context within its spacious 1M-token limit.
## Pareto frontier
We’re able to get some data on cost effectiveness here. At face value, the results seem consistent with OpenAI’s claims about occupying the a good share of the frontier.
Note that I couldn’t get token data for Cursor in these experiments. I was limited in resources and of course, making use of subscriptions for these runs. Thus, we're working off the traces I could get.
[Cost–performance comparison of 11 model–harness combinations over 53 runs. The frontier steps through Codex models, from Luna at $0.86 and 14 mean advancements to Astra at $24.53 and 22.4.](https://massiminoe.github.io/minetrials/assets/cost-performance.png)
Figure 2 Cost on a logarithmic axis; performance is the mean one-hour advancement count. Cursor is excluded. [View full size](https://massiminoe.github.io/minetrials/assets/cost-performance.png)
## Conclusion
For anyone grappling with the anxieties of rapid progress in AI, I will say that spending some time playing alongside these agents can offer some temporary relief. It’s both magical and frustrating as they struggle to build a home and interact with a live (and hostile!) 3D world.
I first saw this idea two years ago on [Emergent Garden’s YouTube channel](https://www.youtube.com/watch?v=NTHWMk5pcYs). The [Mindcraft project](https://github.com/kolbytn/mindcraft) has since inspired many others to explore what language models can do in Minecraft. I also draw parallels to the original [Claude Plays Pokémon](https://www.twitch.tv/claudeplayspokemon) and even Typeface's more timely [Doomo.](https://typesafe.ai/blog/introducing-system-one-models-and-jev#:~:text=Doom) | mxls | 2 | 0 |
| 🟧 echo.blog ⭐ | Reports fixed-seed, one-hour Minecraft evaluations across 13 model–harness combinations, with Astra achieving 20–24 advancements versus a ma | mxls | — | — |
| 🟠 reddit | PokeBench: I gave frontier LLMs the Game Boy screen and 1,000 turns to beat Pokemon Red's first gym OpenAI | VibeCodyH | 87 | 28 |
| 🟠 reddit | Astra leads in IKEA furniture assembly singularity | Proper_Actuary2907 | 125 | 16 |
| 🟧 hn | Show HN: StarSkirmish, an arena where LLMs create StarCraft Brood War bots | __cayenne__ | 4 | 1 |
| 🟠 reddit | ChatGPT-6 Astra plays World of Warcraft 'blind' and clears the orc starting zone in 40 minutes with no deaths — AI agent navigates by parsing raw server network packets and SQL files OpenAI | ThereWas | 1349 | 276 |
| 🟠 reddit | ChatGPT-6 Astra plays World of Warcraft 'blind' and clears the orc starting zone in 40 minutes with no deaths — AI agent navigates by parsing raw server network packets and SQL filesa artificial | ThereWas | 40 | 6 |
2026-10-07T14:49:47Z
The Oct 5-8 velocity spikes are all the same object — the WoW demo post's cumulative engagement tail (1349 pts / 276 comments) — not new evidence: current rate is ~0 pts/h at the 12.5th percentile and no new evaluator line has appeared since Sept 26. The magnitude-valve spread reading is one tangential story (Reddit megapost + crosspost + one outlet) already assessed and discounted on Oct 4-5, not eval-genre breadth, so heat drops to low. The case settles as a stable corroborated finding — hour-long interactive arenas show sharp frontier stratification, Astra leading three of four lines — coasting on its named material markers.
2026-10-04T17:02:53Z
The Sept-27 cooldown condition has landed: ~8 days with no new evaluator line, and the fresh burst is a tangential WoW demo — packet/SQL state access, i.e. an LLM-driven bot rather than blind play — that adds Astra-game-agent attention but no eval-genre evidence (and is really one story posted twice plus one outlet). The case steps back from accelerating to corroborated; heat drops to medium on periphery breadth, not on the demo's top-decile single-post velocity.
2026-10-04T16:42:23Z
evidence attached: reddit.post.1wxirdb — Tom's Hardware-covered Astra game run succeeding via raw packet/SQL parsing independently contextualizes the claimed model-and-harness advantage in sustained interactive game tasks.
2026-10-04T16:42:23Z
evidence attached: reddit.post.1wxin4m — Independently covered (Tom's Hardware) Astra game-agent demo adds to the accelerating interactive-task evidence, though packet/SQL state access rather than perception tempers what it shows.
2026-09-26T17:29:45Z
StarSkirmish adds a fourth independent line — another builder independently adopting the one-hour wall-clock time-box, this time for StarCraft bot coding — shifting the case's meaning from 'Astra dominates game evals' to 'hour-long interactive arenas are becoming a grassroots eval genre that consistently shows sharp frontier stratification' (Astra leads three lines, but only ties Opus 5.5 in the fourth). Measured heat flipped cooling→accelerating (10.7 pts/h, 84th percentile) on this expansion; heat holds high on periphery breadth, not single-story velocity, and drops quickly if new evaluator lines stop appearing.
2026-09-26T17:27:35Z
evidence attached: hn.story.49858284 — Independent second hour-long interactive arena (StarCraft bot coding) echoes the frontier-stratification finding — GPT-6 Astra and Opus 5.5 tied, steep drop-off outside the frontier — material corroboration for re-judging that case.
2026-09-26T07:25:16Z
relevance=high case never alerted; deterministic escalation to deliver
2026-09-26T07:23:35Z
evidence attached: reddit.post.1wqj7mu — Independent corroboration from a credible evaluator (Epoch AI): Astra again leads competitors on a physical/interactive assembly benchmark, extending the same model-and-harness advantage pattern.
2026-09-25T20:31:33Z
The Astra-dominance inference graduated from lone report to cross-game pattern: PokeBench (independent creator, Pokemon Red) ranks Astra best of 14 with a local model failing outright, joining Vals AI's 141-hour Minecraft run as second independent line — but engagement is dead (0 comments everywhere), so the case is corroborated-and-cold, kept at medium only because independent game-benchmark replications keep appearing within days and the model-vs-Codex-harness attribution Scott cares about is still open.
2026-09-25T20:24:27Z
evidence attached: reddit.post.1wq6cy4 — Independent game-benchmark result (Astra best of 14, local model fails entirely) corroborates the case's broader inference about model-and-harness advantage in sustained interactive tasks.
2026-09-23T20:06:07Z
grounded: converges/high — MineTrials independently arrives at Scott's Model-Plus-Harness Benchmark Unit — evaluating 13 model–harness combinations rather than raw weights — and its headl
2026-09-22T17:30:33Z
case created — Published traces and bounded comparative results make this a concrete interactive-agent evaluation rather than a generic benchmark proposal.