2026-10-11 17:11 UTC

Independent evaluations will determine whether LongHorizon-Harness provides a reproducible and practically useful framework for assessing and improving agents on extended real-world tasks.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses long-horizon-agents agent-evaluationAMAP-ML

What is this?

LongHorizon-Harness is an open-source computer-use agent harness from AMAP-ML for running extended workflows across desktop applications and the CLI. It uses a Manage-Execute-Audit loop with durable verified task state, fresh-context execution, read-only auditing, recoverable progress, and integrations with Claude Code, Codex, and OpenClaw; its authors report improved results across benchmarks, including OSWorld 2.0. The supplied snippets establish the project’s release and reported benchmark gains, but do not establish genuinely independent reproduction or validation of its practical utility.

Why it matters to Scott

LongHorizon-Harness independently implements Scott’s stateless-worker/stateful-kernel architecture—fresh contexts, durable recoverable state, checkpoints, and separate auditing—and its reported benchmark gains bear directly on his claim that capability belongs to the model-plus-harness unit. It is a strong dated-receipts and testing opportunity, but practical utility and verifier independence remain unvalidated, preventing high relevance.
ip:framework.long-running-agentsip:concept.model-plus-harness-benchmark-unitip:concept.mechanically-different-verifiersdev:concept.resumable-agent-job-control-planeradar:concept.agent-harnessesradar:concept.long-horizon-agentsradar:concept.verificationradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • durable state and fresh-context execution for long-running agents
  • manager executor auditor architecture for agent harnesses
  • independent verification and audit trails in computer-use agents
  • recoverable progress and context management for long-horizon tasks
  • real-world agent evaluation versus benchmark gaming
  • harness-layer improvements versus model capability gains

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (16) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnLongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Taskstingletech11
🟧 echo.github ⭐Released LongHorizon-Harness for advancing and evaluating long-horizon agents on real-world tasks.AMAP-ML——
🟧 hnRewriting a production compiler's IR with AI agents in five weeksCommanderTvis20
🟠 redditClaude Opus5 failing absurd AAA benchmark, producing a functional game anyway (starting with Three.js, empty repo).
ClaudeAI
incoherentian12
🟠 redditI extended Matt Shumer's gauntlet loop so it works for apps + CI + more on CC
ClaudeAI
viborci10
🟠 redditHow do you get claude to work on an issue continuously instead of stopping after a report
ClaudeAI
IamLeperMessiah37
🟠 redditThe cleanest "definition of done" I've seen: the ticket can only be closed by a successful run
ClaudeAI
Frequent-Ad-83621
🟧 hnStateM: Stateful control for long-horizon agentsjohntrob1411
🟧 hnStateM: Stateful Control for Long-Horizon Agentsjohntrob14102
🟧 hnHeadlong: Microharness Featuring Persistent Agencyhandfuloflight31
🟧 hnTerminal-bench: benchmarks for AI agents in terminal environmentsBluestein20
🟠 redditWhat would a fair benchmark for agent architecture look like? [D]
MachineLearning
jonah_omninode10
🟠 redditI built a 600-card mobile roguelike deckbuilder with Claude agents doing most of the implementation. The game is about working with AI agents; building it meant living that.
ClaudeAI
bariyu010
🟧 hnGet woken up in the middle of the night when your agent hits an external blockerwhp_wessel23
🟧 hnAgents Should Be Durable, Not Long-Livedsoasme10
🟠 redditclaudeplaint of the day: they have been tiring out
ClaudeAI
Deathnote_Blockchain11

Interpretation history

Decision trace