LongHorizon-Harness is an open-source computer-use agent harness from AMAP-ML for running extended workflows across desktop applications and the CLI. It uses a Manage-Execute-Audit loop with durable verified task state, fresh-context execution, read-only auditing, recoverable progress, and integrations with Claude Code, Codex, and OpenClaw; its authors report improved results across benchmarks, including OSWorld 2.0. The supplied snippets establish the project’s release and reported benchmark gains, but do not establish genuinely independent reproduction or validation of its practical utility.
2026-09-01T05:38:30Z
Repeated staleness checks have produced only adjacent examples, with no direct reproduction, adoption, or comparative evaluation of LongHorizon-Harness. The project-specific episode has faded and can be reopened if an independent test appears.
2026-08-30T05:30:08Z
No direct evaluation, reproduction, adoption, or comparative result has emerged; the surrounding harness pattern is increasingly well illustrated, but that repetition no longer adds meaning to the project-specific case. LongHorizon-Harness remains a cold testing opportunity rather than a validated framework.
2026-08-28T05:26:37Z
The new field report reinforces context loss and declining focus as practical long-horizon failure modes, but it is another anecdotal adjacent signal rather than a test of LongHorizon-Harness. Repeated architectural corroboration has not become project-specific reproduction, adoption, or comparative evidence.
2026-08-28T05:23:10Z
evidence attached: reddit.post.1w0h6qm — This field report describes persistent context loss and declining task focus in a long-running coding project despite handoff summaries.
2026-08-26T23:26:06Z
The durable-versus-long-lived framing strengthens the surrounding architectural rationale for recoverable state and restartable workers, but adds no direct evaluation, reproduction, or adoption of LongHorizon-Harness. The project-specific utility and benchmark claims remain unvalidated.
2026-08-26T23:23:17Z
evidence attached: hn.story.49456841 — The durable-versus-long-lived distinction materially contextualizes how long-horizon agent systems should persist, recover, and manage state.
2026-08-26T21:27:14Z
The blocker-alert artifact adds a practical human-escalation mechanism for unattended agents, refining the surrounding operational pattern without testing LongHorizon-Harness. Its reproducibility, benchmark gains, and real-world utility still await direct independent evaluation or adoption.
2026-08-26T21:24:05Z
evidence attached: hn.story.49455493 — A usable blocker-alert artifact materially contextualizes how long-running coding agents can remain supervised while unattended.
2026-08-26T06:31:58Z
The real-world agent-fleet retrospective strengthens the evaluation criteria around checkable task boundaries, human review, and failures hidden by automated metrics. It remains adjacent practice evidence rather than an independent reproduction, adoption, or direct evaluation of LongHorizon-Harness, leaving the core claim unsettled.
2026-08-26T06:23:13Z
evidence attached: reddit.post.1vyo6dk — Detailed real-world case study provides useful evidence about human-supervised fleets of coding agents, checkable task boundaries, and failures missed by automated metrics.
2026-08-25T20:42:21Z
The benchmark proposal sharpens the required validation design by separating model capability from harness, tool, context, and acceptance-gate effects. Because it provides no implementation or results and does not test LongHorizon-Harness, the project-specific reproducibility and utility claims remain unvalidated.
2026-08-25T14:25:21Z
evidence attached: reddit.post.1vy0ki7 — The proposed benchmark usefully separates model capability from orchestration, context assembly, tool design, and acceptance-gate effects.
2026-08-25T09:38:15Z
Terminal-Bench expands the adjacent evaluation landscape but supplies no new results, methodology, comparison, or direct test of LongHorizon-Harness. The architectural neighbourhood remains active while the harness-specific reproducibility and practical-utility claims remain unvalidated.
2026-08-25T09:23:15Z
evidence attached: hn.story.49430706 — Terminal-bench is relevant independent evidence for whether terminal-agent evaluation can measure practical autonomous coding performance.
2026-08-25T01:24:33Z
Headlong adds another independent implementation of the persistent-agent harness pattern, but the available evidence offers no technical results, comparison, adoption, or direct evaluation of LongHorizon-Harness. The architectural neighbourhood is active while the project-specific reproducibility and utility claim remains unvalidated.
2026-08-25T01:22:57Z
evidence attached: hn.story.49427736 — A first-party GitHub release of a microharness centered on persistent agency is directly relevant evidence for long-running agent-harness evaluation.
2026-08-24T20:46:47Z
The new StateM item is duplicate adjacent coverage, not an independent reproduction, comparison, or adoption of LongHorizon-Harness. It reinforces the surrounding stateful-harness pattern but leaves the project’s reproducibility and practical utility unvalidated.
2026-08-24T19:26:37Z
evidence attached: hn.story.49423887 — shared external link with case evidence
2026-08-24T15:24:37Z
The staleness check adds no direct evaluation, reproduction, or adoption; surrounding implementations continue to support the architectural pattern without validating LongHorizon-Harness itself. Keep the case open but cold pending an external test.
2026-08-22T14:42:28Z
StateM adds another adjacent implementation signal around stateful long-horizon control, but the available evidence contains no architecture details, results, adoption, or comparison that validates LongHorizon-Harness. The case remains cold and dependent on a direct external reproduction or practical evaluation.
2026-08-22T14:23:03Z
evidence attached: hn.story.49399887 — Stateful control for long-horizon agents directly bears on whether agent harnesses can preserve and manage state over extended tasks.
2026-08-22T08:28:53Z
No new independent evaluation, reproduction, or adoption has appeared; recent evidence remains adjacent support for the broader architecture rather than validation of LongHorizon-Harness. The case stays open but cold while awaiting a direct external test.
2026-08-20T07:36:59Z
The production outcome-gated workflow adds a concrete design criterion for mechanically verified completion, strengthening the surrounding harness pattern. It neither uses nor independently evaluates LongHorizon-Harness, so the project’s reproducibility and practical gains remain unvalidated.
2026-08-20T07:22:46Z
evidence attached: reddit.post.1vtbjwk — A concrete production example shows an agent pipeline using successful external events as its definition of done, offering relevant evidence for long-running harness design.
2026-08-19T21:43:15Z
The refreshed discussion sharpens a known continuation failure into a testable harness requirement: persist status within tool-driven durable state rather than emitting prose that creates a turn boundary. It still provides no independent evaluation, reproduction, or adoption of LongHorizon-Harness, so the core hypothesis remains unsettled.
2026-08-19T19:23:22Z
evidence attached: reddit.post.1vsuq8o — User evidence contradicts the assumption that prompts and subagent loops alone reliably sustain continuous long-running coding-agent execution.
2026-08-19T16:53:07Z
The eight-day backlog-and-CI loop is an independent implementation-level anecdote supporting the broader durable-state, fresh-context architecture, enough to move the pattern into watching. It does not use or evaluate LongHorizon-Harness and lacks artifacts or quality analysis, so the harness’s reproducibility and practical gains remain unvalidated.
2026-08-19T16:24:23Z
evidence attached: reddit.post.1vsr2qm — Reports an eight-day recurring backlog-and-CI loop that bears directly on long-running coding-agent harness reliability.
2026-08-19T01:26:35Z
The 24-hour coding-agent anecdote reinforces the broader need for durable execution and mechanically independent evaluation, but its failed judge and lack of connection to LongHorizon-Harness provide no external validation of the harness itself.
2026-08-18T23:23:12Z
evidence attached: reddit.post.1vs4ssl — An anecdotal 24-hour coding-agent run provides weak but relevant evidence about long-horizon execution, evaluation failure, and autonomous iteration.
2026-08-18T12:33:03Z
The compiler rewrite is an adjacent example of long-horizon agent work, not an independent evaluation, reproduction, or adoption of LongHorizon-Harness. The harness therefore remains a relevant first-party artifact whose practical utility and reported gains are unvalidated.
2026-08-18T12:22:54Z
evidence attached: hn.story.49344173 — A concrete production compiler rewrite by AI agents provides useful evidence about long-horizon coding-agent performance.
2026-08-18T00:28:25Z
Re-evaluation adds no independent testing, implementation uptake, or benchmark reproduction; the case remains a promising first-party artifact awaiting external validation.
2026-08-18T00:26:26Z
grounded: converges/medium — LongHorizon-Harness independently implements Scott’s stateless-worker/stateful-kernel architecture—fresh contexts, durable recoverable state, checkpoints, and s
2026-08-18T00:23:23Z
case created — The first-party repository is a concrete harness artifact, but it has only one low-engagement observation and no independent validation yet.