long-horizon-agents
band: hotmomentum: stable
score: 0.618
Episodes (20)
Trajectory notes
- 2026-09-03T07:26:50Z: swarmworld-persistent-agent-cultures closed (faded) — The claimed finding converges with Scott’s institutional-memory and long-running-agent position that capabilities and conventions can survive replaceable workers, while extending it with the possibility that specialization a
- 2026-09-01T05:38:30Z: longhorizon-harness-validation closed (faded) — LongHorizon-Harness independently implements Scott’s stateless-worker/stateful-kernel architecture—fresh contexts, durable recoverable state, checkpoints, and separate auditing—and its reported benchmark gains bear directly on his
- 2026-08-29T04:28:17Z: recirculation-running-context closed (faded) — Scott’s context-engineering and long-running-agents frameworks already treat active context as scarce and favor bounded context, compaction, and durable external state; the radar also already tracks comparative context-management c
- 2026-08-27T12:26:59Z: mirrorcode-long-horizon-reimplementation closed (faded) — The radar already tracks this same development in `radar:mirrorcode-autonomous-project-scope`. Its eventual reproduction or failure would directly test Scott’s claims that substantial software can be regenerated from spe
- 2026-08-25T22:29:22Z: autodesign-meta-harness-optimization closed (faded) — AutoDesign independently operationalizes Scott’s position that agent capability resides in the model-plus-harness system and that evaluated runs can breed improved scaffolding, closely converging with Replay-Driven Design Ev
- 2026-08-22T09:24:38Z: mw2-agent-decompilation-marathon closed (faded) — If independently verified, the month-long, 200-billion-token reconstruction would be a substantial external test of Scott’s long-running-agent, verification-loop, legacy-reconstruction, and token-discipline claims—not merely ano
- 2026-08-22T03:26:45Z: oh-my-subagents-multiday-refactoring closed (faded) — Scott already holds the relevant position in Long-Running Agents, Evaluation-Driven Development, and Discussed Is Not Deployed: multi-day agent claims require durable state, explicit completion evidence, and repeatable indep
- 2026-08-15T17:29:39Z: evoharnessrl-self-evolving-agent-harness closed (faded) — The radar already tracks this development’s central validation question on `radar:trained-harness-cross-model-transfer`: whether a learned harness produces transferable gains without overfitting. Independent results woul
- 2026-08-12T06:31:24Z: dspy-factorio-rlm-gepa-agents closed (faded) — The implementation independently combines execution-feedback-driven harness evolution with a long-horizon environment, converging with Scott’s Reflexive Agent Design, Replay-Driven Design Evolution, and model-plus-harness evaluatio
- 2026-08-09T19:41:22Z: open-ended-agent-coordination-benchmark closed (faded) — The reported ablation provisionally challenges Scott’s load-bearing claim that durable external state, rather than better inter-agent coordination, is what makes multi-agent loops reliable. If independent follow-ups confi