2026-10-11 17:10 UTC

long-running-agents

band: warmmomentum: stable score: 0.585
temperature history

Episodes (21)

Independent use will determine whether OpenMetaLoop can reliably sustain autonomous long-horizon tasks across multiple sessions.
expiredknownscott: low
Independent deployments will determine whether Supervice reliably supervises, restarts, and controls the lifecycle of long-running agent processes without external dependencies.
expiredknownscott: low
Independent deployments will determine whether Dropstone SDK provides reliable persistence, recovery, and continuity for long-running agents beyond disposable session-based runtimes.
expiredknownscott: low
Independent deployments will determine whether Machine0’s CLI- and MCP-managed persistent CPU and GPU VMs provide reliable, economical infrastructure for multi-day agent workloads.
expiredknownscott: low
Independent use will determine whether the released constitutional governance practices improve the reliability, controllability, and maintainability of long-running personal agent fleets.
expiredconvergesscott: medium
Artifact review and further operation will determine whether Cairn Wake’s file-persisted, scheduled Claude agent can run a small online business reliably over multi-week periods while keeping spending human-gated.
expiredconvergesscott: medium
Independent use and repository review will determine whether Ducklab can reliably automate iterative software construction with local models at costs comparable to its reported 416-run, $176 self-development process.
expiredconvergesscott: high
Gantree’s creator claims its chat-independent harness makes long-horizon agent work persistent, resumable, and manageable outside a conversational session.
expiredknownscott: low
The paper’s authors claim persistent non-decaying state can make safety failures compound across autonomous LLM-agent loops, implying long-running harnesses need lifecycle-level rather than per-step controls.
expiredconvergesscott: medium
PILOT’s authors claim their within-run self-improvement mechanism materially improves long-running agent performance without prohibitive overhead, potentially enabling agents to adapt during a task rather than only between deployments.
watchingconvergesscott: high
Keel’s maintainer claims its released conductor architecture can coordinate agent workflows without embedding a monolithic agent loop, potentially providing a more modular foundation for long-running agent orchestration.
expiredknownscott: low
Prime Intellect claims Prime Agent v0.9.0 makes long-running coding-agent execution more reliable through atomic REPL state restoration, safer interruption handling, and improved preservation of background output and artifacts.
expiredconvergesscott: high
Attemory System claims its released Spire agent can sustain long-horizon Slay the Spire play by delegating deterministic actions to domain tools while using an LLM for reasoning, with about 40% of runs reaching Act 3.
expiredconvergesscott: low
Moadim’s creator claims its open-source local Rust daemon can manage recurring, resumable workflows across multiple agent runners through Git-controlled routines, potentially providing an agent-agnostic orchestration layer for long-running operations.
expiredknownscott: low
DeepMem's maintainers claim their released agent-memory repository combines vector retrieval, BM25, and time decay, potentially giving persistent agents a retrieval layer that accounts for semantic similarity, lexical matches, and recency.
expiredknownscott: low
Anthropic documents support for mid-conversation system messages and tool changes in Claude, potentially allowing agent harnesses to reconfigure instructions and available tools within an ongoing conversation.
watchingknownscott: medium
Pizza Bot's maintainers claim their Apache-2.0 release combines checkpointed DeepAgents/LangGraph runs, scheduling, and durable approval queues in a local-first inbox, enabling users to supervise asynchronous agent work across client disconnects while its backend remains running.
corroboratedconvergesscott: medium
Trigora's Omar Abdelrahman claims its demonstrated Transparent Continuation Checkpointing prototype restores durable executions from live continuations rather than replaying history, potentially decoupling long-running agent recovery costs from accumulated execution history.
watchingconvergesscott: medium
Redditor skeole reports that Qwen3.8-27B on one RTX 3090 sustained a roughly 21-day CUDA-engine development run with about 12 human messages, producing working kernels but no llama.cpp performance win and spending roughly 83 hours on compaction, suggesting local long-running agents are feasible but context maintenance is a major bottleneck.
seed
anglepoiselife claims a deterministic harness ran Qwen3.8-27B unattended for roughly 24 hours on one RTX 5090 to build and browser-test a PostgreSQL, Spring Boot, and React spreadsheet application within a 32K context limit, suggesting local orchestration can sustain substantial multi-file development without hosted inference.
watchingconvergesscott: medium
jbsalles claims the released SelMem engine's selective reconstructive memory — deliberate forgetting, distortion, and sleep-time consolidation — gives LLM entities persistent, path-dependent behavioral divergence, positioning memory design around identity and divergence rather than fidelity for long-running agents.
seedconvergesscott: medium

Trajectory notes