2026-10-11 17:13 UTC

Autonomous Production claims its released AutoBot harness combines persistent task graphs, disk-backed memory, and separate completion validation with native ChatGPT to improve long-running computer-use work, reporting 32.41% OSWorld 2.0 accuracy and 50.70% AssistantBench accuracy.

state: seedheat: mediumuncertainty: mediumknownscott: lowagent-harnesses agent-memory long-running-orchestration computer-useAutonomous Productiondemeyer1

What is this?

The supplied case describes AutoBot as a released harness from Autonomous Production that uses native ChatGPT, persistent task graphs, disk-backed memory, and separate completion validation for long-running computer-use work; its evidence titles also advertise live voice control. None of the returned web snippets directly documents AutoBot, establishes demeyer1's role, or verifies the claimed 32.41% OSWorld 2.0 and 50.70% AssistantBench accuracies. The OSWorld 2.0 paper snippet does identify lost constraints, incomplete state, and failed outcome verification as failure modes, while AssistantBench's site describes realistic, time-consuming web tasks. These sources establish the evaluation context, not AutoBot's results or their comparability to published scores.

Why it matters to Scott

AutoBot’s claimed durable state and separate completion validation repeat positions already held in Scott’s Long-Running Agents and Same Session Supervision ebook, with operational overlap in Proposal Compiler’s persisted job identity and recovery. No supplied radar hit tracks AutoBot itself, but the material establishes neither a consequential new adopter nor verified benchmark evidence that would extend Scott’s model-plus-harness argument or change his implementation choices.
ip:framework.long-running-agentsip:source.same-session-supervision-ebookip:concept.model-plus-harness-benchmark-unitdev:project.proposalradar:concept.agent-harnessesradar:concept.long-running-orchestrationradar:concept.agent-verificationradar:concept.computer-use-agents
queries asked of Scott's wikis
  • persistent task graphs long-running agent orchestration recovery
  • disk-backed hierarchical agent memory context management
  • independent completion validation agent supervision
  • harness versus model capability long-horizon reliability
  • computer-use benchmark methodology partial credit task completion
  • native ChatGPT automation harness voice intervention

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 575h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 17:44 (minted)⭐ origin echo-reconstructedAutoBot publishes a native-ChatGPT operating harness with persistent state, hierarchical memory, supervision, and benchmark methodology; its
Autonomous Production / demeyer1 on github (echo) · attributed from hn.story.49743478 · published time unknown
—
09-17 16:54first on hacker news · published · lag ?Show HN: AutoBot – live voice control for long-running AI work
demeyer1
—
09-17 16:54amplified on hacker news 👑hn.story.49743478
demeyer1
peak 20 · 3 comments · 100% of case engagement
09-17 17:22our radar first saw it · lag ?discovery anchor: hn.story.49743478—
pace: p54 vs 1032 stories at the 336h mark (now 575h old) — ahead of agentdrive-persistent-shared-storage (1.1x), behind aws-project-spend-limits (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: AutoBot – live voice control for long-running AI work
Retrieved article excerpt

Open article · Retrieved 2026-09-17T17:26:10.529506+00:00

# Autobot

**AutoBot: a self-improving agentic harness that makes frontier AI better at finishing complex knowledge work.**

AutoBot achieved **18.5% higher task completion than the published OpenAI Sol Max baseline**, surpassing **Anthropic’s Claude Opus 5 Max** on OSWorld 2.0, a benchmark of long, multi-application workflows. It also reached **#1 on the official AssistantBench hidden-test leaderboard**. [Results and methodology](https://github.com/demeyer1/Autobot/tree/main/benchmarks)

**Hard workflows become upgrades to the agent itself.** AutoBot repairs its own harness, independently validates the changes, and carries them forward. The next workflow inherits the improvement. **Compounding capability, without retraining the model.**

**Your knowledge outgrows the context window.** Hierarchical memory lives on disk; task-specific retrieval builds the working context. Nightly consolidation integrates new knowledge and corrections. Your agent accumulates institutional memory across projects and conversations.

**Your project can outlive the agent working on it.** Persistent task graphs, atomic checkpoints and independent supervision let a replacement worker resume the assignment. Completion is bound to current requirements and verified destination evidence.

**Local compute makes persistent intelligence economical.** Your CPU handles orchestration, state and integrity checks. Compiled context and reusable proofs reduce repeated inference, directing the model’s budget toward the difficult judgments that move work forward.

Open source. Native ChatGPT on your Mac. Built by [Autonomous Production](https://autoprod.ai). [Get AutoBot](https://github.com/demeyer1/Autobot/releases).

## Benchmarks

### AssistantBench

**#1 on the recorded official hidden-test leaderboard: 50.70% accuracy across 181 tasks.**

[AssistantBench leaderboard with AutoBot’s result highlighted in red](https://github.com/demeyer1/Autobot/blob/main/benchmarks/assistantbench)

AutoAssist was the harness's previous brand name, changed to AutoBot two weeks after this screenshot.

[Results and verification](https://github.com/demeyer1/Autobot/blob/main/benchmarks/assistantbench)

### OSWorld 2.0

**32.41% binary accuracy and 64.28% partial accuracy across 108 tasks.** Final best-valid-per-task aggregate, compared with the published September 10, 2026 leaderboard snapshot.

[OSWorld 2.0 comparison with AutoBot’s result highlighted in red](https://github.com/demeyer1/Autobot/blob/main/benchmarks/osworld-2.0)

[Results and methodology](https://github.com/demeyer1/Autobot/blob/main/benchmarks/osworld-2.0)

## What it changes

- **Privacy by relationship and purpose.** Personal, family and friends, work, and deliberately shared context have separate zones. Nothing moves into `SHARED` automatically.
- **Independent completion.** The local runtime rejects validation under the producer label. The operating contract also requires a separate validator to inspect current evidence, including the real destination when the task changes something outside the workspace.
- **Durable follow-through.** Native ChatGPT Goals keep the active work moving. Autobot keeps the objective, stages, evidence, and recovery state on disk so an interrupted chat does not silently erase the commitment.
- **Clean external actions.** Recipient-visible writes use the signed-in first-party app through Computer Use. The operating contract requires the active ChatGPT workflow to verify the account, destination, and visible content before the action, then check the rendered result for a duplicate, failure, or unwanted AI attribution. If the clean route cannot be verified, the workflow must stop.
- **Tone that stays in its lane.** Communication profiles are separated by channel and audience. A family text profile does not become a work email profile. Autobot stores compact, user-approved patterns rather than raw message archives by default.

## What Autobot includes

Autobot is the combination of two layers:

1. **Native ChatGPT:** the desktop app, local Projects, Voice, Goals, skills and plugins, scheduled work, notifications, and Computer Use.
2. **The Autobot workspace:** the operating contract in `AGENTS.md`, privacy zones, selective memory, project status, a local objective state machine, a one-minute liveness supervisor, external-action policy, and first-time setup.

Autobot is not a separate model, chatbot, or agent gateway. It depends on current ChatGPT capabilities and their plan, region, usage, sandbox, and permission limits. See [Architecture](https://github.com/demeyer1/Autobot/blob/main/docs/ARCHITECTURE.md) and [Permissions](https://github.com/demeyer1/Autobot/blob/main/docs/PERMISSIONS.md).

## Autobot compared with OpenClaw and Hermes

This comparison evaluates Autobot plus native ChatGPT for one person managing personal and work tasks on a Mac. "Better" means a more explicit default for the stated user need, based on the published design; it does not mean a measured advantage in accuracy, speed, or reliability. "Doesn't do" means the specific built-in requirement was not found in the primary documentation reviewed on September 10, 2026. Both alternatives can be extended.

|  | Top 3 shared capabilities: Autobot's approach and user benefits | Top 3 additional Autobot capabilities and user benefits |
| --- | --- | --- |
| [**OpenClaw**](https://github.com/openclaw/openclaw) | **1. Remember context with explicit boundaries.** OpenClaw persists and searches memory. Autobot adds default personal, family/friends, work, and opt-in shared zones. This makes the rules for reusing private context in a work task more explicit. [Memory](https://docs.openclaw.ai/concepts/memory) / [Autobot privacy](https://github.com/demeyer1/Autobot/blob/main/PRIVACY.md).  **2. Track the promised outcome.** OpenClaw records background tasks and delivery state. Autobot assigns each promised output an owner, destination, and completion gate. This gives the user a clearer record of what remains owed after an interruption. [Tasks](https://docs.openclaw.ai/automation/tasks) / [Autobot contract](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#projects-and-unfinished-work).  **3. Check the result of an app action.** OpenClaw supports tools and outbound audit history. Autobot requires account, destination, and content checks before a write, followed by rendered inspection. This adds an explicit check for a wrong destination, failed save, or duplicate action. [Audit](https://docs.openclaw.ai/gateway/audit) / [Autobot writes](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#external-reads-and-writes). | **1. Require a separate completion validator.** Autobot's five-stage process requires current evidence and a validator label different from the producer, with destination readback for external work. This gives users evidence beyond the worker's own success report. [Completion](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#completion-integrity).  **2. Block writes when clean delivery cannot be verified.** Autobot's default policy requires first-party Computer Use and stops a route that forces unwanted attribution. This gives users an explicit publication rule instead of relying on each connector's behavior. [Write policy](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#external-reads-and-writes).  **3. Keep learned tone separate by channel and audience.** Autobot requires authorized examples and compact, separate communication profiles. This helps keep a family-text style from shaping a customer email. [Profiles](https://github.com/demeyer1/Autobot/blob/main/00_CONTEXT/COMMUNICATION-PROFILES.md). |
| [**Hermes**](https://github.com/NousResearch/hermes-agent) | **1. Remember context with purpose-specific retrieval.** Hermes has curated memory, user profiles, and session search. Autobot makes relationship and purpose part of its default retrieval rules. This gives users clearer control over which personal facts may inform professional work. [Memory](https://hermes-agent.nousresearch.com/docs/user-guide/features/memory) / [Autobot privacy](https://github.com/demeyer1/Autobot/blob/main/PRIVACY.md).  **2. Follow work through to its deliverable.** Hermes schedules jobs and delivers their outputs. Autobot also retains the objective, ordered evidence stages, and unfinished outputs. This makes it easier to distinguish a job that ran from a requested artifact that was saved and checked. [Scheduling](https://hermes-agent.nousresearch.com/docs/user-guide/features/cron) / [Autobot completion](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#completion-integrity).  **3. Verify user-visible changes.** Hermes provides tool approvals and execution controls. Autobot adds a standard before-and-after inspection of the signed-in destination. This gives the user a specific check that an authorized action produced the intended visible result. [Security](https://hermes-agent.nousresearch.com/docs/user-guide/security) / [Autobot writes](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#external-reads-and-writes). | **1. Require independent acceptance of every completion stage.** Autobot separates the producer and validator labels and requires fresh destination evidence for external outputs. This makes an unsupported "done" report insufficient under the operating contract. [Completion](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#completion-integrity).  **2. Require a clean first-party write route by default.** Autobot checks the exact account and visible content, then rejects forced-attribution or unverifiable delivery routes. This gives users a consistent rule for communications across services. [Write policy](https://github.com/demeyer1/Autobot/blob/main/AGENTS.md#external-reads-and-writes).  **3. Require audience-specific tone-learning boundaries.** Hermes supports personality and user-style preferences; Autobot specifies separately authorized profiles for each channel and audience. This gives users more explicit control over which examples shape each kind of message. [Hermes memory](https://hermes-agent.nousresearch.com/docs/user-guide/features/memory) / [Autobot profiles](https://github.com/demeyer1/Autobot/blob/main/00_CONTEXT/COMMUNICATION-PROFILES.md). |

These are operating-contract differences, not guarantees of error-free execution. Autobot's privacy zones and validator separation are procedural within one Mac; they are not OS isolation or cryptographic identities. OpenClaw and Hermes offer broader standalone deployment, messaging, and model choices. The absence findings above concern the exact required workflows, not an absence of memory, privacy controls, voice, verification tools, or automation in either project.

## What is new in 0.3.0

The project-local first-time skill now acts as the setup operator. It checks current state, applies safe local defaults, runs supported actions, preserves the earliest unresolved user or capability gate, and returns to the user's original task only after current marker and status readback.

The base folder installs without Node or sign-in. The advanced runtime needs Node 22 or newer. Native app capabilities are discovered separately; no model, account, notification recipient or external write is configured for you. See [capabilities and availability](https://github.com/demeyer1/Autobot/blob/main/docs/CAPABILITIES.md).

## Privacy zones

| Zone | Intended context | Default boundary |
| --- | --- | --- |
| `PRIVATE` | Sensitive personal facts, preferences, health, finances, and private plans | Never shared automatically |
| `FAMILY_FRIENDS` | Relationships, events, logistics, and authorized personal communication patterns | Never used for work without a current explicit need |
| `WORK` | Organizations, projects, teammates, customers, vendors, and professional communication patterns | Never used for personal communication unless explicitly relevant |
| `SHARED` | The minimum facts you deliberatel
demeyer1203
🟧 echo.github ⭐AutoBot publishes a native-ChatGPT operating harness with persistent state, hierarchical memory, supervision, and benchmark methodology; itsAutonomous Production / demeyer1——

Interpretation history

Decision trace