Retrieved article excerpt
Open article · Retrieved 2026-09-17T02:22:32.502322+00:00
# OVERLORD
An agent hypervisor — the trust kernel for delegated computing.
[OVERLORD demo: rm -rf inside a transactional shell, then rollback — everything comes back](https://github.com/B1tR0n1n/overlord/blob/master/assets/demo.gif)
Hand a program — or an AI agent — a fully writable copy of a directory. Let it
run. Then review every change it made as a hashed, attributed manifest and
**commit or roll back**, all or nothing. Transactional isolation, provenance,
and arbitration for untrusted execution — in dependency-free Python, on the
kernel's own primitives. It ships with a built-in agent (`overlord agent`) that
runs a model with its hands jailed, so you can watch the whole loop happen
inside the transaction and sign off on the diff. For people who just want to
use it, `overlord ui` opens a chat workspace where you talk to that agent and
press Commit when you like what it did — nothing on disk changes until you do.
## Thesis
Computing is transitioning to a new operator: machine agents. The OS has no native
concept of a machine actor. Every agent today runs with its principal's full authority
on infrastructure that cannot distinguish the principal's intent from the agent's
behavior. Every harness vendor duct-tapes around this independently and badly.
OVERLORD is the missing layer between the agent harness and the operating system.
Not an AI. Model-agnostic. A boring, load-bearing primitive — the SQLite pattern,
not the Windows pattern.
## The Four Primitives
1. **Capability, not identity** — an agent receives a scoped grant (paths, budget,
network egress, time window), not a user account. Commander's intent expressed
as kernel-enforced constraints.
2. **Provenance** — every mutation traceable to actor, instruction, and the reasoning
artifact that caused it. A flight recorder for machine action.
3. **Reversibility** — agent action is transactional: snapshot, execute, inspect,
commit or roll back. The keystone. Delegation is blocked on "what if it breaks
something"; this removes the question.
4. **Arbitration** — when N agents contend for a resource, authority is scheduled
the way CPU is scheduled.
## Install
```
sudo bash packaging/install.sh
overlord doctor
```
Installs the engine to `/usr/local/lib/overlord/`, a compiled ELF launcher to
`/usr/local/bin/overlord` (the AppArmor attachment point), the AppArmor profile
that enables the kernel backend on Ubuntu 24.04+, and the runtime deps
(`fuse-overlayfs`, `strace`).
### Windows and macOS
The engine is Linux kernel machinery — overlayfs, user namespaces,
`mount(2)` — so OVERLORD does not run natively on Windows or macOS, and a
port would be a different product. It runs in the Linux those systems
ship or host:
- **WSL2** (Windows): `wsl --install -d Ubuntu`, then the install above
inside Ubuntu (the AppArmor step skips itself) and `overlord ui`; WSL2
forwards localhost, so a Windows browser opens `http://127.0.0.1:7777`.
Keep the project folder in the Linux filesystem (`~/projects/…`, seen
from Windows as `\\wsl$\Ubuntu\home\…`), not under `/mnt/c/`, where
the overlay is slow and `doctor` may fall back to the fuse backend.
- **Docker Desktop** (Windows, macOS): the `Dockerfile` builds an image on
the fuse backend; its header has the run line (`--device /dev/fuse --cap-add SYS_ADMIN`, a data volume, accounts and a certificate first).
`overlord doctor` is the ground truth on any machine: it names the
backend it found and what each grant will mean there.
## Use
```
overlord run -t /srv/app -- some-agent --do-things # transactional execution
overlord run --jail --net none --timeout 300 -t /srv/app -- <cmd> # scoped grants
overlord run --manifest cap.json -t /srv/app -- <cmd> # grants from file
overlord run --trace -t /srv/app -- <cmd> # + syscall flight recorder (strace)
overlord run --trace ebpf -t /srv/app -- <cmd> # kernel-side recorder (root-only)
overlord run --merge-base -t /srv/app -- <cmd> # keep base copy for commit --merge
overlord shell -t /srv/app # interactive transactional shell
overlord agent -t /srv/app "add a Makefile with a test target" # jailed by default
overlord agent --net none -t /srv/app "<task>" # ...and offline too
overlord agent --net proxy --net-allow github.com --net-allow "*.pypi.org" -t /srv/app "<task>" # recorded, allowlisted egress
overlord agent --no-jail -t /srv/app "<task>" # opt out: tools reach the real fs
overlord agent --provider openai-compatible --base-url http://127.0.0.1:11434/v1 \
--model llama3 -t /srv/app "<task>" # a local model, same jail
overlord agent --effort xhigh --max-tokens 64000 -t /srv/app "<task>" # generation knobs
overlord models --provider anthropic # what the endpoint serves right now
overlord mcp add github --command npx --arg -y --arg @modelcontextprotocol/server-github \
--env GITHUB_TOKEN=… # register an MCP connector (stdio)
overlord mcp add docs --url https://host/mcp --header 'Authorization: Bearer …' # (http)
overlord agent --connector github -t /srv/app "<task>" # grant it; actions ask you first
overlord agent --audit -t /srv/app ["focus"] # containment audit of its own jail
overlord memory show -t /srv/app # what the agent is told before message one
overlord memory user --add "Prefers pytest." # a note about you, across every folder
overlord memory accept <session> --all # keep the notes an agent proposed
overlord users add alice --role admin # accounts: the UI now asks who you are
overlord tls selfsign --host overlord.lan # a certificate for --bind
overlord sso set --issuer https://login.example.com --client-id … --client-secret-stdin \
--domain example.com --admin [email protected] # OpenID Connect; accounts provisioned on sign-in
overlord ui --bind 0.0.0.0 --tls-cert ~/.overlord/tls/cert.pem --tls-key ~/.overlord/tls/key.pem
overlord skills add skills/python-testing # packaged know-how, loaded when it fits
overlord skills new release -t /srv/app # a project skill: part of the tree, reviewed like code
overlord webhooks add team https://hooks.slack.com/… # tell the channel when work waits for a person
overlord secrets set-command 'vault kv get -field=value secret/overlord/{name}' # then secret://NAME anywhere
overlord export <session> -o review.ovl # one signed file: record, retained versions, pending changes
overlord import review.ovl -t /srv/app # replay its changes here as a new pending session
overlord cost # what the models spent, by model / account / day
overlord cost budget --day-usd 20 --month-usd 500 # lines no conversation crosses
overlord audit verify # walk the signed chain
overlord audit checkpoint pin.json # witness the head off-box
overlord audit verify --pin pin.json # prove the live log still carries it
overlord audit witness https://witness.example/log --auto # send a signed head off-box on every act
overlord audit verify --witness # check the live log against the witnessed head
overlord audit verify # the hash chain of every consequential act
overlord gc --dry-run # what retention would prune
overlord sessions # pending/committed history with command provenance
overlord diff <session> # added / modified / deleted / replaced-dir
overlord log <session> # per-change sha256 before -> after, syscall count
overlord savepoints <session> # the layer stack: one savepoint per command that wrote
overlord rewind <session> --to 3 # discard everything above savepoint @3
overlord resume <session> --note "..." # agent carries on from there, reading the note
overlord fork <session> --at 3 # a second continuation of the same moment, as a new session
overlord compare <sess-a> <sess-b> # where two continuations diverge, per path
overlord review <session> --provider openai # a second model countersigns the diff
overlord commit <session> # verify no external drift, replay onto real tree
overlord commit --countersigned <sess> # ...only with a fresh approval on record
overlord commit --drop tool:shell <sess> # replay all but the shell tool's layers
overlord commit --only turn:2-4 <sess> # replay only what turns 2–4 did
overlord commit --merge <sess> # three-way merge non-overlapping drift (needs --merge-base)
overlord commit --force <sess> # commit despite drift (explicit override)
overlord rollback <session> # discard — target byte-identical
overlord blame <path> # which session, turn, tool call, instruction put each line here
overlord doctor # backend / dependency diagnostics
```
The wrapped command sees a fully writable tree and exits believing everything
happened. Nothing touches the real tree until `commit`. Commit re-verifies the
snapshot fingerprints (size + mtime\_ns of every file) and **refuses to clobber
external changes** made while the session was pending. If the session was run
with `--merge-base`, `commit --merge` three-way merges non-overlapping drift
(git merge-file against the kept base) and still refuses overlapping edits.
## Models
The model thinks on your machine; only its hands are jailed. Every way to
reach one is in `providers.py`, stdlib only, behind one contract, and the
workspace, the CLI, the daemon and the SDK all share it:
| provider | reaches | auth |
| --- | --- | --- |
| `anthropic` | Claude, native Messages API: streaming, adaptive thinking, `effort`, refusal handling, opt-in server-side refusal fallbacks | `ANTHROPIC_API_KEY` or the key store |
| `openai` | OpenAI Responses API (`/v1/responses`): streaming, `reasoning.effort`, function tools with reasoning, encrypted reasoning carried across tool turns, nothing stored server-side; `--api chat` for the old shape | `OPENAI_API_KEY` |
| `azure` | Azure OpenAI deployments (`--base-url https://<resource>.openai.azure.com`, deployment as the model, `--azure-api-version`) | `AZURE_OPENAI_API_KEY` |
| `openai-compatible` | anything speaking the Chat Completions shape behind a base URL: Ollama, vLLM, LiteLLM, Groq, Together, your gateway; `--api responses` once it grows the new shape | optional |
| `gemini` | Google Gemini REST: streaming, function calling | `GEMINI_API_KEY` |
Every provider takes a **base URL** and **extra headers**, which is how a
proxy or an enterprise gateway sits in front of it. `overlord models` lists
what an endpoint serves right now, and the workspace's Settings shows the same
list. Keys live in `~/.overlord/keys.json` (mode 600); environment variables
win over the store.
**Generation knobs** — `--max-tokens`, `--temperature`, `--top-p`, `--stop`,
`--effort low|medium|high|xhigh|max`, `--thinking summarized|off`,
`--api responses|chat`, `--system` (appended instructions), `--no-stream`,
`--no-fallbacks` — are
model-aware: nothing is sent unless you set it. That matters because the
current Claude family rejects `temperature` and `top_p` outright and takes
its depth from `effort`; blank means the model's own default. Replies stream
as they are generated (the CLI prints them live; the workspace renders them
into the bubble), tool inputs that stream in are parsed strictly and handed
back as an error rather than run when malformed, and a reply cut off at
`max_tokens` or declined by a safety classifier never executes its tool
calls. The knobs and the endpoint a conversation ran with are recorded on
the session, so `resume` uses the same model the same way.
## Connectors (MCP)
`mcp.py` is a Model Context Protocol client, stdlib only: stdio servers
(a command) and streamable-HTTP servers (a URL), configured in
`~/.overlord/mcp.json`. Th