Retrieved article excerpt
Open article Ā· Retrieved 2026-10-05T03:24:03.418036+00:00
# OptChat: an endless chat where the AI remembers everything
## What this is
AI agents forget. A chat session fills up, gets compacted or thrown away,
and the next one starts from zero. If you work with agents every day, your
life ends up scattered over hundreds of sessions in different tools, and
the agent can't find any of it: not the decision you made last month, not
the algorithm you designed in March, not the PR you forgot to answer.
Long sessions have a second problem: **context rot**. The longer the
context, the worse the model works. Compaction ("summarize and keep
going") throws detail away for good, and the result keeps decaying.
OptChat fixes both with one idea: **the chat history itself is the
memory**, stored as a compressed tree.
- There is ONE chat, and it never ends. Every message (yours, the
agent's replies, its tool calls and their results) is appended to a log
and kept forever, word for word.
- In the background, a cheap model compresses the log into a binary tree
of one-line summaries: each message becomes a line, two adjacent lines
merge into one line covering both, two of those merge again, and so on.
- Each time you send a message, the agent starts FRESH: no leftover
context. It sees a fixed-size "view" (about 64k tokens) of the WHOLE
chat: recent messages one line each, older ones many per line, the
older the coarser. Then it sees your message.
- When a line is too vague, the agent "zooms": it opens the line into the
two lines it was made from, down to the original message. Any fact from
your whole history is a few zooms away.
What you get:
- **Infinite context with a constant size.** Nothing is ever deleted; only
the resolution of the distant past fades.
- **No context rot and no manual compaction.** Every turn starts clean.
- **Low cost.** Most tokens of an agent turn are its own tool loop, and
that is prompt-cached. Across turns, the view changes only near its
end, so most of it is cached too.
- **Your instructions stick.** Corrections and preferences you give in
chat stay in the memory (the compressor ranks your own words first), so
most of an AGENTS.md becomes unnecessary: change your mind in a message
and the latest ruling wins.
- **You can browse it too.** The tree is a plain file you can walk from
the root down to any message.
OptChat grew out of OptMem (github.com/VictorTaelin/OptMem), a memory
tool: an append-only log of short notes plus a summary tree, which an
agent reads at the start of each session. OptChat turns that around:
instead of a tool the agent calls, the memory IS the chat, and the
harness builds every turn from it.
**The rest of this document is a complete technical spec, for anyone
(human or AI) who wants to build it.** It describes a working
implementation in full, including the reasons behind each choice. Many
choices that look natural are wrong (they break the cache, or make the
memory decay, or make the agent act on partial text), and the spec says
which and why. If you are an AI building this: follow it exactly, and
when you deviate, have a reason that the spec does not already refute.
---
# Technical specification
## 1. Overview
Components:
1. **The log** (ROOT): every message, verbatim, append-only.
2. **The tree**: summaries. Node `(l, i)` covers messages
`[iĀ·2^l, (i+1)Ā·2^l)`. Level 0 summarizes one message; level `l > 0`
merges its two children `(l-1, 2i)` and `(l-1, 2i+1)`.
3. **The compactor**: a background worker that builds tree nodes with a
cheap model, in a strict order.
4. **The view**: a list of tree nodes that covers the whole chat, oldest
first, kept under a byte budget, changed incrementally.
5. **The turn loop**: each user message starts a fresh model call whose
input is `[system prompt] [view] [new message]`, plus tools `zoom` and
`date`.
6. **Caching**: the request layout and cache breakpoints that make all of
this cheap.
Constants used in the reference implementation:
| name | value | meaning |
| --- | --- | --- |
| `NODE` | 512 bytes | target size of one summary line |
| `VIEW` | 128,000 bytes | budget of the view (ā 62-64k tokens) |
| `JOBS` | 8 | compactor calls running at once |
| `TRIES` | 5 | attempts per node to get under `NODE` |
| `RETRY` | 10 s | wait before retrying a failed node |
| `CAP` | 30,000 chars | max size of one tool result (head + tail kept) |
| `MARKS` | 50,000 / 80,000 / 100,000 chars | cache breakpoints inside the view |
All sizes are **UTF-8 bytes** (or characters, for the cache marks), never
tokens: a tokenizer changes between models, a byte count never does.
## 2. Storage
Two append-only streams in one directory:
```
chat/
main/YYYY-MM-DD.jsonl one message per line: {i, kind, text, size, date}
tree/YYYY-MM-DD.jsonl one node per line: {l, i, text, size}
```
- `i` is the message index (0, 1, 2, ...), its permanent id. `kind` is one
of: `user` (the user's words; also subagent reports, see §9), `talk`
(the agent's replies), `tool` (the agent's tool calls, as text: name and
JSON input), `echo` (tool results), `note` (memories imported from an
older system). `date` is ISO time. `size` = bytes of `kind + ": " + text`.
- A line goes to the file of the local day it was written. Files split by
day only to keep them manageable; the ids are global.
- **Durability**: each line is written with one `write` and then `fsync`,
before the function returns. A crash loses nothing written.
- **Torn lines**: at load, a line that is not valid JSON (a crash
mid-write) is reported and skipped, and a file not ending in `\n` gets
one appended, so the next write starts on its own line.
- **One writer**: two processes on the same chat would corrupt it. Hold a
lock for the life of the process. The reference uses a Unix socket: the
process listens on `lock`; a second process that can connect to it
exits; a socket that refuses connections is stale (the OS frees it when
the owner dies), so it is deleted and taken over. No PID files, no
timeouts.
- The tree is a cache in principle (rebuildable from the log), but it
costs model calls to rebuild, so it is stored and never recomputed.
Never edit or delete anything in these files. The log is history.
Thoughts (model reasoning) are shown to the user but **never logged**.
Reason: the compactor would have to summarize them, and Claude Sonnet's
reasoning-extraction safeguard refused the compactor on thoughts in 112
of 630 first tries; without thoughts, 0 of 523. Thoughts also add little
that the replies and tool calls don't already show.
## 3. The tree
```
node(0, i) = message i, in ⤠NODE bytes
node(l, i) = merge(node(l-1, 2i), node(l-1, 2i+1)), in ⤠NODE bytes
covers(l,i) = messages [iĀ·2^l, (i+1)Ā·2^l)
```
The tree is **purely binary**. (OptMem built small blocks of up to 16
memories straight from the raw notes, and only bigger ones from two
halves. That had no real justification and made the structure confusing;
OptChat dropped it. Every parent comes from exactly its two children.)
**Free nodes.** If the source already fits in `NODE` bytes, it IS the
node, with no model call:
- level 0: `kind + ": " + text` of a short message, verbatim (so short
user messages stay word for word forever, up to the level where they
get merged);
- level > 0: `childA + "\n" + childB`, if that fits.
**Why 512 bytes.** It started at 128 and the model couldn't write useful
lines that short (and overshot often). OptMem used 280. 512 bytes is a
dense paragraph: room for several items with names and numbers. With an
average real line around 250 bytes, a 128 KB view holds ~500 lines.
**Addressing.** A node is named `id+n`: `id` = its first message, `n` =
`2^l` = how many messages it covers. So `node(l, i)` is
`(iĀ·2^l)+(2^l)`. Real message ids, not tree coordinates: the agent reads
`2184+8` in the view and calls `zoom(2184, 8)` directly.
## 4. The compactor
### 4.1 When a node is built
A background loop ("pump") scans all levels and starts every node that:
1. is not built and not already running;
2. has its sources: level 0, the message exists; level > 0, both
children are built;
3. has its whole context summarized: every line of the current view that
lies before the node's end is a built summary (see the `first`
function below).
Up to `JOBS` run at once. After each finishes (or fails), pump again.
Rule 3 is essential, and it gives OptMem's order for free: **messages are
compressed one at a time, in order**, while merges of finished parts run
alongside. The compactor never sees a line that isn't a summary.
```
function first(mem): # first message whose view line is unbuilt
for part in mem.view:
if not built(part): return start(part)
return len(mem.root)
function pump(mem):
T = len(mem.root)
for l = 0 while 2^l <= T:
for i = 0 while (i+1)Ā·2^l <= T:
if len(busy) >= JOBS: return
end = (l == 0) ? i : (i+1)Ā·2^l
if built(l,i) or busy(l,i) or not ready(l,i) or end > first(mem): continue
busy.add((l,i))
build(l,i).then(
ok -> busy.remove((l,i)); fail.remove((l,i)); pump(mem),
err -> report err once per node;
after RETRY: busy.remove((l,i)); pump(mem))
```
For level 0 node `i`, `end = i` means "every view line before message i
is a summary", so node `i` is the first unbuilt one. For a merge,
`end = (i+1)Ā·2^l` means everything up to the node's last message is
summarized.
**Retry**: a failed node waits `RETRY` (10 s) and is tried again, forever;
only its first failure is reported. Don't use long exponential backoff:
the next turn waits for these summaries (§6), so the compactor must catch
up as fast as possible.
### 4.2 What a compactor call sees
One call per node, with no tools, on a cheap model (the reference uses
Claude Sonnet at medium effort; at low effort it overshot the size limit
much more). Input:
- **system**: the `COMPACT` prompt (§4.4), constant.
- **user message**, two text blocks:
1. **Context**: the current view's lines up to the node, wrapped in
`<chat> ... </chat>`. For a level-0 node: lines covering messages
before it (the message itself comes whole in block 2). For a merge:
lines up to the node's last message.
2. **The step**:
```
For scale, this line is exactly 512 bytes:
<SCALE: a realistic summary line of exactly 512 bytes>
Compress this message into one line, in at most 512 bytes:
<kind>: <the message, whole, newlines kept>
```
or, for a merge:
```
For scale, this line is exactly 512 bytes:
<SCALE>
Merge these two lines into one, in at most 512 bytes:
<child A text, newlines flattened to spaces>
<child B text, newlines flattened to spaces>
```
Why each piece:
- **The context block.** The first version gave the compactor only the
message (or the two lines) and nothing else. A summarizer that doesn't
know what's going on writes useless summaries: it can't resolve "do
it", "the other one", "that file". With the view, it knows the project,
the people, the open question, and can even recover detail its input
lost. It is ~64k tokens per call, but it is the same prefix across
calls, so put it first and let it cache.
- **NO IDS anywhere in a compactor call.** View lines are shown bare
(text only, one per line), and the step's lines too. When lines were
shown as `id+n|text`, the model copied the format and began its own
output with an id (6 of 16 tries on one big message). Without ids:
0 of 32. The `<chat>` lines have no markers at all, and the two lines to
merge are written out again, whole, under the instruction, so the model
never has to "find" them.
- **SCALE.** Models can't count bytes. A real example line of exactly
`NODE` bytes gives them a sense of the size. Use a realistic, dense,
multi-item line, tagged with kinds like a real summary.
- **The message goes whole.** Never truncate the compactor's input. Tool
results