2026-10-11 18:01 UTC

Victor Taelin claims his OptChat setup — the entire chat history kept verbatim in an append-only log, compressed in the background into a binary tree of 512-byte summary lines, with every turn served a fixed ~64k-token zoomable view and fresh context — gives agents unbounded, non-decaying memory without context rot or manual compaction, and the pattern becomes a real episode if other builders replicate his published spec and adopt it, while a quiet fade closes it.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagent-memory agent-harnesses llm-toolingVictor Taelin

What is this?

Victor Taelin — a Brazilian programmer and founder of Higher Order Company, previously known for the HVM runtime, Bend, and Kind — has published a first-party writeup titled 'OptChat: an endless chat where the AI remembers everything — HOW TO REPLICATE MY SETUP', describing a setup where the full chat history is kept verbatim in an append-only log while a background process compresses it into a binary tree of 512-byte summary lines, with every turn served a fixed ~64k-token zoomable view and fresh context; he claims this yields unbounded, non-decaying memory without context rot or manual compaction. Notably, the supplied web snippets do not include the OptChat post itself — the only name matches ('OptiChat') are an unrelated academic dialogue system for optimization models — so the technical specifics rest on the first-party evidence rather than independent corroboration. What the snippets do establish is Taelin's standing and a relevant prior: his GitHub lists an older viral post, 'GPT-4: Increase context via compression', showing compression-for-context is a theme he has been pushing since the GPT-4 era, and his starred projects show active interest in high-performance LLM inference.

Why it matters to Scott

Taelin's published OptChat spec — verbatim append-only history, background compression into a navigable summary tree, fixed ~64k zoomable view served on fresh context — independently converges with Scott's Keep-the-Bronze / Read-Ladder / prefix-caching-economics stack and is nearly a reimplementation of his own pointer-backed transcript-compression concept, giving dated receipts from a high-signal builder plus a replicable spec for his search-conversations and ask harnesses. The overclaim to press is that 'non-decaying memory without context rot' doesn't obviously survive Scott's own canon: 512-byte auto-summaries are exactly where hallucinated consolidation bites, and a fixed resident window is still an attention budget — a concrete, publishable critique that stands even before replication decides whether the episode becomes real.
ip:concept.keep-the-bronzeip:concept.prefix-caching-economicsip:concept.read-ladderip:concept.context-rotip:concept.hallucinated-consolidationip:concept.transcript-distillationip:framework.long-running-agentsdev:concept.pointer-backed-transcript-compressionradar:concept.agent-memoryradar:concept.context-compactionradar:concept.context-managementradar:concept.long-running-agentsradar:concept.persistent-memoryradar:lossless-temporal-agent-memoryradar:futureos-context-compaction-recallradar:toucan-transcript-session-memory
queries asked of Scott's wikis
  • context rot degradation long agent sessions
  • memory compaction summarization harness strategy
  • agent-maintained wiki vs automatic summaries
  • verbatim transcript append-only agent memory
  • hierarchical summary tree progressive disclosure
  • fixed context prefix prompt caching economics

Measured heat

now 0 pts/hpeak 22 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 157h
points/hour across evidence Ā· reading as of 2026-10-12 02:59:37.977291+11:00 Ā· deterministic, not a model opinion

How the heat travelled

10-05 03:24 (minted)⭐ origin echo-reconstructed"OptChat: an endless chat where the AI remembers everything — HOW TO REPLICATE MY SETUP": one chat that never ends, every message appended t
Victor Taelin on github (echo) Ā· attributed from hn.story.49960158 Ā· published time unknown
—
10-05 02:44first on hacker news Ā· published Ā· lag ?OptChat: An endless chat where the AI remembers everything
simonpure
—
10-06 22:56first on r/LocalLLaMA Ā· published Ā· lag ?pi-optchat: never compact again - endless chat as a memory tree
voxvoxboy
—
10-05 02:44amplified on hacker newshn.story.49960158
simonpure
peak 1 Ā· 0 comments Ā· 2% of case engagement
10-06 22:56amplified on r/LocalLLaMAreddit.post.1wzgv0u
voxvoxboy
peak 2 Ā· 2 comments Ā· 4% of case engagement
10-06 23:04amplified on r/LocalLLaMA šŸ‘‘reddit.post.1wzh1i7
TheVoxcraft
peak 65 Ā· 41 comments Ā· 95% of case engagement
10-05 03:21our radar first saw it Ā· lag ?discovery anchor: hn.story.49960158—
pace: p68 vs 1247 stories at the 96h mark (now 157h old) — ahead of openai-ignored-security-warnings (1.0x), behind artificial-analysis-mobile-llm-benchmark (1.0x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOptChat: An endless chat where the AI remembers everything
Retrieved article excerpt

Open article Ā· Retrieved 2026-10-05T03:24:03.418036+00:00

# OptChat: an endless chat where the AI remembers everything

## What this is

AI agents forget. A chat session fills up, gets compacted or thrown away,
and the next one starts from zero. If you work with agents every day, your
life ends up scattered over hundreds of sessions in different tools, and
the agent can't find any of it: not the decision you made last month, not
the algorithm you designed in March, not the PR you forgot to answer.

Long sessions have a second problem: **context rot**. The longer the
context, the worse the model works. Compaction ("summarize and keep
going") throws detail away for good, and the result keeps decaying.

OptChat fixes both with one idea: **the chat history itself is the
memory**, stored as a compressed tree.

- There is ONE chat, and it never ends. Every message (yours, the
  agent's replies, its tool calls and their results) is appended to a log
  and kept forever, word for word.
- In the background, a cheap model compresses the log into a binary tree
  of one-line summaries: each message becomes a line, two adjacent lines
  merge into one line covering both, two of those merge again, and so on.
- Each time you send a message, the agent starts FRESH: no leftover
  context. It sees a fixed-size "view" (about 64k tokens) of the WHOLE
  chat: recent messages one line each, older ones many per line, the
  older the coarser. Then it sees your message.
- When a line is too vague, the agent "zooms": it opens the line into the
  two lines it was made from, down to the original message. Any fact from
  your whole history is a few zooms away.

What you get:

- **Infinite context with a constant size.** Nothing is ever deleted; only
  the resolution of the distant past fades.
- **No context rot and no manual compaction.** Every turn starts clean.
- **Low cost.** Most tokens of an agent turn are its own tool loop, and
  that is prompt-cached. Across turns, the view changes only near its
  end, so most of it is cached too.
- **Your instructions stick.** Corrections and preferences you give in
  chat stay in the memory (the compressor ranks your own words first), so
  most of an AGENTS.md becomes unnecessary: change your mind in a message
  and the latest ruling wins.
- **You can browse it too.** The tree is a plain file you can walk from
  the root down to any message.

OptChat grew out of OptMem (github.com/VictorTaelin/OptMem), a memory
tool: an append-only log of short notes plus a summary tree, which an
agent reads at the start of each session. OptChat turns that around:
instead of a tool the agent calls, the memory IS the chat, and the
harness builds every turn from it.

**The rest of this document is a complete technical spec, for anyone
(human or AI) who wants to build it.** It describes a working
implementation in full, including the reasons behind each choice. Many
choices that look natural are wrong (they break the cache, or make the
memory decay, or make the agent act on partial text), and the spec says
which and why. If you are an AI building this: follow it exactly, and
when you deviate, have a reason that the spec does not already refute.

---

# Technical specification

## 1. Overview

Components:

1. **The log** (ROOT): every message, verbatim, append-only.
2. **The tree**: summaries. Node `(l, i)` covers messages
   `[iĀ·2^l, (i+1)Ā·2^l)`. Level 0 summarizes one message; level `l > 0`
   merges its two children `(l-1, 2i)` and `(l-1, 2i+1)`.
3. **The compactor**: a background worker that builds tree nodes with a
   cheap model, in a strict order.
4. **The view**: a list of tree nodes that covers the whole chat, oldest
   first, kept under a byte budget, changed incrementally.
5. **The turn loop**: each user message starts a fresh model call whose
   input is `[system prompt] [view] [new message]`, plus tools `zoom` and
   `date`.
6. **Caching**: the request layout and cache breakpoints that make all of
   this cheap.

Constants used in the reference implementation:

| name | value | meaning |
| --- | --- | --- |
| `NODE` | 512 bytes | target size of one summary line |
| `VIEW` | 128,000 bytes | budget of the view (ā‰ˆ 62-64k tokens) |
| `JOBS` | 8 | compactor calls running at once |
| `TRIES` | 5 | attempts per node to get under `NODE` |
| `RETRY` | 10 s | wait before retrying a failed node |
| `CAP` | 30,000 chars | max size of one tool result (head + tail kept) |
| `MARKS` | 50,000 / 80,000 / 100,000 chars | cache breakpoints inside the view |

All sizes are **UTF-8 bytes** (or characters, for the cache marks), never
tokens: a tokenizer changes between models, a byte count never does.

## 2. Storage

Two append-only streams in one directory:

```
chat/
  main/YYYY-MM-DD.jsonl   one message per line: {i, kind, text, size, date}
  tree/YYYY-MM-DD.jsonl   one node per line:    {l, i, text, size}
```

- `i` is the message index (0, 1, 2, ...), its permanent id. `kind` is one
  of: `user` (the user's words; also subagent reports, see §9), `talk`
  (the agent's replies), `tool` (the agent's tool calls, as text: name and
  JSON input), `echo` (tool results), `note` (memories imported from an
  older system). `date` is ISO time. `size` = bytes of `kind + ": " + text`.
- A line goes to the file of the local day it was written. Files split by
  day only to keep them manageable; the ids are global.
- **Durability**: each line is written with one `write` and then `fsync`,
  before the function returns. A crash loses nothing written.
- **Torn lines**: at load, a line that is not valid JSON (a crash
  mid-write) is reported and skipped, and a file not ending in `\n` gets
  one appended, so the next write starts on its own line.
- **One writer**: two processes on the same chat would corrupt it. Hold a
  lock for the life of the process. The reference uses a Unix socket: the
  process listens on `lock`; a second process that can connect to it
  exits; a socket that refuses connections is stale (the OS frees it when
  the owner dies), so it is deleted and taken over. No PID files, no
  timeouts.
- The tree is a cache in principle (rebuildable from the log), but it
  costs model calls to rebuild, so it is stored and never recomputed.

Never edit or delete anything in these files. The log is history.

Thoughts (model reasoning) are shown to the user but **never logged**.
Reason: the compactor would have to summarize them, and Claude Sonnet's
reasoning-extraction safeguard refused the compactor on thoughts in 112
of 630 first tries; without thoughts, 0 of 523. Thoughts also add little
that the replies and tool calls don't already show.

## 3. The tree

```
node(0, i)  = message i, in ≤ NODE bytes
node(l, i)  = merge(node(l-1, 2i), node(l-1, 2i+1)), in ≤ NODE bytes
covers(l,i) = messages [iĀ·2^l, (i+1)Ā·2^l)
```

The tree is **purely binary**. (OptMem built small blocks of up to 16
memories straight from the raw notes, and only bigger ones from two
halves. That had no real justification and made the structure confusing;
OptChat dropped it. Every parent comes from exactly its two children.)

**Free nodes.** If the source already fits in `NODE` bytes, it IS the
node, with no model call:

- level 0: `kind + ": " + text` of a short message, verbatim (so short
  user messages stay word for word forever, up to the level where they
  get merged);
- level > 0: `childA + "\n" + childB`, if that fits.

**Why 512 bytes.** It started at 128 and the model couldn't write useful
lines that short (and overshot often). OptMem used 280. 512 bytes is a
dense paragraph: room for several items with names and numbers. With an
average real line around 250 bytes, a 128 KB view holds ~500 lines.

**Addressing.** A node is named `id+n`: `id` = its first message, `n` =
`2^l` = how many messages it covers. So `node(l, i)` is
`(iĀ·2^l)+(2^l)`. Real message ids, not tree coordinates: the agent reads
`2184+8` in the view and calls `zoom(2184, 8)` directly.

## 4. The compactor

### 4.1 When a node is built

A background loop ("pump") scans all levels and starts every node that:

1. is not built and not already running;
2. has its sources: level 0, the message exists; level > 0, both
   children are built;
3. has its whole context summarized: every line of the current view that
   lies before the node's end is a built summary (see the `first`
   function below).

Up to `JOBS` run at once. After each finishes (or fails), pump again.

Rule 3 is essential, and it gives OptMem's order for free: **messages are
compressed one at a time, in order**, while merges of finished parts run
alongside. The compactor never sees a line that isn't a summary.

```
function first(mem):            # first message whose view line is unbuilt
  for part in mem.view:
    if not built(part): return start(part)
  return len(mem.root)

function pump(mem):
  T = len(mem.root)
  for l = 0 while 2^l <= T:
    for i = 0 while (i+1)Ā·2^l <= T:
      if len(busy) >= JOBS: return
      end = (l == 0) ? i : (i+1)Ā·2^l
      if built(l,i) or busy(l,i) or not ready(l,i) or end > first(mem): continue
      busy.add((l,i))
      build(l,i).then(
        ok   -> busy.remove((l,i)); fail.remove((l,i)); pump(mem),
        err  -> report err once per node;
                after RETRY: busy.remove((l,i)); pump(mem))
```

For level 0 node `i`, `end = i` means "every view line before message i
is a summary", so node `i` is the first unbuilt one. For a merge,
`end = (i+1)Ā·2^l` means everything up to the node's last message is
summarized.

**Retry**: a failed node waits `RETRY` (10 s) and is tried again, forever;
only its first failure is reported. Don't use long exponential backoff:
the next turn waits for these summaries (§6), so the compactor must catch
up as fast as possible.

### 4.2 What a compactor call sees

One call per node, with no tools, on a cheap model (the reference uses
Claude Sonnet at medium effort; at low effort it overshot the size limit
much more). Input:

- **system**: the `COMPACT` prompt (§4.4), constant.
- **user message**, two text blocks:
  1. **Context**: the current view's lines up to the node, wrapped in
     `<chat> ... </chat>`. For a level-0 node: lines covering messages
     before it (the message itself comes whole in block 2). For a merge:
     lines up to the node's last message.
  2. **The step**:

     ```
     For scale, this line is exactly 512 bytes:
     <SCALE: a realistic summary line of exactly 512 bytes>

     Compress this message into one line, in at most 512 bytes:
     <kind>: <the message, whole, newlines kept>
     ```

     or, for a merge:

     ```
     For scale, this line is exactly 512 bytes:
     <SCALE>

     Merge these two lines into one, in at most 512 bytes:
     <child A text, newlines flattened to spaces>
     <child B text, newlines flattened to spaces>
     ```

Why each piece:

- **The context block.** The first version gave the compactor only the
  message (or the two lines) and nothing else. A summarizer that doesn't
  know what's going on writes useless summaries: it can't resolve "do
  it", "the other one", "that file". With the view, it knows the project,
  the people, the open question, and can even recover detail its input
  lost. It is ~64k tokens per call, but it is the same prefix across
  calls, so put it first and let it cache.
- **NO IDS anywhere in a compactor call.** View lines are shown bare
  (text only, one per line), and the step's lines too. When lines were
  shown as `id+n|text`, the model copied the format and began its own
  output with an id (6 of 16 tries on one big message). Without ids:
  0 of 32. The `<chat>` lines have no markers at all, and the two lines to
  merge are written out again, whole, under the instruction, so the model
  never has to "find" them.
- **SCALE.** Models can't count bytes. A real example line of exactly
  `NODE` bytes gives them a sense of the size. Use a realistic, dense,
  multi-item line, tagged with kinds like a real summary.
- **The message goes whole.** Never truncate the compactor's input. Tool
  results 
simonpure10
🟧 echo.github ⭐"OptChat: an endless chat where the AI remembers everything — HOW TO REPLICATE MY SETUP": one chat that never ends, every message appended tVictor Taelin——
🟠 redditpi-optchat: never compact again - endless chat as a memory tree
LocalLLaMA
TheVoxcraft6541
🟠 redditpi-optchat: never compact again - endless chat as a memory tree
LocalLLaMA
voxvoxboy12

Interpretation history

Decision trace