2026-10-11 16:37 UTC

FutureOS claims its originals-first context compaction retained 83% of tested session facts versus 47% for OpenCode and 38% for Codex, suggesting that preserving assistant prose and indexing tool evidence can materially improve long-session recall at higher per-turn context cost.

state: watchingheat: lowuncertainty: highconvergesscott: mediumagent-harnesses context-compaction agent-memory agent-orchestration inference-economicsFutureOSOpenCodeCodex

What is this?

FutureOS — a vendor publishing on its own engineering blog (future-os-blog.github.io) with experiment code in a public repo — ran a reproducible benchmark of what three coding-agent harnesses retain after context compaction: its originals-first strategy (keep assistant prose, index tool evidence) answered 147/178 (83%) of session-fact questions versus 83/178 (47%) for OpenCode and 68/178 (38%) for Codex, with a zero-model-call deterministic tier at 71% and open-book rounds showing models never retrieve archived material unless forced (93.3% when forced). The tradeoff is the largest median context projection (~12.1K tokens) and priciest turn of the four strategies, offset by a claimed 99.8% production prefix-cache hit that cuts a ¥7.98 cold compaction to ¥0.53. Third-party material corroborates the problem shape: an openai/codex GitHub issue reports compaction silently discards 100% of tool outputs (a 'telephone game' decaying to ~6.9% after two rounds), an academic Addressable Recall Compaction paper (arXiv:2607.25066) argues summaries destroy addressable evidence, and OpenCode plugin authors ship archive-and-recall layers precisely because 'the original content stays in the database; the agent just stops being able to see it' — though one write-up still voices the opposing prune philosophy that tool output is 'high-volume and low-value after its immediate use.' What these snippets still don't supply: any independent reproduction of the 178-question methodology or scoring, and no trace of the 'Codex Originals' OpenAI page this reground was meant to audit — so the headline figures remain vendor-authored and the first-party characterization of Codex remains unverified.

Why it matters to Scott

Converges with Scott's pointer-backed transcript compression concept and bears directly on a live design choice: `ask --compact` is deliberately lossy (tool args dropped, results truncated to 100 chars), and FutureOS's 83/47/38 recall split plus the prose-90%-of-questions vs tool-output-96%-of-volume asymmetry are exactly the quantified receipts that would justify reworking `--compact` toward originals-plus-index; the forced-retrieval finding (93.3% open-book when forced) independently supports his answers-not-content posture. But the reground found no trace of the 'Codex Originals' page, so FutureOS's characterization of Codex remains unaudited and the headline figures stay vendor-authored — strong supporting evidence for a position he already holds, not yet a forcing change to what he builds.
dev:concept.pointer-backed-transcript-compressiondev:project.askdev:concept.agent-authored-context-compactiondev:concept.answers-not-contentradar:concept.context-compactionradar:concept.prompt-cachingradar:anthropic-context-compaction-cost-reversalradar:compactdiff-agent-compaction-auditradar:lmstudio-bionic-transcript-introspectionradar:tool-output-compression-proxy
queries asked of Scott's wikis
  • ask --compact lossy summary design originals discarded
  • prefix-cache-aware compaction prompt caching agent sessions
  • agent memory pointer-backed original transcript tool-output index
  • compaction recall evaluation session-fact benchmark harness
  • open-book retrieval agent never retrieves forced recall behavior
  • local model context cache economics long-session cost

Measured heat

now 0 pts/hpeak 6 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 14:00⭐ origin echo-reconstructedThe primary artifact is FutureOS's own experiment commit and its accompanying closed/open-book compaction reports. The report states: “All 2
Hui Chen on github (echo) · attributed from hn.story.49795948
—
09-22 02:04first on hacker news · published · +108.1hContext compaction, measured: FutureOS vs. Codex vs. OpenCode
thetechlead
—
09-22 16:32first on r/LocalLLaMA · published · +122.5hI built a cache-friendly context compacting plugin for OpenCode
schennardo
—
09-22 02:04amplified on hacker newshn.story.49795948
thetechlead
peak 1 · 0 comments · 4% of case engagement
09-22 16:32amplified on r/LocalLLaMA 👑reddit.post.1wnedrc
schennardo
peak 17 · 19 comments · 85% of case engagement
09-24 06:48amplified on r/LocalLLaMAreddit.post.1wougsk
Icy-Stay-1004
peak 1 · 2 comments · 7% of case engagement
10-03 06:20amplified on hacker newshn.story.49941792
MehrdadKhnzd
peak 1 · 0 comments · 4% of case engagement
09-22 02:20our radar first saw it · +108.3hdiscovery anchor: hn.story.49795948—
pace: p60 vs 1032 stories at the 336h mark (now 578h old) — ahead of code-world-model-coding-agents (1.0x), behind intigriti-support-agent-authentication-failures (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnContext compaction, measured: FutureOS vs. Codex vs. OpenCode
Retrieved article excerpt

Open article · Retrieved 2026-09-22T02:22:27.484928+00:00

# Context compaction, measured: FutureOS vs Codex vs OpenCode

2026-08-04
FutureOS

[agent](https://future-os-blog.github.io/tags/agent.html)[compaction](https://future-os-blog.github.io/tags/compaction.html)[llm](https://future-os-blog.github.io/tags/llm.html)[experiments](https://future-os-blog.github.io/tags/experiments.html)

A long agent session fills the context window. Something has to go. Every compaction strategy answers the same question differently: what do you keep?

We tested three answers — FutureOS's default, OpenCode, and Codex — on the same model, the same questions, the same call shape. The only variable was what a compaction keeps. Everything below is reproducible from [`compaction_experiment/`](https://github.com/futuregene/future-os/tree/main/scripts/compaction_experiment) in the FutureOS repo. External implementations are pinned to Codex `b13164d8` and OpenCode `e03db9bc` (package 1.18.31); a later commit invalidates them.

## The experiment

Grow a session until the context is full. Force a compaction. Ask 178 questions that can only be answered from memory — line three of some script's output, an ID a tool returned, a constraint the user laid down hours earlier. Eight questions are decoys whose answer never appeared anywhere, there to catch a system that guesses.

The 178 questions, and how many each system could still answer

*Figure 1: the 178 questions, and how many each system could still answer after compaction.*

FutureOS retained 147 (83%). OpenCode retained 83 (47%). Codex retained 68 (38%). None fell for a decoy.

The gap isn't about who wrote a better summary. The three systems mean different things by "compact":

|  | What it thinks should be kept |
| --- | --- |
| **Codex** | What the user said. All user messages (capped at 20,000 tokens) plus a whole-history summary; assistant prose and tool output are dropped outright. |
| **OpenCode** | A summary plus a recent tail (capped at 15,000 tokens). Detail is expected to survive inside the summary. |
| **FutureOS** | The originals. Protected user and assistant originals come first; the summary and the tool-evidence index are compensation for what doesn't fit. |

A summary is lossy and you can't verify what it dropped. Codex goes further — it never meant to keep the agent's own output at all. That choice, more than any difference in summary quality, is where the score comes from.

## What fills the context window

Compaction is a budget problem. It helps to know what filled the window. We tallied five frozen real-session chains:

What actually fills a context window

*Figure 2: left — what fills the window; right — where the follow-up questions point.*

| Record type | Records | Share of records | Characters | **Share of characters** |
| --- | --- | --- | --- | --- |
| **Tool output** | 4,370 | 45.2% | 6,622,624 | **96.1%** |
| Assistant text | 784 | 8.1% | 256,060 | 3.7% |
| User text | 141 | 1.5% | 10,927 | 0.2% |
| Tool calls (arguments) | 4,370 | 45.2% | — | — |

9,665 records, 6,889,611 characters. Per chain, tool output is 99.4% / 98.4% / 96.8% / 96.7% / 93.2% of the characters — never below 93%.

The window fills with tool output, not conversation. Everything you and the model said to each other is 3.9%. But when someone asks "what happened earlier", 80% of those follow-ups point at what the assistant itself said, and only 10% each at user turns and tool output:

|  | Share of the bulk | Share of the questions |
| --- | --- | --- |
| Tool output | **96.1%** | 10% |
| Assistant + user text | 3.9% | **90%** |

That mismatch is the whole design in one table. Tool output is almost everything by volume and almost nothing by reference, so it can be compressed down to signposts. The prose is tiny by volume but is where 90% of the questions point, so compressing it saves nearly nothing and costs exactly the part people ask about.

Read that way, the losses above make sense. Codex and OpenCode compress prose and tools together. Codex keeps no assistant text at all, giving up 80% of the questions by construction; OpenCode folds prose into a summary, which is lossy. Most of the 83% vs 38% / 47% gap comes from this one decision.

## What each approach is good at

three-way comparison

*Figure 3: left — how much information is retained; middle — how many tokens that costs; right — the money behind each point of recall.*

| Strategy | Recall | Median projection | Compression | Compaction (cold / cached) | Per turn |
| --- | --- | --- | --- | --- | --- |
| **`summarized`** (ours, default) | **147/178 (83%)** | 12,113 tok | 5.7% | 7.98 / **0.53** | 0.000480 |
| `deterministic` (ours, no model call) | 127/178 (71%) | 9,945 tok | 4.7% | **0 / 0** | 0.000394 |
| **`opencode`** | 83/178 (47%) | 5,152 tok | 2.1% | 0.99 / 0.99 | 0.000204 |
| **`codex`** | 68/178 (38%) | 1,706 tok | 0.6% | 7.64 / **0.42** | 0.000068 |
| *(no compaction)* | — | 232,777 tok | — | — | 0.009219 |

Costs are CNY; `Per turn` is what re-sending the projection costs on every later turn.

compression vs recall

*Figure 4: x is projection size (log), y is recall. Compression and quality are negatively correlated here.*

Codex is the cheapest. Smallest projection, cheapest turn, fastest break-even (46 turns). Its compaction request shares the prefix, so it compacts for ¥0.42. The price is retention scope — no assistant prose, no tool output.

OpenCode is the balanced one. Summary plus tail keeps a little of all three kinds (47% at 5,152 tokens), and a dedicated compaction system prompt keeps the compaction logic out of the session prompt. It gives up prefix sharing (a flat ¥0.99) and the summary is still lossy.

FutureOS keeps the most. Originals-first earns 83%, and because the summary request reuses the session's own system prompt and tool definitions it still hits the prefix cache (99.8% in production), so ¥7.98 cold becomes ¥0.53. The `deterministic` tier makes no model call at all — zero cost, 71% recall. The cost is the largest projection and the most expensive turn: we spend tokens for recall.

cost, cold vs cache-served

*Figure 5: compaction cost, cold vs cache-served.*

cost per point of recall

*Figure 6: one compaction plus 100 turns, per point of recall.*

The cache cuts the price by an order of magnitude: ¥7.98 cold against ¥0.53 cached — 15× — which turns "8× more expensive than OpenCode" into "half the price of OpenCode". Per point of recall the arms cost 0.0070 / 0.0112 / 0.0217 (FutureOS / Codex / OpenCode). A cheap compaction isn't cheap if it drops what later turns need. And compaction only pays for itself over tens of turns (61 for `summarized`, 46 for `codex`, 110 for `opencode`); for a short session it's a net cost, buying recall and headroom.

> ⚠️ The cache figures are modelled, not measured: the runs' own cache counters are contaminated by arm and run ordering, so the 98% is anchored to the one production measurement (99.8%).

## How our compaction works

one complete projection

*Figure 7: one complete projection. Original messages are never deleted or rewritten; compaction only changes what the next request sees.*

The keyword is projection. Compaction doesn't delete the journal or rewrite history; it recomputes, for the next request only, what the model gets to see. The originals stay on disk — searchable, exportable, forkable.

There are two algorithms and one fallback. `algorithm_version` writes exactly two values: `deterministic-evidence-v1` (protected originals + a recent tail + a deterministic tool-evidence index, no model call) and `summarized-evidence-v1` (the same plus a handoff summary, one model call). `summarized` is the default; `deterministic` is also the fallback — with no provider reachable, or when the summary fails, the deterministic projection is committed.

The trigger runs before every model step:

```
economic_trigger  = floor(W × 0.8)                      # 1M window -> 800,000
effective_trigger = min(economic_trigger, W − O − margin)
margin            = min(2048, W / 16)
```

`O` is the model's declared output ceiling; the `min` means a model that reserves a large output is bounded by its own limit, not a fixed number. If the input still fits, nothing is cut.

trigger and budget

*Figure 8: top — how the trigger is derived; bottom — the target budgets inside the projection.*

What's in the projection: user text is always kept, and the remaining room prefers assistant originals. Tool output is compressed to a 2,048-token evidence index, a recent tail of ≈8K tokens is kept, and the history target is about 32K (up to 128K, never past real capacity). A demoted assistant output is labelled "omitted / not summarised" — never dressed up as a summary — and its original stays queryable.

Why does user text win? User constraints and goals can't be regenerated; tool output and assistant prose usually can, from the originals. It's an information-recoverability argument, not a fairness one.

The evidence index is fully deterministic — no model call. Error results first; then grouped by tool and target, prioritising config / schema / validation / test targets; the latest and first of each group; remaining space by recency. Each selected record becomes one bounded JSON row (entryId, blockIndex, sourceOrder, tool/target, error flag, head/tail excerpts). A JSON row is never cut in half, and the index is priority-ordered rather than a timeline — `sourceOrder` keeps the chronology visible so an older error can't be mistaken for a current one.

The summary request has no separate prompt, and that's deliberate. The provider's prefix cache compares from token 0: `[system prompt][tool definitions][messages…]`. As long as those three segments match the conversation that drove the request, the whole prefix is served from cache — so the summary request carries the session's own system prompt and tool definitions. Measured on a primed prefix: identical shape 93.7% cache hit; system prompt substituted 0%; tool definitions dropped 0%; one line added 0%. On the real path, a session grown to 212,911 tokens compacted with `cache_read = 212,548` — 99.8% from cache, ¥0.003 against ¥0.53 cold. (Cache read is 50× cheaper than fresh input: 0.02 vs 1.0 per 1M tokens.)

Checkpoints are idempotent. A compaction commits atomically with its completion receipt; a checkpoint written by a retired algorithm isn't a checkpoint at all and gets re-covered once; a same-key success is reused without another request; an unresolved `started` operation keeps its concurrency fence — no claim-stealing, no fabricated success.

## Can retrieval recover the loss?

If the originals are all still there, why not let the model look things up? We ran the open-book exam: one arm, `summarized`, 18 cases, 178 values.

| # | What changed | Design | Recall | Tool calls |
| --- | --- | --- | --- | --- |
| — | *(closed baseline)* | no tools | 147/178 (82.6%) | 0 |
| 1 | production prompt + tools | autonomous | 148/178 (83.1%) | **0** |
| 2 | + stronger recall guidance | autonomous | 149/178 (83.7%) | **0** |
| 3 | + wording that stops calling the projection "the record" | autonomous | 145/178 (81.5%) | **0** |
| 4 | closed re-scored under round 3's wording (control) | no tools | 147/178 (82.6%) | 0 |
| 5 | + explicit verification instruction, first call forced | required | **166/178 (93.3%)** | **100** |

open book

*Figure 9: where the 178 values went. The orange 27 is constant across every round where searching is only allowed — the model simply never looked.*

Rounds 1–3 are the same result three times: zero tool calls, and the 145–149 spread is one-flip noise on 178 values. Round 4 is the control for round 3 — if the wording change moved the score, the closed arm would move with it; it didn't (147 both ways). The 27 is constant: those values are in the archive and the model never looked. The required round recovered 20 of the 28 absent ones by searching — so the gap is behavioural, not capability. Round 5 is an upper bound, not a result: no such in
thetechlead10
🟧 echo.github ⭐The primary artifact is FutureOS's own experiment commit and its accompanying closed/open-book compaction reports. The report states: “All 2Hui Chen——
🟠 redditI built a cache-friendly context compacting plugin for OpenCode
LocalLLaMA
schennardo1719
🟠 redditContext compaction, measured: FutureOS vs Codex vs OpenCode
LocalLLaMA
Icy-Stay-100402
🟧 hnCodex OriginalsMehrdadKhnzd10

Interpretation history

Decision trace