Retrieved article excerpt
Open article · Retrieved 2026-09-22T02:22:27.484928+00:00
# Context compaction, measured: FutureOS vs Codex vs OpenCode
2026-08-04
FutureOS
[agent](https://future-os-blog.github.io/tags/agent.html)[compaction](https://future-os-blog.github.io/tags/compaction.html)[llm](https://future-os-blog.github.io/tags/llm.html)[experiments](https://future-os-blog.github.io/tags/experiments.html)
A long agent session fills the context window. Something has to go. Every compaction strategy answers the same question differently: what do you keep?
We tested three answers — FutureOS's default, OpenCode, and Codex — on the same model, the same questions, the same call shape. The only variable was what a compaction keeps. Everything below is reproducible from [`compaction_experiment/`](https://github.com/futuregene/future-os/tree/main/scripts/compaction_experiment) in the FutureOS repo. External implementations are pinned to Codex `b13164d8` and OpenCode `e03db9bc` (package 1.18.31); a later commit invalidates them.
## The experiment
Grow a session until the context is full. Force a compaction. Ask 178 questions that can only be answered from memory — line three of some script's output, an ID a tool returned, a constraint the user laid down hours earlier. Eight questions are decoys whose answer never appeared anywhere, there to catch a system that guesses.
The 178 questions, and how many each system could still answer
*Figure 1: the 178 questions, and how many each system could still answer after compaction.*
FutureOS retained 147 (83%). OpenCode retained 83 (47%). Codex retained 68 (38%). None fell for a decoy.
The gap isn't about who wrote a better summary. The three systems mean different things by "compact":
| | What it thinks should be kept |
| --- | --- |
| **Codex** | What the user said. All user messages (capped at 20,000 tokens) plus a whole-history summary; assistant prose and tool output are dropped outright. |
| **OpenCode** | A summary plus a recent tail (capped at 15,000 tokens). Detail is expected to survive inside the summary. |
| **FutureOS** | The originals. Protected user and assistant originals come first; the summary and the tool-evidence index are compensation for what doesn't fit. |
A summary is lossy and you can't verify what it dropped. Codex goes further — it never meant to keep the agent's own output at all. That choice, more than any difference in summary quality, is where the score comes from.
## What fills the context window
Compaction is a budget problem. It helps to know what filled the window. We tallied five frozen real-session chains:
What actually fills a context window
*Figure 2: left — what fills the window; right — where the follow-up questions point.*
| Record type | Records | Share of records | Characters | **Share of characters** |
| --- | --- | --- | --- | --- |
| **Tool output** | 4,370 | 45.2% | 6,622,624 | **96.1%** |
| Assistant text | 784 | 8.1% | 256,060 | 3.7% |
| User text | 141 | 1.5% | 10,927 | 0.2% |
| Tool calls (arguments) | 4,370 | 45.2% | — | — |
9,665 records, 6,889,611 characters. Per chain, tool output is 99.4% / 98.4% / 96.8% / 96.7% / 93.2% of the characters — never below 93%.
The window fills with tool output, not conversation. Everything you and the model said to each other is 3.9%. But when someone asks "what happened earlier", 80% of those follow-ups point at what the assistant itself said, and only 10% each at user turns and tool output:
| | Share of the bulk | Share of the questions |
| --- | --- | --- |
| Tool output | **96.1%** | 10% |
| Assistant + user text | 3.9% | **90%** |
That mismatch is the whole design in one table. Tool output is almost everything by volume and almost nothing by reference, so it can be compressed down to signposts. The prose is tiny by volume but is where 90% of the questions point, so compressing it saves nearly nothing and costs exactly the part people ask about.
Read that way, the losses above make sense. Codex and OpenCode compress prose and tools together. Codex keeps no assistant text at all, giving up 80% of the questions by construction; OpenCode folds prose into a summary, which is lossy. Most of the 83% vs 38% / 47% gap comes from this one decision.
## What each approach is good at
three-way comparison
*Figure 3: left — how much information is retained; middle — how many tokens that costs; right — the money behind each point of recall.*
| Strategy | Recall | Median projection | Compression | Compaction (cold / cached) | Per turn |
| --- | --- | --- | --- | --- | --- |
| **`summarized`** (ours, default) | **147/178 (83%)** | 12,113 tok | 5.7% | 7.98 / **0.53** | 0.000480 |
| `deterministic` (ours, no model call) | 127/178 (71%) | 9,945 tok | 4.7% | **0 / 0** | 0.000394 |
| **`opencode`** | 83/178 (47%) | 5,152 tok | 2.1% | 0.99 / 0.99 | 0.000204 |
| **`codex`** | 68/178 (38%) | 1,706 tok | 0.6% | 7.64 / **0.42** | 0.000068 |
| *(no compaction)* | — | 232,777 tok | — | — | 0.009219 |
Costs are CNY; `Per turn` is what re-sending the projection costs on every later turn.
compression vs recall
*Figure 4: x is projection size (log), y is recall. Compression and quality are negatively correlated here.*
Codex is the cheapest. Smallest projection, cheapest turn, fastest break-even (46 turns). Its compaction request shares the prefix, so it compacts for ¥0.42. The price is retention scope — no assistant prose, no tool output.
OpenCode is the balanced one. Summary plus tail keeps a little of all three kinds (47% at 5,152 tokens), and a dedicated compaction system prompt keeps the compaction logic out of the session prompt. It gives up prefix sharing (a flat ¥0.99) and the summary is still lossy.
FutureOS keeps the most. Originals-first earns 83%, and because the summary request reuses the session's own system prompt and tool definitions it still hits the prefix cache (99.8% in production), so ¥7.98 cold becomes ¥0.53. The `deterministic` tier makes no model call at all — zero cost, 71% recall. The cost is the largest projection and the most expensive turn: we spend tokens for recall.
cost, cold vs cache-served
*Figure 5: compaction cost, cold vs cache-served.*
cost per point of recall
*Figure 6: one compaction plus 100 turns, per point of recall.*
The cache cuts the price by an order of magnitude: ¥7.98 cold against ¥0.53 cached — 15× — which turns "8× more expensive than OpenCode" into "half the price of OpenCode". Per point of recall the arms cost 0.0070 / 0.0112 / 0.0217 (FutureOS / Codex / OpenCode). A cheap compaction isn't cheap if it drops what later turns need. And compaction only pays for itself over tens of turns (61 for `summarized`, 46 for `codex`, 110 for `opencode`); for a short session it's a net cost, buying recall and headroom.
> ⚠️ The cache figures are modelled, not measured: the runs' own cache counters are contaminated by arm and run ordering, so the 98% is anchored to the one production measurement (99.8%).
## How our compaction works
one complete projection
*Figure 7: one complete projection. Original messages are never deleted or rewritten; compaction only changes what the next request sees.*
The keyword is projection. Compaction doesn't delete the journal or rewrite history; it recomputes, for the next request only, what the model gets to see. The originals stay on disk — searchable, exportable, forkable.
There are two algorithms and one fallback. `algorithm_version` writes exactly two values: `deterministic-evidence-v1` (protected originals + a recent tail + a deterministic tool-evidence index, no model call) and `summarized-evidence-v1` (the same plus a handoff summary, one model call). `summarized` is the default; `deterministic` is also the fallback — with no provider reachable, or when the summary fails, the deterministic projection is committed.
The trigger runs before every model step:
```
economic_trigger = floor(W × 0.8) # 1M window -> 800,000
effective_trigger = min(economic_trigger, W − O − margin)
margin = min(2048, W / 16)
```
`O` is the model's declared output ceiling; the `min` means a model that reserves a large output is bounded by its own limit, not a fixed number. If the input still fits, nothing is cut.
trigger and budget
*Figure 8: top — how the trigger is derived; bottom — the target budgets inside the projection.*
What's in the projection: user text is always kept, and the remaining room prefers assistant originals. Tool output is compressed to a 2,048-token evidence index, a recent tail of ≈8K tokens is kept, and the history target is about 32K (up to 128K, never past real capacity). A demoted assistant output is labelled "omitted / not summarised" — never dressed up as a summary — and its original stays queryable.
Why does user text win? User constraints and goals can't be regenerated; tool output and assistant prose usually can, from the originals. It's an information-recoverability argument, not a fairness one.
The evidence index is fully deterministic — no model call. Error results first; then grouped by tool and target, prioritising config / schema / validation / test targets; the latest and first of each group; remaining space by recency. Each selected record becomes one bounded JSON row (entryId, blockIndex, sourceOrder, tool/target, error flag, head/tail excerpts). A JSON row is never cut in half, and the index is priority-ordered rather than a timeline — `sourceOrder` keeps the chronology visible so an older error can't be mistaken for a current one.
The summary request has no separate prompt, and that's deliberate. The provider's prefix cache compares from token 0: `[system prompt][tool definitions][messages…]`. As long as those three segments match the conversation that drove the request, the whole prefix is served from cache — so the summary request carries the session's own system prompt and tool definitions. Measured on a primed prefix: identical shape 93.7% cache hit; system prompt substituted 0%; tool definitions dropped 0%; one line added 0%. On the real path, a session grown to 212,911 tokens compacted with `cache_read = 212,548` — 99.8% from cache, ¥0.003 against ¥0.53 cold. (Cache read is 50× cheaper than fresh input: 0.02 vs 1.0 per 1M tokens.)
Checkpoints are idempotent. A compaction commits atomically with its completion receipt; a checkpoint written by a retired algorithm isn't a checkpoint at all and gets re-covered once; a same-key success is reused without another request; an unresolved `started` operation keeps its concurrency fence — no claim-stealing, no fabricated success.
## Can retrieval recover the loss?
If the originals are all still there, why not let the model look things up? We ran the open-book exam: one arm, `summarized`, 18 cases, 178 values.
| # | What changed | Design | Recall | Tool calls |
| --- | --- | --- | --- | --- |
| — | *(closed baseline)* | no tools | 147/178 (82.6%) | 0 |
| 1 | production prompt + tools | autonomous | 148/178 (83.1%) | **0** |
| 2 | + stronger recall guidance | autonomous | 149/178 (83.7%) | **0** |
| 3 | + wording that stops calling the projection "the record" | autonomous | 145/178 (81.5%) | **0** |
| 4 | closed re-scored under round 3's wording (control) | no tools | 147/178 (82.6%) | 0 |
| 5 | + explicit verification instruction, first call forced | required | **166/178 (93.3%)** | **100** |
open book
*Figure 9: where the 178 values went. The orange 27 is constant across every round where searching is only allowed — the model simply never looked.*
Rounds 1–3 are the same result three times: zero tool calls, and the 145–149 spread is one-flip noise on 178 values. Round 4 is the control for round 3 — if the wording change moved the score, the closed arm would move with it; it didn't (147 both ways). The 27 is constant: those values are in the archive and the model never looked. The required round recovered 20 of the 28 absent ones by searching — so the gap is behavioural, not capability. Round 5 is an upper bound, not a result: no such in