2026-10-11 18:03 UTC

Independent replication will determine whether Distil’s statistical decision-equivalence gates can reduce coding-agent context usage without materially changing tool calls or SWE-bench task outcomes.

state: expiredheat: lowuncertainty: highnovelscott: nonecontext-compression coding-agents agent-evaluation

What is this?

Distil is presented as a context-compression method for coding agents that uses statistical non-inferiority or decision-equivalence gates to preserve consequential information while reducing context. The supplied evidence titles claim its “surprise-preserving digest” matched or exceeded full-context performance on 500 SWE-bench tasks using the official harness (42.0% versus 39.2%), but the web results do not identify its creators, substantiate those figures, or document an independent replication. The available snippets only establish that SWE-bench measures agentic repair in real repositories and that coding-agent behavior can be evaluated separately from final outcomes.

Why it matters to Scott

No intersection found: the supplied Scott wiki and radar searches returned no hits linking Distil’s compression gates, SWE-bench claims, or replication question to Scott’s existing work or tracked stories.
queries asked of Scott's wikis
  • coding-agent context compression and token economics
  • lossy agent memory with decision-equivalence gates
  • surprise-preserving summaries for tool-using agents
  • non-inferiority evaluation for agent behavior
  • SWE-bench harness replication and statistical power
  • context reduction effects on tool calls and task outcomes

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDistil: context compression for agents, gated by a statistical non-inferiority test — compressed context matched full context on 500 SWE-bench tasks
LocalLLaMA
chandu122107
🟧 echo.github ⭐The originating commit is titled “results(E14): the surprise-preserving digest beats full context — 42.0% vs 39.2% (n=500, official harness)Shekhar Mudarapu——

Interpretation history

Decision trace