2026-10-11 17:09 UTC

Yolanda and Spencer claim a fine-tuned Qwen proxy that trims Codex tool-call output cut their tokens 29.6% without breaking the prompt cache, and adoption as a standard coding-agent cost lever β€” or replication showing trajectory damage β€” resolves whether trajectory-preserving tool-output compression becomes routine.

state: seedheat: lowuncertainty: mediumconvergesscott: highcoding-agent-cost tool-output-compression agent-harnesses

What is this?

A Show HN post by Yolanda and Spencer presents a CLI that runs as a proxy in front of Codex-style coding agents, using a fine-tuned Qwen model to trim tool-call output before it re-enters the context window, with a claimed 29.6% token reduction that preserves the provider's prompt cache. The supplied web snippets do not surface the post itself or the 29.6% figure β€” the specific claims rest on the case's own evidence and are not independently corroborated here. What the snippets do establish is the surrounding territory: Qwen models are in wide use as cheap local pre-processors in coding workflows (one snippet shows a 'spencer_kw' describing a local-Qwen code-review step that catches ~60% of mistakes and saves ~$80/month in API costs β€” plausibly the same Spencer, but not established), and coding-agent inference costs are under visible price pressure (OpenAI's 80% input-price cuts on GPT-5.6 Luna/Terra).

Why it matters to Scott

Converges with his answers-not-content / signal-extraction doctrine β€” a fine-tuned local Qwen compressing bulky tool output outside the agent's attention before it re-enters context β€” and its cache-preserving insertion-time trim directly stresses his prefix-caching-economics claim that full transcript retention is the cost-optimal shape (compress at first write, never rewrite). Immediately actionable on his own stack: a learned, cache-aware upgrade to ask's deliberately lossy --compact truncation, droppable into the LiteLLM proxy path his agents already route through β€” and the open trajectory-damage question bears on the scout–senior zero-loss rule, with GitHub's over-compression counter-claim as the live counter-hypothesis the adoption-or-replication test would settle.
dev:concept.answers-not-contentip:concept.signal-extractionip:concept.prefix-caching-economicsip:framework.scout-senior-splitdev:project.askdev:technology.litellmradar:tokencompress-agent-context-pruningradar:github-tool-output-cost-tradeoffradar:distil-decision-equivalent-context-compressionradar:concept.context-managementradar:concept.prompt-cachingradar:concept.inference-economicsradar:concept.small-language-models
queries asked of Scott's wikis
  • prompt cache invalidation agent harness
  • tool output compression context window management
  • local model proxy preprocessing cost savings
  • fine-tuned small model narrow task in pipeline
  • coding agent token cost economics
  • context compaction agent trajectory memory

Measured heat

now 0 pts/hpeak 8 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 314h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-28 14:00⭐ origin echo-reconstructedThe repo is the primary artifact launched via Show HN. README: "everest reduces Codex tool output before it enters the model context. Everes
Everest (everestagi.com) β€” Spencer McKe & Yolanda, published under GitHub user spenmcke on github (echo) Β· attributed from hn.story.49911910
β€”
09-30 17:30first on hacker news Β· published Β· +51.5hShow HN: Token compression CLI to save Codex/Astra costs
yolandac
β€”
09-30 17:30amplified on hacker news πŸ‘‘hn.story.49911910
yolandac
peak 14 Β· 5 comments Β· 100% of case engagement
09-30 18:21our radar first saw it Β· +52.4hdiscovery anchor: hn.story.49911910β€”
pace: p50 vs 1188 stories at the 168h mark (now 314h old) β€” ahead of android-editable-graph-agent (1.1x), behind ai-vuln-reports-oss-disclosure (0.9x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Token compression CLI to save Codex/Astra costsyolandac145
🟧 echo.github ⭐The repo is the primary artifact launched via Show HN. README: "everest reduces Codex tool output before it enters the model context. EveresEverest (everestagi.com) β€” Spencer McKe & Yolanda, published under GitHub user spenmckeβ€”β€”

Interpretation history

Decision trace