Yolanda and Spencer claim a fine-tuned Qwen proxy that trims Codex tool-call output cut their tokens 29.6% without breaking the prompt cache, and adoption as a standard coding-agent cost lever β or replication showing trajectory damage β resolves whether trajectory-preserving tool-output compression becomes routine.
state: seedheat: lowuncertainty: mediumconvergesscott: highcoding-agent-cost tool-output-compression agent-harnesses
What is this?
A Show HN post by Yolanda and Spencer presents a CLI that runs as a proxy in front of Codex-style coding agents, using a fine-tuned Qwen model to trim tool-call output before it re-enters the context window, with a claimed 29.6% token reduction that preserves the provider's prompt cache. The supplied web snippets do not surface the post itself or the 29.6% figure β the specific claims rest on the case's own evidence and are not independently corroborated here. What the snippets do establish is the surrounding territory: Qwen models are in wide use as cheap local pre-processors in coding workflows (one snippet shows a 'spencer_kw' describing a local-Qwen code-review step that catches ~60% of mistakes and saves ~$80/month in API costs β plausibly the same Spencer, but not established), and coding-agent inference costs are under visible price pressure (OpenAI's 80% input-price cuts on GPT-5.6 Luna/Terra).
Why it matters to Scott
Converges with his answers-not-content / signal-extraction doctrine β a fine-tuned local Qwen compressing bulky tool output outside the agent's attention before it re-enters context β and its cache-preserving insertion-time trim directly stresses his prefix-caching-economics claim that full transcript retention is the cost-optimal shape (compress at first write, never rewrite). Immediately actionable on his own stack: a learned, cache-aware upgrade to ask's deliberately lossy --compact truncation, droppable into the LiteLLM proxy path his agents already route through β and the open trajectory-damage question bears on the scoutβsenior zero-loss rule, with GitHub's over-compression counter-claim as the live counter-hypothesis the adoption-or-replication test would settle.
dev:concept.answers-not-contentip:concept.signal-extractionip:concept.prefix-caching-economicsip:framework.scout-senior-splitdev:project.askdev:technology.litellmradar:tokencompress-agent-context-pruningradar:github-tool-output-cost-tradeoffradar:distil-decision-equivalent-context-compressionradar:concept.context-managementradar:concept.prompt-cachingradar:concept.inference-economicsradar:concept.small-language-models
queries asked of Scott's wikis
- prompt cache invalidation agent harness
- tool output compression context window management
- local model proxy preprocessing cost savings
- fine-tuned small model narrow task in pipeline
- coding agent token cost economics
- context compaction agent trajectory memory
Measured heat
now 0 pts/hpeak 8 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 314h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p50 vs 1188 stories at the 168h mark (now 314h old) β ahead of android-editable-graph-agent (1.1x), behind ai-vuln-reports-oss-disclosure (0.9x)
Evidence (2) β β canonical anchor
Interpretation history
2026-09-30T19:44:59Z
origin walked (opencode/cheap-glm, conf 0.9): anchor hn.story.49911910 -> echo.github.cb151dfa2a by Everest (everestagi.com) β Spencer McKe & Yolanda, published under GitHub user spenmcke
2026-09-30T19:32:07Z
grounded: converges/high β Converges with his answers-not-content / signal-extraction doctrine β a fine-tuned local Qwen compressing bulky tool output outside the agent's attention before
2026-09-30T19:25:30Z
case created β Quantified, mechanism-specific inference-economics artifact (cache-preserving compression) with no existing case covering the approach.
Decision trace
- 10-04 09:54review_screenjev screen: no material development (noul=0.09)
- 10-02 20:58review_screenjev screen: no material development (noul=0.15)
- 10-01 05:44promote_anchororigin walk conf 0.9
- 10-01 05:32groundConverges with his answers-not-content / signal-extraction doctrine β a fine-tuned local Qwen compressing bulky tool output outside the agent's attention before it re-enters context β and its cac
- 10-01 05:25createQuantified, mechanism-specific inference-economics artifact (cache-preserving compression) with no existing case covering the approach.