Shunt is a Claude Code plugin published by Spotify Engineering as part of its portal-ai-plugins suite: a ~33-line PreToolUse hook blocks bulk reads of any file over 350 lines and delegates that context ingestion to a lightweight worker model (e.g. Gemini 2.5 Flash) via Spotify's Portal gateway, claiming 82-94% savings on bulk reads in Spotify's Java monorepo. Third-party teardowns (Wavect, YouTube analyses) show the headline denominator is Claude's context intake — estimated by characters/4 and excluding worker-model tokens, Portal costs and cached-read pricing — with a worked example where Claude's reads shrink 90% while total cross-model tokens rise from 40k to 48k; the adjacent RTK tool was independently tested by JetBrains reporting 96.2M tokens 'saved' while median session cost rose 7.6%. The supplied web material uniformly attributes the plugin to Spotify's portal-ai-plugins, which conflicts with the case's echo-derived repo attribution to 'sorantis' (uncorroborated), and a similarly named inference-routing proxy (pleaseai/shunt) is a distinct project. Shunt sits in a widening category of Claude Code token-saving harnesses (context-mode, code knowledge graphs, Superpowers, RTK/Headroom) where measured whole-session effects run far below headline claims.
Known territory: Scott's Mature Token Law canon already holds that tokens saved are not the score without task outcomes and that usage-based claims must be audited at the whole-session unit, so a month of new category entrants repeats his position rather than arriving at it from a consequential party. The one usable nugget is Extra Headroom's production denominator (~-33% across 183 users) as a citable reality anchor against 82–94% headline claims — but it is vendor-reported and lands exactly where the radar's own RTK cost-regression and GitHub over-compression cases already scrutinize, so it confirms rather than changes what Scott argues or builds.
ip:framework.the-mature-token-lawip:source.the-mature-token-law-ebookip:concept.ai-unit-economicsradar:github-tool-output-cost-tradeoffradar:rtk-coding-agent-cost-regressionradar:multi-model-orchestrator-worker-agentsradar:tura-token-efficient-agentradar:halv-coding-token-savings
queries asked of Scott's wikis
- mature token law savings without task outcomes
- deterministic hook guardrails vs advisory prompt routing
- worker model delegation offloading whole-session cost
- priced benchmark token-saving coding-agent tools
- targeted retrieval vs bulk context ingestion agent
- Claude Code hooks plugin harness engineering
now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 818h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-10-06T17:32:46Z
grounded: known/low — Known territory: Scott's Mature Token Law canon already holds that tokens saved are not the score without task outcomes and that usage-based claims must be audi
2026-10-06T17:24:44Z
Extra Headroom's production figures (input -34%, output -33% across 183 users) turn this from a lone unverified Shunt claim into a forming category of Claude Code cost harnesses: corroboration now attaches to the category (Spotify Portal's implementation, JetBrains' independent RTK test, Extra Headroom's deployment) while Shunt's 82-94% itself remains unevidenced and looks inflated against the only production denominator. The magnitude-valve spread reflects a month of accumulated breadth across three platforms, not current velocity (0.33 pts/h, latest entrant at 2 points/0 comments), so heat stays low.
2026-10-06T16:42:15Z
evidence attached: hn.story.49978145 — A second vendor (Extra Headroom) with production measurements across 183 users enters the same token-savings episode, showing category formation — company-reported, not independent corroboration.
2026-09-11T12:34:58Z
The RTK test and Ory Lumen attachments broaden the adjacent efficiency-tool landscape but supply neither results nor a demonstrated connection to Shunt. Prior routing commentary overstates the supplied RTK/Headroom excerpt: it reports a test setup and correct answers, not the cost outcome, so it cannot substantiate negative savings or disprove Shunt.
2026-09-11T12:22:45Z
evidence attached: hn.story.49656845 — This is another concrete attempt to reduce Claude Code latency and token cost through semantic task handling.
2026-09-11T11:23:10Z
evidence attached: hn.story.49656275 — JetBrains' independent test of another Claude Code skill claiming 60-90% token savings is independent evidence bearing on whether token-savings skills deliver material reductions.
2026-09-11T01:29:24Z
The refreshed Portal comments are repetitive discussion of model delegation, not a new implementation result or evidence linking Portal to Shunt. Shunt’s claimed savings remain unverified, with no demonstrated quality-preserving reduction in whole-task cost.
2026-09-10T03:26:52Z
The refreshed Portal discussion adds familiar model-delegation commentary, not evidence connecting Portal to Shunt or validating Shunt’s savings. Shunt remains an unverified efficiency lead; neither the adjacent release nor testing of other tools establishes quality-preserving whole-task savings for it.
2026-09-09T14:32:26Z
The Spotify Portal post adds a concrete model-offloading evaluation lead, but the supplied excerpt does not establish its relationship to Shunt despite earlier attachment notes treating that link as settled. RTK/Headroom testing and the new comments motivate checking all-model task cost and workflow tradeoffs; neither independently validates or disproves Shunt’s advertised savings.
2026-09-09T14:24:00Z
evidence attached: reddit.post.1wbmcgw — The post points to the released Spotify plugins underlying the open case's large token-saving claim, though the performance claim remains vendor-reported.
2026-09-09T14:24:00Z
evidence attached: reddit.post.1wbmtw2 — Independent priced testing materially qualifies the token-savings claim, finding modest RTK gains but higher average cost when the tools are combined.
2026-09-07T14:35:43Z
The GitHub echo repeats the HN pitch rather than independently establishing an implementation or validating savings. Shunt remains an unverified efficiency lead, with no evidence distinguishing selected-output token reductions from quality-preserving whole-task savings.
2026-09-07T14:28:58Z
grounded: known/low — Shunt’s token-saving pitch repeats territory Scott already covers in The Mature Token Law—tokens saved are not the score without task outcomes—and targets Claud
2026-09-07T14:26:48Z
case created — The linked implementation and quantified savings claim form a bounded engineering episode, although workload coverage and quality effects are not established.