Anthropic's Opus 5.5 repricing cut cache-read rates 60% (to $0.20/M) versus 20% for input/output tokens, materially changing the economics of cache-heavy long-context agent workloads.
state: resolvedheat: lowuncertainty: lowconvergesscott: highinference-economics prompt-caching anthropicAnthropic
What is this?
Anthropic launched Claude Opus 5.5 on September 22, 2026, cutting input/output prices 20% ($5โ$4 and $25โ$20 per million tokens) while slashing cache reads 60% ($0.50โ$0.20) โ a line item Anthropic itself frames as 'the majority of costs in agentic and coding workloads' โ while cache writes fell only 20% ($6.25โ$5); developer comments in the coverage report long-context interactive sessions running 95โ98% cache reads. The model carries a 1M-token context window, is the default Opus in Claude Code 2.1.280, and ships via API, Bedrock, Vertex, and Foundry. Anthropic claims ~40% lower cost 'at default settings' on typical workloads, leaning on the model finishing jobs with fewer tokens; daily.dev instead measured ~120k tokens per task at max reasoning vs ~73k for Opus 5 (~27k for GPT-6 Astra), calling it among the least token-efficient models tested โ the two claims measure different effort regimes, so realized savings look workload-dependent. Pricing is corroborated across five-plus independent sources; the token-efficiency picture is the contested part.
Why it matters to Scott
Anthropic now prices cache reads as 'the majority of costs in agentic and coding workloads' โ a dated receipt for the cache-read-dominated agent economics behind his Agent Token Manifesto and cost-tiered routing, and the 60% cut directly reprices the opus tier on his LiteLLM gateway and the numbers in his published LLM pricing guide. It also rewrites the write/read tradeoff his caching radar cases (cache warming, TTL analyzer) compute against, while the contested '~40% cheaper' claim versus daily.dev's least-token-efficient measurement is a live instance of the per-task-cost-vs-published-price divergence his budget-calculator canon already argues.
dev:project.llmreportdev:concept.cost-tiered-llm-routingdev:technology.litellmradar:concept.inference-economicsradar:concept.prompt-cachingradar:concept.model-routingradar:concept.token-efficiencyradar:concept.open-weight-modelsradar:gpt6-prompt-cache-controlsradar:openai-gpt56-sol-api-price-cutradar:cache-tax-idle-session-warmingradar:claude-code-cache-ttl-analyzerradar:anthropic-context-compaction-cost-reversal
queries asked of Scott's wikis
- per-task cost math token efficiency vs per-token price
- prompt caching strategy cache-read dominated context design
- agent harness budget calculator pricing assumptions
- agent-maintained wiki long context maintenance cost
- open-weights local inference cost case vs frontier api pricing
- frontier release triage model selection criteria
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
Evidence (4) โ โญ canonical anchor
Interpretation history
2026-09-26T14:42:59Z
The Artificial Analysis attachment rounds out the quality side โ Opus 5.5 tops the Intelligence Index (58) while remaining token-hungry at max effort โ but adds nothing the repricing claim still needed. With pricing corroborated across 5+ sources, the agent-economics impact independently demonstrated by the 203-PR measurement (68% cache reads; ~half Sonnet 5's cost per line), and attention collapsed to 0.17 pts/h at day 5 on 2 platforms with no expanding periphery, the episode is established and closes: durable facts pass to Scott's gateway/guide upkeep and the sibling caching-strategy cases.
2026-09-26T14:25:52Z
evidence attached: reddit.post.1wqrmpa โ Independent Artificial Analysis data (via heise) on Opus 5.5 quality and per-task cost vs GPT-6 Astra ($5.98 vs $3.26) directly contextualizes Opus 5.5's agent-workload economics in the repricing case.
2026-09-26T06:42:31Z
The 203-PR independent measurement (68% of agent spend was cache reads at $0.20/M; Opus 5.5 โ half Sonnet 5's cost per line of code) upgrades the case from a corroborated price-list observation to a demonstrated cost-structure shift for coding agents โ two independent lines (Anthropic's own framing plus real-workload measurements) now support the hypothesis, even as engagement cools well past its peak (0.33 pts/h vs 17 peak).
2026-09-26T06:23:00Z
evidence attached: reddit.post.1wqis66 โ Independent 203-PR cost measurement showing 68% cached reads at $0.20/M, real-workload corroboration that the repricing flips per-line economics toward Opus.
2026-09-24T17:28:22Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wp5kec -> echo.blog.0c90c73c08 by Anthropic
2026-09-24T17:11:11Z
grounded: converges/high โ Anthropic now prices cache reads as 'the majority of costs in agentic and coding workloads' โ a dated receipt for the cache-read-dominated agent economics behin
2026-09-24T17:02:23Z
case created โ Concrete flagship repricing with a disproportionate cache-read cut that lands directly on Scott's agent cost math, and no open case tracks Anthropic API pricing.
Decision trace
- 09-27 00:43resolveThe Artificial Analysis attachment rounds out the quality side โ Opus 5.5 tops the Intelligence Index (58) while remaining token-hungry at max effort โ but adds nothing the repricing claim still neede
- 09-27 00:25attachIndependent Artificial Analysis data (via heise) on Opus 5.5 quality and per-task cost vs GPT-6 Astra ($5.98 vs $3.26) directly contextualizes Opus 5.5's agent-workload economics in the repricing
- 09-27 00:23propose_attachIndependent Artificial Analysis data (via heise) on Opus 5.5 quality and per-task cost vs GPT-6 Astra ($5.98 vs $3.26) directly contextualizes Opus 5.5's agent-workload economics in the repricing
- 09-26 16:42repriceThe 203-PR independent measurement (68% of agent spend was cache reads at $0.20/M; Opus 5.5 โ half Sonnet 5's cost per line of code) upgrades the case from a corroborated price-list observation t
- 09-26 16:23attachIndependent 203-PR cost measurement showing 68% cached reads at $0.20/M, real-workload corroboration that the repricing flips per-line economics toward Opus.
- 09-26 16:22propose_attachIndependent 203-PR cost measurement showing 68% cached reads at $0.20/M, real-workload corroboration that the repricing flips per-line economics toward Opus.
- 09-26 09:51review_screenjev screen: no material development (noul=0.13)
- 09-25 04:21sensor_dirtycomment_update
- 09-25 03:28promote_anchororigin walk conf 0.85
- 09-25 03:11groundAnthropic now prices cache reads as 'the majority of costs in agentic and coding workloads' โ a dated receipt for the cache-read-dominated agent economics behind his Agent Token Manifesto an
- 09-25 03:02createConcrete flagship repricing with a disproportionate cache-read cut that lands directly on Scott's agent cost math, and no open case tracks Anthropic API pricing.