2026-10-11 17:12 UTC

Anthropic's Opus 5.5 repricing cut cache-read rates 60% (to $0.20/M) versus 20% for input/output tokens, materially changing the economics of cache-heavy long-context agent workloads.

state: resolvedheat: lowuncertainty: lowconvergesscott: highinference-economics prompt-caching anthropicAnthropic

What is this?

Anthropic launched Claude Opus 5.5 on September 22, 2026, cutting input/output prices 20% ($5โ†’$4 and $25โ†’$20 per million tokens) while slashing cache reads 60% ($0.50โ†’$0.20) โ€” a line item Anthropic itself frames as 'the majority of costs in agentic and coding workloads' โ€” while cache writes fell only 20% ($6.25โ†’$5); developer comments in the coverage report long-context interactive sessions running 95โ€“98% cache reads. The model carries a 1M-token context window, is the default Opus in Claude Code 2.1.280, and ships via API, Bedrock, Vertex, and Foundry. Anthropic claims ~40% lower cost 'at default settings' on typical workloads, leaning on the model finishing jobs with fewer tokens; daily.dev instead measured ~120k tokens per task at max reasoning vs ~73k for Opus 5 (~27k for GPT-6 Astra), calling it among the least token-efficient models tested โ€” the two claims measure different effort regimes, so realized savings look workload-dependent. Pricing is corroborated across five-plus independent sources; the token-efficiency picture is the contested part.

Why it matters to Scott

Anthropic now prices cache reads as 'the majority of costs in agentic and coding workloads' โ€” a dated receipt for the cache-read-dominated agent economics behind his Agent Token Manifesto and cost-tiered routing, and the 60% cut directly reprices the opus tier on his LiteLLM gateway and the numbers in his published LLM pricing guide. It also rewrites the write/read tradeoff his caching radar cases (cache warming, TTL analyzer) compute against, while the contested '~40% cheaper' claim versus daily.dev's least-token-efficient measurement is a live instance of the per-task-cost-vs-published-price divergence his budget-calculator canon already argues.
dev:project.llmreportdev:concept.cost-tiered-llm-routingdev:technology.litellmradar:concept.inference-economicsradar:concept.prompt-cachingradar:concept.model-routingradar:concept.token-efficiencyradar:concept.open-weight-modelsradar:gpt6-prompt-cache-controlsradar:openai-gpt56-sol-api-price-cutradar:cache-tax-idle-session-warmingradar:claude-code-cache-ttl-analyzerradar:anthropic-context-compaction-cost-reversal
queries asked of Scott's wikis
  • per-task cost math token efficiency vs per-token price
  • prompt caching strategy cache-read dominated context design
  • agent harness budget calculator pricing assumptions
  • agent-maintained wiki long context maintenance cost
  • open-weights local inference cost case vs frontier api pricing
  • frontier release triage model selection criteria

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-21 14:00โญ origin echo-reconstructedAnthropic's "Introducing Claude Opus 5.5" announcement (Sep 22, 2026) is the primary source of both the numbers and the framing: "Opus 5.5 r
Anthropic on blog (echo) ยท attributed from reddit.post.1wp5kec
โ€”
09-24 15:59first on r/ClaudeAI ยท published ยท +74.0hOpus 5.5's cache-read price dropped a lot more than its token prices
Storagee4
โ€”
09-24 15:59amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wp5kec
Storagee4
peak 10 ยท 16 comments ยท 79% of case engagement
09-26 05:58amplified on r/ClaudeAIreddit.post.1wqis66
spy4x
peak 2 ยท 3 comments ยท 15% of case engagement
09-26 14:15amplified on r/ClaudeAIreddit.post.1wqrmpa
Puzzled-Ad-6854
peak 1 ยท 1 comments ยท 6% of case engagement
09-24 16:20our radar first saw it ยท +74.3hdiscovery anchor: reddit.post.1wp5kecโ€”

Evidence (4) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditOpus 5.5's cache-read price dropped a lot more than its token prices
ClaudeAI
Storagee41016
๐ŸŸง echo.blog โญAnthropic's "Introducing Claude Opus 5.5" announcement (Sep 22, 2026) is the primary source of both the numbers and the framing: "Opus 5.5 rAnthropicโ€”โ€”
๐ŸŸ  redditI priced 203 PRs from my Claude Code agents: Opus 5.5 cost about half of Sonnet 5 per line of code
ClaudeAI
spy4x23
๐ŸŸ  redditOpus 5.5 benchmarks from two outside sources side-by-side
ClaudeAI
Puzzled-Ad-685411

Interpretation history

Decision trace