Anthropic opened its Claude 5.5 family with Opus 5.5 on September 22, 2026 β $4/$20 per million input/output tokens (20% under Opus 5), cache reads cut 60% to $0.20, ~40% claimed workload savings β with Sonnet 5.5 and Haiku 5.5 explicitly promised for 'the coming weeks.' The case concerns that promised Sonnet 5.5 release (case dated Sept 28): its evidence titles record a same-day fight over whether Sonnet 5.5 matches Opus 5.5 on coding at roughly half the per-token price, with community measurement (Artificial Analysis / Vals AI style token accounting) countering that it emits ~62% more tokens, so realized cost-per-task at max effort can exceed Opus 5.5. The supplied web results do not cover the Sonnet 5.5 launch itself β they only establish the direct Opus 5.5 precedent that headline price cuts invert at max effort (Artificial Analysis: +37% total tokens, +143% output tokens, per-task cost rising to $13.04 vs Opus 5's $10.79; another benchmark reads $5.98 vs $5.86), so the Sonnet-specific figures (70.6% score, 62% token inflation, exact pricing) rest on the echo-testimony titles and same-day community measurement and cannot be verified from these snippets.
Converges on two of Scott's own positions with fresh dated receipts: same-day independent token accounting re-derives the Mature Token Law's audit-the-conversion claim (realized cost-per-task decides, not headline per-token price), and the 'xhigh not max' finding β including Anthropic's own docs scoring max below xhigh β independently corroborates High, Not Max. It is also directly actionable for his stack: whether Sonnet 5.5 displaces Opus 5.5 as default coding model is the same model/effort tuning his LiteLLM tier aliases and Claude Code usage encode, and his trace-backed fixture comparison is exactly the instrument that could settle the conflicting echo-testimony numbers; on the radar this continues the inference-cost lineage (Brinvik's Anthropic-guide cost reversal, hidden-reasoning real-task costs, Opus 5.5 repricing).
ip:framework.the-mature-token-lawip:concept.high-not-maxip:concept.model-barbelldev:concept.cost-tiered-llm-routingdev:concept.trace-backed-agent-comparisondev:technology.litellmradar:anthropic-context-compaction-cost-reversalradar:hidden-reasoning-real-task-costsradar:claude-code-effort-controlsradar:claude-sonnet-5-permanent-pricingradar:anthropic-opus55-cache-read-repricingradar:concept.inference-costsradar:concept.token-economics
queries asked of Scott's wikis
- realized cost per task vs headline token price in coding agents
- reasoning effort levels xhigh vs max configuration tuning
- vendor benchmark claims vs independent measurement evals
- default model choice in Claude Code-style harnesses Opus vs Sonnet
- prompt cache read pricing and agentic cost accounting
- day-one frontier release adoption policy for agent stack
| source | object | author | score | comments |
| π reddit | Sonnet 5.5 vs Opus 5.5 vs Sonnet 5 in Claude Code and pi: results and behavioral differences on 10 coding tasks ClaudeAI Retrieved article excerptOpen article Β· Retrieved 2026-09-28T21:36:42.524269+00:00 # Prove your humanity
Weβre committed to safety and security. But not for bots. Complete the challenge below and let us know youβre
a real person.
[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)
[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us) | Fabulous_Pollution10 | 18 | 14 |
| π reddit | Sonnet 5.5 vs Opus 5.5 on a from-scratch Rust decompressor: same correctness, a quarter of the price ClaudeAI | _Duex | 32 | 2 |
| π reddit | Sonnet 5.5 is by far the best free model available right now ClaudeAI | DynaBeast | 0 | 32 |
| π reddit | Sonnet 5.5 is 50% cheaper, but produces 62% more tokens singularity | OnAGoat | 147 | 33 |
| π reddit | CLAUDE.md for Sonnet 5.5 based on Anthropic's official platform docs. (their own coding benchmark scored max effort below xhigh) ClaudeAI | Puzzled-Ad-6854 | 7 | 2 |
| π reddit | Sonnet 5.5 better then Opus 5.5 in Agentic coding? ClaudeAI | kpripper | 6 | 9 |
| π reddit | Sonnet 5.5 has been out for an hour. Has anyone gotten a chance to stress test it? ClaudeAI | thedirewulf | 255 | 78 |
| π reddit | Sonnet is more expensive than Opus ClaudeAI | Blake08301 | 0 | 17 |
| π reddit | Sonnet 5.5 on Vals AI benchmark, if these hold true the $20 is insane value right now ClaudeAI | software-boulder | 511 | 76 |
| π reddit | Every Sonnet 5.5 effort level has a cheaper Sol or Opus alternative with an equal or higher Artificial Analysis score singularity | OnAGoat | 193 | 22 |
| π reddit | Use Sonnet 5.5 on xhigh not max effort singularity | theimposingshadow | 28 | 8 |
| π reddit | GPT-6 Sol vs Sonnet 5.5 at the same cost per task: Sol is more efficient, Sonnet 5.5 has the higher ceiling singularity | AMBNNJ | 56 | 17 |
| π§ echo.blog β | Anthropic's Sonnet 5.5 launch post and official prompting guide: per the echoes, its benchmark tables show Sonnet 5.5 at 70.6% vs Opus 5.5's | Anthropic | β | β |
| π reddit | Anthropic sets a new AA record with sonnet 5.5 singularity | Gohab2001 | 97 | 11 |
| π§ hn | Claude Sonnet 5.5 (Max Effort) Intelligence, Performance and Price Analysis | Topfi | 4 | 0 |
| π reddit | sonnet 5.5 vs GPT-6 Sol on the same 5 SaaS builds, one JS framework each (React, Vue, Svelte, Angular, Solid) ClaudeAI | marvijo-software | 2 | 0 |
| π reddit | sonnet 5.5 vs opus 5.5 on the same frontend spec: both passed every hidden test, sonnet did it in 2 minutes for a fifth of the price ClaudeAI | _Duex | 6 | 3 |
| π§ hn | Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index | spenvo | 8 | 3 |
| π§ hn | When to choose Sonnet over Opus, what it costs, and how to tune it | pretext | 2 | 0 |
| π reddit | I replayed Sonnet 5.5's starter pick 21 times with the exact same prompt. It grabbed the closest Poke Ball 13 times. Opus 5.5 never did. OpenAI | VibeCodyH | 5 | 4 |
| π reddit | Sonnet 5.5 ranks #2 on our writing benchmark! ClaudeAI | OnlyProggingForFun | 5 | 1 |
| π reddit | Sonnet 5.5, 410M Output tokens from Intelligence Index making it the most verbose model. still worth it? singularity | Roflxd88 | 17 | 5 |
| π reddit | Sonnet 5.5 vs Opus 5.5: the Terminal-Bench "win" is Max vs Xhigh. Here's the fair comparison. ClaudeAI | Intelligent-Lynx-953 | 4 | 15 |
| π reddit | Sonnet 5.5 has the same problem as Sonnet 5 ClaudeAI | YakFull8300 | 0 | 13 |
| π reddit | Opus 5.5 vs Sonnet 5.5 : 3D steampunk whale modeling ClaudeAI | Fun-Meaning-6474 | 748 | 65 |
2026-09-29T13:53:48Z
No new meaning this window: the velocity spike is the launch-day Vals thread (489 pts) decaying above a shrinking cohort β the magnitude-valve reading is residual accrual on the big launch threads, not fresh periphery (HN legs sit at 2β8 pts) β the comment tick lands on the already-incorporated effort-artifact post, and the flagged whale comparison was attached last window. With the displacement question settled by converging independent replications, the launch episode resolves as absorbed; a final-weights AA re-run or an official Anthropic efficiency response can re-open the cost side as its own evidence.
2026-09-29T13:26:12Z
evidence attached: reddit.post.1wt9cdl β Same-day measured Opus 5.5 vs Sonnet 5.5 comparison ($156 vs $109, token counts) on a real workload showing the cost-per-task question is workload-dependent β material context for the case's decisive economics.
2026-09-29T09:07:49Z
No new meaning this window: the flagged post is a zero-score restatement of the already-incorporated AA max-effort cost inversion, and its own comments push back on the high-effort framing. With periphery static for multiple windows (no new platforms, communities, or implementations) and rate down to ~15% of peak, this is launch tail, not spread β heat steps to low; the magnitude-valve reading reflects residual accrual on the big launch threads against a decaying cohort rather than fresh periphery, and the AA re-run on final weights stays the re-ignition trigger that will arrive as its own evidence.
2026-09-29T08:25:12Z
evidence attached: reddit.post.1wt3o8m β Independent user measurement echoing the case's counter-claim that Sonnet 5.5's token volume makes it cost more per task than Opus 5.5.
2026-09-29T07:51:37Z
The same-effort table demotes the headline parity claim to an effort artifact: at equal xhigh, Sonnet trails Opus ~5 points (61.5 vs 66.4) on Anthropic's own benchmark, so the case's meaning shifts from 'does Sonnet displace Opus' to settled terms β Opus keeps the long-horizon/equal-effort crown, and Sonnet 5.5 is a price-performance play at xhigh on short, well-specified tasks. The velocity spike is the launch-day Vals post accruing, not new spread; with periphery static (no new platforms, communities, or implementations this window) state steps down from accelerating, while heat holds medium on top-percentile residual rates and the still-unlanded AA re-run on final weights β the one live item that could flip the efficiency narrative.
2026-09-29T07:26:16Z
evidence attached: reddit.post.1wt3baa β Directly contradicts the Sonnet-beats-Opus headline by showing the Terminal-Bench gap was an effort-setting artifact, with a full same-effort table β material input to whether Sonnet 5.5 displaces Opus.
2026-09-29T07:07:27Z
This window adds no new substance β the flagged verbosity datum (410M output tokens) was already attached and incorporated, and all engagement deltas are trivial amplification of the settled effort-and-horizon rule. What changed is the attention phase: launch peak is clearly passed (56 vs 297 pts/h, cooling, no new implementations or outlets in hours), so heat steps down to medium β not low, because rates are still top-percentile across three platforms and the AA re-run on final weights, the one live item that could flip the efficiency narrative, has not landed.
2026-09-29T06:24:58Z
evidence attached: reddit.post.1wt1tzt β Independent Artificial Analysis verbosity data (410M output tokens, most verbose model) feeds the case's decisive cost-per-task question.
2026-09-29T03:33:24Z
The writing benchmark extends the token-bloat corroboration beyond coding but inverts the expected conclusion: even at ~130k tokens/script on max, Sonnet 5.5 beats Fable 5.1 at less than half the per-script price β so realized-cost inversion is horizon/domain-specific (long agentic trajectories), not a blanket max-effort penalty, sharpening the effort-and-horizon tuning rule into its near-final form. Engagement is past launch peak (149 vs 294 points/h) but still top-percentile across three platforms; heat stays high on the magnitude of independent replications and because the AA re-run on final weights β the one live item that could flip the efficiency narrative β has not yet landed.
2026-09-29T02:30:36Z
evidence attached: reddit.post.1wsw08f β Independent third-party measurement corroborating the token-bloat side: Sonnet 5.5 at max effort emits ~130k tokens/script and costs more than Opus, matching the community counter-measurement.
2026-09-29T01:12:55Z
The new independent replications split the realized-cost question BY TASK TYPE rather than resolving it one way: Sonnet's discount holds decisively on short, well-specified tasks (35/35 hidden-test pass at $0.31 vs Opus $1.42 and 4x faster; ~quarter price on the Rust DEFLATE task) but inverts on long-horizon agentic runs ($10.49 vs Opus $8.44, 470 vs 271 turns), sharpening the shaping conclusion from 'no blanket displacement' into an effort-and-horizon tuning rule that Anthropic's own selection guidance now echoes. Live contradiction kept open: AA's token-inflation figures reportedly came from a pre-release snapshot and are being re-run on final weights.
2026-09-29T00:31:57Z
evidence attached: reddit.post.1wsuaiq β Independent same-day controlled measurement of Sonnet 5.5 vs Opus 5.5 cost-per-task ($10.49 vs $8.44, 470 vs 271 turns) directly bears on whether Sonnet 5.5's headline discount survives realized task economics.
2026-09-29T00:31:57Z
evidence attached: hn.story.49884820 β Anthropic's own Sonnet-vs-Opus selection and cost guidance materially contextualises whether Sonnet 5.5 displaces Opus as the default coding model.
2026-09-29T00:31:57Z
evidence attached: hn.story.49885725 β Independent Artificial Analysis measurement placing Sonnet 5.5 just behind Opus 5.5 is third-party quality-side evidence the displacement question needs.
2026-09-29T00:31:57Z
evidence attached: reddit.post.1wsrb1z β Independent hidden-test run measuring realized cost-per-task (Sonnet $0.31 vs Opus $1.42 at equal 35/35 pass) is exactly the decisive evidence this case hinges on.
2026-09-29T00:31:56Z
evidence attached: reddit.post.1wstxgz β Controlled five-framework build-off (Sonnet 244/250 vs Sol 206/250, faster) extends the realized cost-per-task measurement question to a cross-lab comparison.
2026-09-28T23:47:06Z
The day-one price-parity fight has crystallized into a structured answer rather than a binary: small practitioner replications confirm task-level parity (30/30 in Claude Code; same correctness at ~quarter price on a Rust task), while AA measurement shows the headline discount inverting at max effort (+62% tokens, record output, no effort level dominating Sol/Opus) with community consensus landing on xhigh as the value point β a live replication of High-Not-Max. A single comment flags AA's figures as pre-release-snapshot and being re-run, so the decisive cost-per-task numbers stay unsettled; heat held high on cross-platform top-decile spread (98th percentile, magnitude valve) plus the pending AA re-run, not on settled proof.
2026-09-28T23:28:17Z
evidence attached: hn.story.49882688 β Artificial Analysis's measured intelligence/price analysis of Sonnet 5.5 max effort is exactly the independent measurement this case is waiting on β independent corroboration, not company-reported.
2026-09-28T23:28:17Z
evidence attached: reddit.post.1wsn8oh β Artificial Analysis data (record output tokens, Astra cheaper to run) is third-party evidence directly bearing on the realized cost-per-task question at the heart of the case.
2026-09-28T21:48:03Z
grounded: converges/high β Converges on two of Scott's own positions with fresh dated receipts: same-day independent token accounting re-derives the Mature Token Law's audit-the-conversio
2026-09-28T21:39:38Z
case created β Four scout proposals and their orphan attaches are one day-old Anthropic frontier release echoing across twelve same-day objects with directly conflicting price-performance readings, so they consolidate into a single high-heat case anchored on the official docs.