This case tracks a rumor that Anthropic's Sonnet 5.5 β the unreleased mid-tier of the Claude 5.5 family β received a last-minute upgrade after circulating benchmarks showed it beating OpenAI's GPT-6 Sol at coding, with release expected Monday. The supplied coverage establishes the context but not the claim: Simon Willison (Sept 22) reports Anthropic saying Sonnet 5.5 and Haiku 5.5 are 'coming soon,' the same day Opus 5.5 launched opposite OpenAI's GPT-6 Sol and Luna amid an active mid-tier price war. Anthropic's recent form makes the claim plausible β Sonnet 5 beat GPT-5.5 on all six comparable benchmarks at 60β67% less cost, and Anthropic claims Opus 5.5 beats GPT-5.6 Sol on a coding benchmark at ~one-third the price β but nothing here corroborates the 'last-minute upgrade' or the specific beat-GPT-6-Sol figure, and the benchmark picture is conflicting: BenchLM has GPT-5.6 Sol far ahead of current Sonnet 5 (80.7 vs 69.88, non-overlapping intervals), and AI Catchup warns even GPT-6 Sol doesn't clearly beat its predecessor on hard coding. One framing caveat: Sol is OpenAI's mid-tier model under the GPT-6 Astra flagship, so calling it 'OpenAI's coding lead' is the case's gloss, not something the snippets establish.
If Sonnet 5.5 ships Monday at anything like claimed standing, it lands directly in Scott's pinned-default decision β the LiteLLM cheap/medium/opus aliases, ask's codex default, and Claude Code's model slot β and the 'beats GPT-6 Sol' claim sits in exactly the contested-vendor-benchmark zone (BenchLM's intervals point the other way) that his evaluation-driven-development gate exists to arbitrate before any swap. It also hands dated receipts to his model-perishability argument: a vendor giving a model a last-minute upgrade because rival benchmarks circulated, mid price war, is the repricing-and-re-evaluation cadence that page predicts β though the rumor itself is unconfirmed, so the bearing stays conditional.
ip:concept.model-perishabilityip:concept.evaluation-driven-developmentip:concept.model-barbelldev:technology.litellmdev:technology.claude-codedev:project.askradar:openai-gpt6-sol-luna-releaseradar:concept.model-releasesradar:concept.inference-economicsradar:concept.benchmark-integrityradar:haiku-sonnet-task-length-gap
queries asked of Scott's wikis
- pinned coding model switching cost and eval re-runs
- mid-tier model price-per-task in agent loops
- vendor benchmark claims vs independent coding evals
- frontier release cadence same-day counter-launches price war
- coding agent harness default model choice
- cache pricing effect on agentic workload cost
2026-09-28T20:18:20Z
The Monday fuse resolved on schedule: Sonnet 5.5 shipped with an official release, published system card, and cross-platform spread (HN front page plus multiple Reddit threads), proving out the release prediction β but the leak's headline standing is only partially borne out (2nd on Artificial Analysis, near Astra at max effort, ~Sol 5.6 at high, no clean independent coding win over GPT-6 Sol), and the 'last-minute upgrade' mechanism is now unverifiable and moot. Case closes absorbed; the live question moves downstream to whether the model earns a slot in Scott's pinned defaults, where the early economics (Opus 5.5 outperforms it per dollar in most configs) don't yet make a swap case.
2026-09-28T18:36:54Z
evidence attached: hn.story.49881910 β The published Sonnet 5.5 system card is direct confirmation the rumored release actually shipped at its claimed standing.
2026-09-28T18:36:54Z
evidence attached: hn.story.49881889 β Anthropic's own release page for Claude Sonnet 5.5 is the confirmed-ship evidence the open case is watching for.
2026-09-28T18:36:53Z
evidence attached: hn.story.49881850 β Independent HN front-page spread of the official release the same day β cross-community corroboration for the open case.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wslxcu β Primary release post linking Anthropic's official page; the direct confirmation of the release this open case anticipated.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wslzj5 β Confirms Sonnet 5.5 actually shipped with official speed/cost claims (30% faster, 30% cheaper) the case needs for re-judging.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wsmeu8 β Third-party Artificial Analysis ranking directly evidences the 'claimed standing' the open case is waiting to confirm.
2026-09-28T18:36:53Z
evidence attached: reddit.post.1wslxzs β Confirmed Sonnet 5.5 launch with material specifics (30% faster/cheaper usage, first Sonnet with frontier-grade cyber safeguards) that the watching release case must now re-judge against the beating-GPT-6-Sol framing.
2026-09-28T14:12:43Z
The 'substantive evidence' trigger is a second Reddit post ('Sonnet 5.5 may release today', 21 pts) linking the same kimmonismus X post already held as a second-hand echo β derivative amplification of the single Lyra-rooted rumor chain, not a second witness or new fact; its only content is timing consistency with the priced Monday window. The case stays a single-witness watch ~4h from its 11:00 PT resolution event; heat holds high as look-cadence for that fuse, not a spread call β the numbers line is post-peak (~3.5 pts/h, 54th percentile, steady) and the magnitude-valve flag still reads one dominant thread plus its own echoes, not periphery expansion.
2026-09-28T13:35:16Z
evidence attached: reddit.post.1wsduu7 β Imminent-release timing signal (21-score rumor post) bearing directly on the watched Sonnet 5.5 shipping case.
2026-09-28T01:38:42Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-27T14:42:32Z
Pre-event quiet, not a new wave: the velocity-spike trigger reflects the thread's earlier crest β the post crept 360β404 pts / 60β69 comments on low-substance chatter (Haiku impatience, DevDay-Tuesday speculation) while current rate is ~8.5 pts/h and cooling. No new evidence, second witness, or vendor word; the case's meaning is unchanged and stays priced to Monday 11:00 PT.
2026-09-27T07:57:55Z
Provenance, not substance, moved: the Reddit post is an OCR'd screenshot of verified leaker Lyra's tweet carrying an explicit Monday 11:00 PT release time (origin walk conf 0.8, prior confirmed calls), upgrading the case from anonymous title to sourced single-witness rumor β seed gives way to watching on that plus Anthropic's official 'coming soon' anchor. Attention crested (~103 pts/h peak, now ~20/h, cooling) with zero new evidence or derivatives; the magnitude-valve flag reads one dominant Reddit thread plus its own source echo, not periphery expansion, so medium heat is a watch priced to Monday's resolution event, not a spread call.
2026-09-26T22:42:11Z
origin walked (opencode/cheap-glm, conf 0.8): anchor reddit.post.1wr27ro -> echo.x.dd76ff4dd0 by lyra (@lyraxana), verified individual AI leaker account
2026-09-26T22:33:16Z
grounded: converges/medium β If Sonnet 5.5 ships Monday at anything like claimed standing, it lands directly in Scott's pinned-default decision β the LiteLLM cheap/medium/opus aliases, ask'
2026-09-26T22:25:08Z
case created β A time-boxed, high-ratio rumor (release expected Monday, concrete beat-GPT-6-Sol claim) not covered by any open case; evidence is thin β the linked screenshot is unreadable and the gist rests on the title alone β so seed at medium heat rather than high.