Anthropic released Claude Haiku 5.5 on October 7, 2026 as its fastest, cheapest small model — priced at $0.10/$0.50 per million input/output tokens for prompts ≤100k tokens (roughly 75% cheaper than Haiku 4.5), with a 5× price tier above 100k, a new adjustable 'effort' dial (default medium, adaptive thinking on), 1M context / 128k output, and a tokenizer that uses ~30% more tokens for the same text. First-party benchmarks claim large gains over Haiku 4.5 (Terminal-Bench 39.2% vs 0%, OSWorld 72.4% vs 15.7%), but independent evidence splits: a hobbyist pipeline swap confirmed ~10× cheaper with equal results for ≤100k work; a structured blind test (86 questions, 3 runs) found quality merely tying rival flash-class models (Luna, DeepSeek Flash, Gemini 3.8 Flash); multiple users report agentic sub-agent runs blowing past the 100k cliff (90–200k tokens with ~30k tool stacks) where tokenizer inflation and thinking tokens erode the savings; and one API-equivalent cost test showed 12× GPT-6 Luna's cost on a context-heavy generative task (methodology contested). The case's current assessment: cheap-default is established for bounded high-volume ≤100k work, but the hypothesized default status for agentic sub-agents looks increasingly doubtful pending Artificial Analysis long-context numbers and visible harness/routing migration.
The Haiku 5.5 release and its independent reception converge on multiple load-bearing Scott positions: the Model Barbell's cheap/strong split (Haiku 5.5 as the cheap breadth model), Micro-Agents Architecture's sub-agent routing and model-tier cascades, Token Economics/Mature Token Law's treatment of token spend as investment not score, High Not Max's effort-dial economics, and Context Engineering's attention-budget framing where tokenizer inflation and thinking-token overhead erode savings in agent loops. The evidence split — cheap-default confirmed for ≤100k bounded work, contested for agentic sub-agents past the 100k cliff — directly tests whether Anthropic's pricing mechanics and effort dial make Haiku 5.5 the winning cheap slot in the barbell for Scott's harness patterns, or whether the 5× long-context tier and ~30% tokenizer inflation push sub-agent workloads to stronger models. This would change routing defaults in ask, all-in-one-software, and other production systems.
ip:concept.model-barbellip:framework.micro-agents-architectureip:concept.token-economicsip:framework.the-mature-token-lawip:concept.high-not-maxip:framework.context-engineeringip:concept.model-perishabilitydev:concept.cost-tiered-llm-routingdev:project.askradar:concept.inference-economicsradar:concept.agent-harnessesradar:concept.model-routingradar:anthropic-opus55-prompting-guideradar:claude-code-effort-controlsradar:haiku-sonnet-task-length-gapradar:frontierharness-17x-cost-variationradar:anthropic-context-compaction-cost-reversalradar:claude-code-auto-mode-defaultradar:sonnet-55-release-economicsradar:openai-gpt6-sol-luna-releaseradar:concept.open-weight-modelsradar:vercel-september-open-weight-majorityradar:token-warden-memory-rentradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
- inference economics: sub-agent routing and model-tier cascades
- tokenizer inflation and thinking-token overhead in agent loops
- effort-dial / adaptive-thinking economics: when smaller models at max effort cost more than larger at low effort
- model sovereignty: closed-model pricing power vs open-weight alternatives for high-volume workloads
- agent harness defaults: what drives migration from Haiku 4.5-class routing in production systems
- local inference economics: cost floor for self-hosted small models vs API pricing tiers
| source | object | author | score | comments |
| 🟠 reddit | Introducing Claude Haiku 5.5 singularity Retrieved article excerptOpen article · Retrieved 2026-10-07T18:33:06.890057+00:00 Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.¹
Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.²
Along with this launch, we’re making improvements to the value of our model range. We’re halving the price of Claude Sonnet 5.5’s cache reads, which means Sonnet 5.5 now runs around 20% cheaper on most agentic work. And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.
## Performance
Here’s how Claude Haiku 5.5 performs across a range of benchmarks:
| | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | | Sonnet 5.5For reference |
| --- | --- | --- | --- | --- | --- |
| Knowledge workGDPval-AA v2.1 | | | | |
| --- | --- | --- | --- | --- |
| Knowledge workGDPval-AA v2.1 | 1620 | 735 | 1437 | | 1840 |
| Knowledge workAA-Briefcase v1.1 | | | | |
| Knowledge workAA-Briefcase v1.1 | 1578 | 614 | 1336 | | 1824 |
| Computer useOSWorld 2.1 | | | | |
| Computer useOSWorld 2.1 | 72.4%Offline subset | 15.7%Offline subset | 48.9%Offline subset | | 83.9%Offline subset |
| Multidisciplinary reasoningHumanity’s Last Exam | | | | |
| Multidisciplinary reasoningHumanity’s Last Exam | 45.9%no tools | 10.2%no tools | — | | 56.9%no tools |
| 57.4%with tools | 18.7%with tools | — | | 64.5%with tools |
| Agentic codingTerminal-Bench 4.0 | | | | |
| Agentic codingTerminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | | 70.6% |
| Agentic codingFrontierCode 1.1 (Main) | | | | |
| Agentic codingFrontierCode 1.1 (Main) | 46.4% | — | 42.4% | | 52.1%Xhigh |
| Visual reasoningChartography | | | | |
| Visual reasoningChartography | 46.4%no tools | 6.4%no tools | 29.1%no tools | | 61.6%no tools |
For details on how we run our evaluations, see the [Haiku 5.5 System Card](https://www.anthropic.com/claude-haiku-5-5-system-card).
Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting:
Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam
Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam
OSWorld 2.1 (offline subset)Accuracy vs. cost
- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**
- **GPT-6 Luna**
0102030405060708090Partial-credit score (%)0.050.100.200.50125Cost per attempt (USD, log scale)LowMedHighXhighMax
OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks.
GDPval-AA v2.1Accuracy vs. cost
- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**
- **GPT-6 Luna**
8001000120014001600180020000Elo, as reported0.0050.010.020.050.100.200.50125Cost per task (USD, log scale)LowMedHighXhighMax
Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations.
Humanity’s Last Exam (no tools)Accuracy vs. cost
- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**
010203040506070Score (%)0.0050.010.020.050.100.200.501Cost per attempt (USD, log scale)LowMedHighXhighMax
Humanity’s Last Exam (HLE) is a test of expert-level academic knowledge and reasoning.
In early testing, our customers reported results consistent with the performance and cost improvements shown above. Here’s what they told us about the new model:
AsanaHubSpotAlphaSenseBoxRogoCognition
AsanaHubSpotAlphaSenseBoxRogoCognition
Quote
> “We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.”
CompanyAsana
AuthorAaron Vinh, Staff Software Engineer
Quote
> “At HubSpot, we use simulated portals to evaluate new models on CRM tasks like reporting on deals. We mostly test the smaller, more efficient models, and Claude Haiku 5.5 got the best score we’ve seen on this suite yet, at 92.8% averaged over three runs. One CRM audit task asks models to identify stale but ambiguous records. Across all of the models we tested, Haiku 5.5 was fastest to complete the task, and had the highest hit rate and the lowest false positive rate.”
CompanyHubSpot
AuthorZe’ev Klapow, Distinguished Software Engineer
Quote
> “Ask in Document is one of our big sources of spend, doing about 8M calls a week in production. It answers very specific questions on top of one or a few documents. We ran 400 queries, and Claude Haiku 5.5 was a statistically significant improvement over Haiku 4.5: 0.84 vs. 0.76.”
CompanyAlphaSense
AuthorDaniel Campos, Distinguished Engineer
Quote
> “Our customers use Box AI across large volumes of their enterprise content. With widespread usage comes the need to manage efficiency and cost, and to find the best model to suit the task at hand. In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. We’d put it to use on analytical work that runs at scale, from cost reports to financial summaries and weekly recurring reviews.”
CompanyBox
AuthorYashodha Bhavnani, VP of AI Products
Quote
> “The short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries. While a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs. It’s accurate enough that we’d trust it there, and fast and cheap enough that we can run it a lot.”
CompanyRogo
AuthorAlex Wang, Applied AI
Quote
> “Claude Haiku 5.5 joins the sidekick lineup in Devin Fusion as an excellent option. With Haiku 5.5 as the sidekick, Fusion holds a top-tier FrontierCode score of 66.2 while cutting cost and latency. You can try it today in the Devin CLI with Opus 5.5 as the lead.”
CompanyCognition
AuthorWalden Yan, Co-Founder & CPO
## Pricing
The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model.
| Price per 1 million tokens | **Haiku 5.5** prompts up to / over 100k | **Haiku 4.5** | **Sonnet 5.5** |
| --- | --- | --- | --- |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Input tokens | $0.10 / $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 / $2.50 | $5.00 | $10.00 |
## Safety
**Alignment.** Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The model’s [system card](https://www.anthropic.com/claude-haiku-5-5-system-card) describes our evaluation process and results in more detail.
**Safeguards.** Consistent with its capabilities, Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.
Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our [Life Sciences Verification Program](https://www.anthropic.com/news/life-sciences-verification-program) and [Cyber Verification Program](https://www.anthropic.com/news/cyber-verification-program).
## Availability
Claude Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with `claude-haiku-5-5`.
See our [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) for details.
## Further updates
Alongside our new pricing for Claude Haiku 5.5, we’re making further improvements to the value of our models and products.
First, starting today, we’re **lowering the price of cache reads on Claude Sonnet 5.5**. Cache reads now cost 50% less: $0.10 per million tokens rather than $0.20. Because cache reads make up a large share of models’ token consumption, this reduces the cost of Sonnet 5.5 on most agentic tasks by around 20%.
For instance, here’s what the price cut means for Sonnet 5.5’s performance relative to cost on Terminal-Bench 4.0:
Terminal-Bench 4.0Accuracy vs. cost
- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5** ($0.10 cache reads)
- **Sonnet 5.5** ($0.20 cache reads)
010203040506070Score (pass@1, %)0.5012510Cost per attempt (USD, log scale)LowMedHighXhighMax
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface.
This chart illustrates an important difference between Haiku 5.5 and our larger models. Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0. By contrast, Haiku 5.5 is best suited to more narrowly scoped tasks that might otherwise have been cost-prohibitive with previous versions of Claude—like compaction, summarization, or subagent work.
Second, this week, we’ll roll out **a new monthly API credit to all Max and Team subscribers for use on the Claude Platform**. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, [see our Help Center article](https://support.claude.com/en/articles/17154008).
For developers, we’re also **updating our Claude Python and TypeScript SDKs to add support for computer use and browser use** in beta. Haiku 5.5 is especially well-suited to these tasks, given its combination of speed, capability, and price. You can read more about this [in our Claude Platform docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-sdk).
## Footnotes
1 Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.
2 Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. This calculation also accounts for changes between Haiku 4.5 and Haiku 5.5 in how many tokens are used to complete a given piece of work: Haiku 5.5 has an updated tokenizer (similar to Sonnet 5.5’s and Opus 5.5’s), which means it uses slightly more tokens per task. | AMBNNJ | 602 | 113 |
| 🟠 reddit | Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released ClaudeAI | ClaudeOfficial | 2364 | 334 |
| 🟧 hn | Claude Haiku 5.5Retrieved article excerptOpen article · Retrieved 2026-10-07T18:33:09.235870+00:00 [@claudeai](https://x.com/claudeai)
[Claude](https://x.com/claudeai)
[Anthropic](https://x.com/AnthropicAI)
[@claudeai](https://x.com/claudeai)
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.

00:00
[6:01 PM · Oct 7, 2026](https://x.com/claudeai/status/2107894039626277339)·[336.8K
Views](https://x.com/claudeai/status/2107894039626277339)
[514](https://x.com/compose/post?in_reply_to=2107894039626277339)
761
9.3K
972 | denysvitali | 6 | 0 |
| 🟧 hn | Claude Haiku 5.5Retrieved article excerptOpen article · Retrieved 2026-10-07T18:33:10.287194+00:00 [Models & pricing](https://platform.claude.com/docs/en/models/overview)Models
# Claude Haiku 5.5Latest
For high-volume, latency-sensitive tasks such as classification, extraction, and routing
Copy page
claude-haiku-5-5[Try in playground](https://platform.claude.com/playground?model=claude-haiku-5-5)
Copy page
Context window
: 1Mtokens
Max output
: 128Ktokens
Input pricing
: From $0.10/ MTok
Output pricing
: From $0.50/ MTok
## Overview
Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.
For code changes, see the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide). For model IDs, pricing, and limits, see the [Claude Haiku 5.5 overview](https://platform.claude.com/docs/en/models/haiku-5-5/overview). For prompting guidance, see [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5).
[What's new in Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/whats-new-haiku-5-5)
## How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
| --- | --- | --- | --- | --- | --- | --- | --- |
| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
| Claude Haiku 5.5This model | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |
## Specifications
### Model IDs
Claude API
: claude-haiku-5-5
[Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
: anthropic.claude-haiku-5-5
[Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
: claude-haiku-5-5
[Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
: claude-haiku-5-5
[Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
: claude-haiku-5-5
### Pricing
Input
: $0.10 / MTok for prompts up to 100,000 tokens$0.50 / MTok for prompts over 100,000 tokens
Output
: $0.50 / MTok for prompts up to 100,000 tokens$2.50 / MTok for prompts over 100,000 tokens
[5m cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
: $0.125 / MTok for prompts up to 100,000 tokens$0.625 / MTok for prompts over 100,000 tokens
[1h cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
: $0.20 / MTok for prompts up to 100,000 tokens$1 / MTok for prompts over 100,000 tokens
[Cache read](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
: $0.01 / MTok for prompts up to 100,000 tokens$0.05 / MTok for prompts over 100,000 tokens
[Batch API](https://platform.claude.com/docs/en/build-with-claude/batch-processing)
: 50% discount on input and output
[Full price list](https://platform.claude.com/docs/en/about-claude/pricing)
### Capabilities
[Context window](https://platform.claude.com/docs/en/build-with-claude/context-windows)
: 1M tokens
Max output
: 128K tokens
[Max output (Batch API, beta)](https://platform.claude.com/docs/en/build-with-claude/batch-processing#extended-output-beta)
: 300K tokens
[Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking)
: Adaptive
[Default effort](https://platform.claude.com/docs/en/build-with-claude/effort)
: `medium`
Comparative latency
: Fastest
Input → output
: Text and images → text
Reliable knowledge cutoff
: Jun 2026
Training data cutoff
: Jun 2026
### Availability
[Status](https://platform.claude.com/docs/en/about-claude/model-deprecations)
: Active (latest)
Released
: October 7, 2026
Retirement
: Not sooner than October 7, 2027
Platforms
: Claude API[Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)[Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)[Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)[Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
## Good to know
- Adaptive thinking is on by default. Control thinking depth with the [effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort).
- Omit `temperature`, `top_p`, and `top_k`, since a non-default value for any of them returns a 400 error.
- On the [Message Batches API](https://platform.claude.com/docs/en/build-with-claude/batch-processing#extended-output-beta), Claude Haiku 5.5 supports up to 300k output tokens with the `output-300k-2026-03-24` beta header.
- Query limits and capabilities programmatically with the [Models API](https://platform.claude.com/docs/en/api/models/list).
## Resources
[Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5)
Behavioral differences and prompting patterns specific to Claude Haiku 5.5.
[Reduce latency](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency)
Choose a model and effort level, shape prompts, and stream output for faster responses.
[Adaptive thinking](https://platform.claude.com/docs/en/build-with-claude/thinking)
Claude Haiku 5.5 decides when and how much to think. Steer depth with `effort`.
[Context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows)
1M tokens. How the window is counted and managed.
## Reference
[System prompt](https://platform.claude.com/docs/en/release-notes/system-prompts/claude-haiku-5-5)
The system prompt Claude Haiku 5.5 uses on claude.ai and the Claude apps.
[System card](https://www.anthropic.com/document/claude-haiku-5-5-system-card)
Safety evaluations and deployment decisions for Claude Haiku 5.5.
[Pricing](https://platform.claude.com/docs/en/about-claude/pricing)
Full price list, including batch discounts and prompt caching rates.
[Model IDs and versioning](https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions)
How model IDs, aliases, and pinned snapshots work.
[Model deprecations](https://platform.claude.com/docs/en/about-claude/model-deprecations)
Lifecycle status and retirement commitments for every Claude model.
Was this page helpful?
| denysvitali | 13 | 1 |
| 🟧 hn | Claude Haiku 5.5 | sfkgtbor | 1040 | 484 |
| 🟧 echo.blog ⭐ | "Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we've ever released" — on average around 75% less to run | Anthropic | — | — |
| 🟠 reddit | Claude Haiku 5.5 official prompting guide ClaudeAI | BuffaloConscious7919 | 4 | 4 |
| 🟧 hn | Claude Haiku 5.5: Intelligence, Performance and Price Analysis | theanonymousone | 1 | 0 |
| 🟠 reddit | I Tested Haiku 5.5 vs Luna vs DeepSeek Flash vs Gemini 3.8 Flash on real-ish work stuff. Basically a tie on quality, big differences in speed ClaudeAI | NiagaraPeloton | 152 | 20 |
| 🟠 reddit | Haiku 5.5 token usage ClaudeAI | ceramgcf | 26 | 20 |
| 🟠 reddit | Haiku 5.5 is a game-changer (at least for me). Some real benchmarks ClaudeAI | shadow_nik21 | 2 | 5 |
| 🟠 reddit | Haiku 5.5 price to performance if there was no long context upcharge singularity | NoFaithlessness951 | 35 | 8 |
| 🟠 reddit | Claude Haiku 5.5 cost 12x more than GPT-6 Luna for the same voxel pagoda ClaudeAI | Fun-Meaning-6474 | 1183 | 132 |
| 🟠 reddit | DeepSeek V4.1 Flash beat Haiku 5.5 as my research sub-agent and title model (small eval) LocalLLaMA | dergachoff | 1 | 3 |
| 🟠 reddit | Haiku 5.5 vs DeepSeek V4.1 Flash on my 2 boring agent jobs: DeepSeek kept both ClaudeAI | dergachoff | 8 | 3 |
| 🟠 reddit | Haiku 5.5 vs DeepSeek V4.1 Flash on a real agent pipeline: ~3x cheaper per video, 2.6x faster, and it's not because of the per-token price ClaudeAI | zzJoeyyy | 2 | 1 |
| 🟠 reddit | Claude Haiku 5.5 is out and it is 75% cheaper than the last one artificial | Alone-Dragonfruit602 | 5 | 7 |
| 🟠 reddit | Same prompt, same 3D voxel scene: Claude Haiku 5.5 vs GPT-6 Luna — cost & time compared ClaudeAI | KeyoAPI | 0 | 2 |
| 🟠 reddit | Is Haiku really 75% cheaper ? ClaudeAI | steph_pop | 0 | 1 |
| 🟠 reddit | Haiku 5.5 solved 50/55 CTF flags for $1.87: speed and cost of 15 models at every effort level ClaudeAI | krauq_com | 3 | 3 |
| 🟠 reddit | Testing every Claude effort level vs GPT effort level (cost vs speed vs accuracy) ClaudeAI | krauq_com | 11 | 3 |
| 🟠 reddit | Haiku is really fast, about 242 tokens/s ClaudeAI | localaisimple | 48 | 17 |
| 🟧 hn | Opus 5.5 Took Fable's Job. Haiku 5.5 Proved Me Wrong | joozio | 2 | 0 |
| 🟠 reddit | Haiku 5.5 Crushes Opus 5.5 when using skills at ~15x lower cost ClaudeAI | NoScene7932 | 4 | 18 |
| 🟠 reddit | I benchmarked Haiku 5.5 against models costing 20x more on math problems ClaudeAI | Hoddmachine | 1 | 1 |
| 🟠 reddit | I tested Haiku 5.5, Sonnet 5.5 and Opus 5.5 as subagents on 6 real tasks ClaudeAI | emarkosov | 4 | 6 |
| 🟠 reddit | Wilson's Survival Guide for October 2-9, 2026 now available! ClaudeAI | ClaudeAI-mod-bot | 5 | 3 |
| 🟠 reddit | Is Haiku 5.5 smarter than Opus 4.6? ClaudeAI | Voiston44 | 18 | 20 |
| 🟠 reddit | Is Haiku 5.5 actually worth using? I didn’t think I’d care, but the under-100K pricing changed my mind for API work ClaudeAI | tjrobertson-seo | 0 | 15 |
| 🟠 reddit | PSA: Claude Code 2.1.289's cost readout prices Haiku 5.5 at Opus rates (17x too high in my runs) ClaudeAI | Sufficient-Storage87 | 5 | 2 |
2026-10-11T16:01:26Z
New material evidence: Claude Code 2.1.289 displays Haiku 5.5 costs at Opus rates (17× inflation), a display bug that directly impedes the migration hypothesis. Cheap-default for ≤100k bounded work is now established across 11 independent lines; sub-agent default remains contested due to 100k context cliff, ~30–40% tokenizer inflation, thinking-token overhead, and now tooling misreporting. Cross-platform spread triggers magnitude valve.
2026-10-11T15:42:37Z
evidence attached: reddit.post.1x3amrp — Claude Code 2.1.289 displays Haiku 5.5 costs at Opus rates (17x inflation), a display bug that directly hinders the 'default cheap high-volume model' migration hypothesis.
2026-10-10T03:38:02Z
New independent benchmarks (1,200-turn math eval, 6-task subagent comparison, CTF 50/55 flags for $1.87, effort-level vs GPT-6-Luna) deepen corroboration across nine lines; cheap-default for ≤100k bounded work is now well-established, but sub-agent default remains contested due to 100k context cliff, ~30-40% tokenizer inflation, and thinking-token overhead in agent loops. Cross-platform evidence spread triggers magnitude valve.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x1x22g — User discussion of Haiku 5.5 pricing and value proposition bears on the open case about its adoption as default cheap model.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x1xxcc — Community benchmark comparison discussion directly addressing Haiku 5.5 vs Opus capabilities — independent signal on the release's positioning.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x20j8j — Community survival guide discussing Haiku 5.5 pricing, context cliff, and usage limits — independent corroboration of the release's practical impact.
2026-10-09T19:58:38Z
evidence attached: reddit.post.1x1nsaz — Subagent comparison across Haiku/Sonnet/Opus 5.5 on 6 real tasks — third-party benchmark confirming Haiku 5.5 cost advantage and quality trade-offs for agent routing.
2026-10-09T19:58:38Z
evidence attached: reddit.post.1x1o50x — Independent benchmark of Haiku 5.5 vs 4 frontier models on 1,200 math turns — third-party confirmation of Haiku 5.5 price/performance for agent workloads.
2026-10-09T19:58:38Z
evidence attached: reddit.post.1x1paqo — Independent user benchmark (20 tasks, blind-judged) showing Haiku 5.5 outperforming Opus 5.5 with skills at ~15x lower cost — direct third-party confirmation of the case's migration/viability hypothesis.
2026-10-09T17:51:10Z
Cheap-default for bounded ≤100k work is now strongly corroborated across 9+ independent lines (hobbyist pipeline, blind test, CTF benchmark, effort-level benchmark, Artificial Analysis speed, production pipeline, research evals, token usage reports, contested cost test). Sub-agent default remains contested: 100k context cliff (5× pricing), ~30-40% tokenizer inflation, and thinking-token overhead erode savings in agent loops; Artificial Analysis long-context price/perf and visible harness/routing migration still pending. Cross-platform spread confirmed via magnitude valve. Momentum cooling (3.7 pts/hr vs 1327 peak) but hot topic neighbourhood.
2026-10-09T13:51:10Z
evidence attached: hn.story.50019927 — Blog evaluation of Haiku 5.5 vs Opus 5.5/Fable, relevant to Haiku 5.5 adoption claims.
2026-10-09T07:54:58Z
New independent benchmarks (CTF 50/55 flags for $1.87, effort-level comparison vs GPT-6-Luna, Artificial Analysis ~242 tok/s speed) strengthen the cheap-default confirmation for bounded ≤100k work. Sub-agent default remains contested: tokenizer inflation (~40% more tokens), 100k cliff, and thinking-token overhead still erode savings in agent loops; Artificial Analysis long-context price/perf and visible harness migration still pending. Cross-platform spread confirmed via magnitude valve.
2026-10-09T04:52:45Z
evidence attached: reddit.post.1x15sug — Artificial Analysis benchmark screenshot confirming Haiku 5.5 ~242 tok/s speed, independent corroboration of release claims.
2026-10-09T04:52:45Z
evidence attached: reddit.post.1x1655c — Independent user benchmark comparing Haiku 5.5 effort levels vs GPT-6-Luna, corroborates release performance claims.
2026-10-08T23:06:43Z
evidence attached: reddit.post.1x0x92o — Independent CTF benchmark evidence (50/55 flags for $1.87) bearing on Haiku 5.5's claim as default cheap high-volume agent model
2026-10-08T23:06:43Z
evidence attached: reddit.post.1x148t9 — User reports personal benchmark showing Haiku 5.5 uses ~40% more input tokens but delivers near-Sonnet-4 quality at ~7x cheaper effective cost, bearing on the release's price-performance claim.
2026-10-08T16:11:10Z
Scott up-voted (feedback_interrupt) and a new head-to-head comparison (Haiku 5.5 vs GPT-6 Luna on creative coding) arrived, but it is a single low-engagement test (score 0) that adds to the contested-comparison bucket without resolving the sub-agent default question. The case remains in its holding pattern: cheap-default confirmed for bounded ≤100k work; sub-agent default contested between motionflare's production win (3× cheaper, 2.6× faster vs DeepSeek V4.1 Flash) and research sub-agent losses plus the 100k cliff/tokenizer/thinking friction. Engagement has cooled to 84 pts/h (95.7th percentile, 3 platforms) but magnitude_valve_eligible confirms broad cross-platform spread. Heat raised to medium on Scott's nudge and topic heat (all four concepts hot), pending Artificial Analysis long-context price/perf and visible harness migration.
2026-10-08T16:03:37Z
evidence attached: reddit.post.1x0sp9c — Independent head-to-head comparison of Haiku 5.5 vs GPT-6 Luna on creative coding task providing real-world cost/token/time data relevant to Haiku 5.5's claim as default cheap sub-agent model.
2026-10-08T13:58:05Z
New evidence since the material reprice is repetitive amplification — a near-zero-engagement repost of the release news and an The Decoder coverage echo — adding no new fact about cost or sub-agent fit. The case's meaning is stable: cheap-default confirmed for bounded ≤100k work, sub-agent default contested between a production pipeline win (motionflare) and research sub-agent losses plus the 100k cliff/tokenizer/thinking friction. It settles into a holding pattern pending Artificial Analysis long-context price/perf, so research attention cools without changing the verdict.
2026-10-08T12:41:57Z
evidence attached: reddit.post.1x0o904 — Independent media coverage (The Decoder) and community discussion of the Haiku 5.5 release and 75% price drop, corroborating the open case's claim about cheaper sub-agent model.
2026-10-08T11:56:57Z
A production agent pipeline benchmark (motionflare.ai) shows Haiku 5.5 3× cheaper and 2.6× faster than DeepSeek V4.1 Flash, adding weight to the cheap-agent-model hypothesis, but research sub-agent evaluations and the 100k context cliff keep the sub-agent default status contested.
2026-10-08T10:42:27Z
evidence attached: reddit.post.1x0ne7g — Independent real-world agent pipeline benchmark shows Haiku 5.5 3x cheaper and 2.6x faster than DeepSeek V4.1 Flash, supporting the case that Haiku 5.5 becomes the default cheap agent model.
2026-10-08T09:17:38Z
A fourth independent line — a head-to-head research sub-agent evaluation finding DeepSeek V4.1 Flash beating Haiku 5.5 on two real tasks — further consolidates the counter-side: cheap-default is established for bounded ≤100k high-volume work, but the hypothesized default status for agentic sub-agents looks increasingly doubtful. The 100k cliff, tokenizer inflation, and thinking-token overhead remain the structural friction. Next concrete trigger is Artificial Analysis long-context-tier price/perf.
2026-10-08T07:01:41Z
evidence attached: reddit.post.1x0iu5x — Cross-post of the same independent evaluation as 269010, further corroborating Haiku 5.5's sub-agent performance positioning.
2026-10-08T07:01:41Z
evidence attached: reddit.post.1x0ixh6 — Independent user evaluation comparing Haiku 5.5 vs DeepSeek V4.1 Flash as research sub-agent, bearing on the case's claim that Haiku 5.5 becomes the default cheap sub-agent model.
2026-10-08T03:29:42Z
grounded: converges/high — The Haiku 5.5 release and its independent reception converge on multiple load-bearing Scott positions: the Model Barbell's cheap/strong split (Haiku 5.5 as the
2026-10-07T23:38:14Z
A fourth independent line — a direct API-equivalent cost comparison showing 12x GPT-6 Luna's cost on a context-heavy generative task (token appetite plus the >100k 5x tier) — consolidates the counter-side that was already accumulating, so the case graduates from 'contested, awaiting verification' to corroborated-contested: cheap-default is established only for bounded ≤100k high-volume work, and the agentic sub-agent default claim the hypothesis names looks increasingly doubtful. Remaining open questions are AA's long-context-tier numbers and whether harness/routing defaults actually migrate at scale.
2026-10-07T23:27:07Z
evidence attached: reddit.post.1x0agoh — Independent cost benchmark showing Haiku 5.5 at 12x GPT-6 Luna's cost via token appetite and 5x >100K-context pricing directly tests the 'default cheap sub-agent model' hypothesis.
2026-10-07T22:51:39Z
First independent results are in and they split the hypothesis: a real Haiku 4.5→5.5 pipeline swap at ~10x lower cost confirms the cheap-default case for ≤100k work, while independent agentic runs blowing past the 100k cliff (plus ~30% tokenizer inflation and thinking tokens) contest it precisely in the sub-agent niche the hypothesis names, and a structured blind test shows quality merely tying rival flash-class models. The case's meaning shifts from 'loud launch awaiting verification' to 'genuinely contested: likely wins narrow high-volume routing, unproven as default agentic sub-agent'; Artificial Analysis's pending long-context-tier numbers are the next concrete trigger.
2026-10-07T22:34:02Z
evidence attached: reddit.post.1x07vqa — The 5x long-context upcharge after 100k materially constrains Haiku 5.5's cheap-default hypothesis, with third-party AA measurement pending.
2026-10-07T22:34:02Z
evidence attached: reddit.post.1x07nnw — Independent migration evidence — Haiku 4.5 to 5.5 swap at ~10x lower cost with equal results — supports the default-cheap-model hypothesis.
2026-10-07T22:34:02Z
evidence attached: reddit.post.1x0829f — Early hands-on counter-signal: agentic token usage far above expectations undercuts the cheap-default-model claim.
2026-10-07T22:34:02Z
evidence attached: reddit.post.1x09xst — Independent structured comparison (86 questions, 3 runs, blind grading) showing Haiku 5.5 merely tying Luna/DeepSeek Flash/Gemini 3.8 Flash — the kind of weak independent result the case says would refute default-model status.
2026-10-07T22:34:02Z
evidence attached: hn.story.49997487 — Artificial Analysis third-party benchmark page is precisely the independent confirmation the case names as its resolution criterion.
2026-10-07T21:37:36Z
Engagement went vertical — official r/ClaudeAI thread at 1155, HN at 388/177, top-percentile velocity across three platforms — but the marginal content is amplification of an already-digested release, not new substance: the first independent data point (a failed complex-UI generation test) sits outside Haiku's target scope, and the pricing-structure criticism was already implicit in the docs. The case's meaning is unchanged: a real, well-documented release whose adoption question now waits on third-party benchmarks and visible routing/harness migration, a daily- not hourly-cadence signal.
2026-10-07T20:34:06Z
evidence attached: reddit.post.1x06d3z — Anthropic's official Haiku 5.5 guide documents the effort dial and thinking-counts-toward-max_tokens mechanics that determine harness migration friction — central facts for the adoption hypothesis.
2026-10-07T18:35:29Z
case created — Four same-day proposals are one release episode — one announcement echoed across r/ClaudeAI, r/singularity, and two HN threads — so they consolidate into a single case carrying all five echoes, with the retrieved Anthropic announcement as anchor.