2026-10-11 16:34 UTC

Anthropic claims Claude Haiku 5.5 — its fastest model, first Haiku with an adjustable effort setting, and roughly 75% cheaper to run than Haiku 4.5 — becomes the default cheap high-volume/sub-agent model for coding and agent workloads; broad migration from Haiku 4.5-class routing and third-party benchmark confirmation resolve it, weak independent results or quiet fade refute it.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highanthropic model-releases inference-economics agent-harnessesAnthropic

What is this?

Anthropic released Claude Haiku 5.5 on October 7, 2026 as its fastest, cheapest small model — priced at $0.10/$0.50 per million input/output tokens for prompts ≤100k tokens (roughly 75% cheaper than Haiku 4.5), with a 5× price tier above 100k, a new adjustable 'effort' dial (default medium, adaptive thinking on), 1M context / 128k output, and a tokenizer that uses ~30% more tokens for the same text. First-party benchmarks claim large gains over Haiku 4.5 (Terminal-Bench 39.2% vs 0%, OSWorld 72.4% vs 15.7%), but independent evidence splits: a hobbyist pipeline swap confirmed ~10× cheaper with equal results for ≤100k work; a structured blind test (86 questions, 3 runs) found quality merely tying rival flash-class models (Luna, DeepSeek Flash, Gemini 3.8 Flash); multiple users report agentic sub-agent runs blowing past the 100k cliff (90–200k tokens with ~30k tool stacks) where tokenizer inflation and thinking tokens erode the savings; and one API-equivalent cost test showed 12× GPT-6 Luna's cost on a context-heavy generative task (methodology contested). The case's current assessment: cheap-default is established for bounded high-volume ≤100k work, but the hypothesized default status for agentic sub-agents looks increasingly doubtful pending Artificial Analysis long-context numbers and visible harness/routing migration.

Why it matters to Scott

The Haiku 5.5 release and its independent reception converge on multiple load-bearing Scott positions: the Model Barbell's cheap/strong split (Haiku 5.5 as the cheap breadth model), Micro-Agents Architecture's sub-agent routing and model-tier cascades, Token Economics/Mature Token Law's treatment of token spend as investment not score, High Not Max's effort-dial economics, and Context Engineering's attention-budget framing where tokenizer inflation and thinking-token overhead erode savings in agent loops. The evidence split — cheap-default confirmed for ≤100k bounded work, contested for agentic sub-agents past the 100k cliff — directly tests whether Anthropic's pricing mechanics and effort dial make Haiku 5.5 the winning cheap slot in the barbell for Scott's harness patterns, or whether the 5× long-context tier and ~30% tokenizer inflation push sub-agent workloads to stronger models. This would change routing defaults in ask, all-in-one-software, and other production systems.
ip:concept.model-barbellip:framework.micro-agents-architectureip:concept.token-economicsip:framework.the-mature-token-lawip:concept.high-not-maxip:framework.context-engineeringip:concept.model-perishabilitydev:concept.cost-tiered-llm-routingdev:project.askradar:concept.inference-economicsradar:concept.agent-harnessesradar:concept.model-routingradar:anthropic-opus55-prompting-guideradar:claude-code-effort-controlsradar:haiku-sonnet-task-length-gapradar:frontierharness-17x-cost-variationradar:anthropic-context-compaction-cost-reversalradar:claude-code-auto-mode-defaultradar:sonnet-55-release-economicsradar:openai-gpt6-sol-luna-releaseradar:concept.open-weight-modelsradar:vercel-september-open-weight-majorityradar:token-warden-memory-rentradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
  • inference economics: sub-agent routing and model-tier cascades
  • tokenizer inflation and thinking-token overhead in agent loops
  • effort-dial / adaptive-thinking economics: when smaller models at max effort cost more than larger at low effort
  • model sovereignty: closed-model pricing power vs open-weight alternatives for high-volume workloads
  • agent harness defaults: what drives migration from Haiku 4.5-class routing in production systems
  • local inference economics: cost floor for self-hosted small models vs API pricing tiers

Measured heat

now 1 pts/hpeak 1344 pts/hcomments 0/hpeers p80momentum: steady3 platformsage 123h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-06 13:00⭐ origin echo-reconstructed"Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we've ever released" — on average around 75% less to run
Anthropic on blog (echo) · attributed from reddit.post.1x03oya, reddit.post.1x03k4j, hn.story.49996438, hn.story.49996358, hn.story.49996437
—
10-07 17:56first on hacker news · published · +28.9hClaude Haiku 5.5
denysvitali
—
10-07 18:01first on r/ClaudeAI · published · +29.0hIntroducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released
ClaudeOfficial
—
10-07 18:06first on r/singularity · published · +29.1hIntroducing Claude Haiku 5.5
AMBNNJ
—
10-08 05:42first on r/LocalLLaMA · published · +40.7hDeepSeek V4.1 Flash beat Haiku 5.5 as my research sub-agent and title model (small eval)
dergachoff
—
10-08 11:16first on r/artificial · published · +46.3hClaude Haiku 5.5 is out and it is 75% cheaper than the last one
Alone-Dragonfruit602
—
10-07 17:56amplified on hacker newshn.story.49996358
denysvitali
peak 13 · 1 comments · 0% of case engagement
10-07 18:01amplified on r/ClaudeAIreddit.post.1x03k4j
ClaudeOfficial
peak 2352 · 331 comments · 34% of case engagement
10-07 18:01amplified on hacker news 👑hn.story.49996437
sfkgtbor
peak 1040 · 484 comments · 35% of case engagement
10-07 18:01amplified on hacker newshn.story.49996438
denysvitali
peak 6 · 0 comments · 0% of case engagement
10-07 18:06amplified on r/singularityreddit.post.1x03oya
AMBNNJ
peak 602 · 113 comments · 9% of case engagement
10-07 19:16amplified on hacker newshn.story.49997487
theanonymousone
peak 1 · 0 comments · 0% of case engagement
23 more amplifiers in ainews.case_chain
10-07 18:20our radar first saw it · +29.3hdiscovery anchor: reddit.post.1x03oya—
pace: p99 vs 1247 stories at the 96h mark (now 123h old) — ahead of anthropic-sonnet-5-5-release (1.1x), behind deepseek-v41-flash-beta (0.9x)

Evidence (30) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditIntroducing Claude Haiku 5.5
singularity
Retrieved article excerpt

Open article · Retrieved 2026-10-07T18:33:06.890057+00:00

Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.

Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.¹

Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.²

Along with this launch, we’re making improvements to the value of our model range. We’re halving the price of Claude Sonnet 5.5’s cache reads, which means Sonnet 5.5 now runs around 20% cheaper on most agentic work. And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.

## Performance

Here’s how Claude Haiku 5.5 performs across a range of benchmarks:

|  | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna |  | Sonnet 5.5For reference |
| --- | --- | --- | --- | --- | --- |
| Knowledge workGDPval-AA v2.1 |  | | | |
| --- | --- | --- | --- | --- |
| Knowledge workGDPval-AA v2.1 | 1620 | 735 | 1437 |  | 1840 |
| Knowledge workAA-Briefcase v1.1 |  | | | |
| Knowledge workAA-Briefcase v1.1 | 1578 | 614 | 1336 |  | 1824 |
| Computer useOSWorld 2.1 |  | | | |
| Computer useOSWorld 2.1 | 72.4%Offline subset | 15.7%Offline subset | 48.9%Offline subset |  | 83.9%Offline subset |
| Multidisciplinary reasoningHumanity’s Last Exam |  | | | |
| Multidisciplinary reasoningHumanity’s Last Exam | 45.9%no tools | 10.2%no tools | — |  | 56.9%no tools |
| 57.4%with tools | 18.7%with tools | — |  | 64.5%with tools |
| Agentic codingTerminal-Bench 4.0 |  | | | |
| Agentic codingTerminal-Bench 4.0 | 39.2% | 0.0% | 16.4% |  | 70.6% |
| Agentic codingFrontierCode 1.1 (Main) |  | | | |
| Agentic codingFrontierCode 1.1 (Main) | 46.4% | — | 42.4% |  | 52.1%Xhigh |
| Visual reasoningChartography |  | | | |
| Visual reasoningChartography | 46.4%no tools | 6.4%no tools | 29.1%no tools |  | 61.6%no tools |

For details on how we run our evaluations, see the [Haiku 5.5 System Card](https://www.anthropic.com/claude-haiku-5-5-system-card).

Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting:

Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam

Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam

OSWorld 2.1 (offline subset)Accuracy vs. cost

- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**
- **GPT-6 Luna**

0102030405060708090Partial-credit score (%)0.050.100.200.50125Cost per attempt (USD, log scale)LowMedHighXhighMax

OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks.

GDPval-AA v2.1Accuracy vs. cost

- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**
- **GPT-6 Luna**

8001000120014001600180020000Elo, as reported0.0050.010.020.050.100.200.50125Cost per task (USD, log scale)LowMedHighXhighMax

Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations.

Humanity’s Last Exam (no tools)Accuracy vs. cost

- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5**

010203040506070Score (%)0.0050.010.020.050.100.200.501Cost per attempt (USD, log scale)LowMedHighXhighMax

Humanity’s Last Exam (HLE) is a test of expert-level academic knowledge and reasoning.

In early testing, our customers reported results consistent with the performance and cost improvements shown above. Here’s what they told us about the new model:

AsanaHubSpotAlphaSenseBoxRogoCognition

AsanaHubSpotAlphaSenseBoxRogoCognition

Quote
> “We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.”

CompanyAsana

AuthorAaron Vinh, Staff Software Engineer

Quote
> “At HubSpot, we use simulated portals to evaluate new models on CRM tasks like reporting on deals. We mostly test the smaller, more efficient models, and Claude Haiku 5.5 got the best score we’ve seen on this suite yet, at 92.8% averaged over three runs. One CRM audit task asks models to identify stale but ambiguous records. Across all of the models we tested, Haiku 5.5 was fastest to complete the task, and had the highest hit rate and the lowest false positive rate.”

CompanyHubSpot

AuthorZe’ev Klapow, Distinguished Software Engineer

Quote
> “Ask in Document is one of our big sources of spend, doing about 8M calls a week in production. It answers very specific questions on top of one or a few documents. We ran 400 queries, and Claude Haiku 5.5 was a statistically significant improvement over Haiku 4.5: 0.84 vs. 0.76.”

CompanyAlphaSense

AuthorDaniel Campos, Distinguished Engineer

Quote
> “Our customers use Box AI across large volumes of their enterprise content. With widespread usage comes the need to manage efficiency and cost, and to find the best model to suit the task at hand. In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. We’d put it to use on analytical work that runs at scale, from cost reports to financial summaries and weekly recurring reviews.”

CompanyBox

AuthorYashodha Bhavnani, VP of AI Products

Quote
> “The short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries. While a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs. It’s accurate enough that we’d trust it there, and fast and cheap enough that we can run it a lot.”

CompanyRogo

AuthorAlex Wang, Applied AI

Quote
> “Claude Haiku 5.5 joins the sidekick lineup in Devin Fusion as an excellent option. With Haiku 5.5 as the sidekick, Fusion holds a top-tier FrontierCode score of 66.2 while cutting cost and latency. You can try it today in the Devin CLI with Opus 5.5 as the lead.”

CompanyCognition

AuthorWalden Yan, Co-Founder & CPO

## Pricing

The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model.

| Price per 1 million tokens | **Haiku 5.5** prompts up to / over 100k | **Haiku 4.5** | **Sonnet 5.5** |
| --- | --- | --- | --- |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Input tokens | $0.10 / $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 / $2.50 | $5.00 | $10.00 |

## Safety

**Alignment.** Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The model’s [system card](https://www.anthropic.com/claude-haiku-5-5-system-card) describes our evaluation process and results in more detail.

**Safeguards.** Consistent with its capabilities, Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.

Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our [Life Sciences Verification Program](https://www.anthropic.com/news/life-sciences-verification-program) and [Cyber Verification Program](https://www.anthropic.com/news/cyber-verification-program).

## Availability

Claude Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with `claude-haiku-5-5`.

See our [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) for details.

## Further updates

Alongside our new pricing for Claude Haiku 5.5, we’re making further improvements to the value of our models and products.

First, starting today, we’re **lowering the price of cache reads on Claude Sonnet 5.5**. Cache reads now cost 50% less: $0.10 per million tokens rather than $0.20. Because cache reads make up a large share of models’ token consumption, this reduces the cost of Sonnet 5.5 on most agentic tasks by around 20%.

For instance, here’s what the price cut means for Sonnet 5.5’s performance relative to cost on Terminal-Bench 4.0:

Terminal-Bench 4.0Accuracy vs. cost

- **Haiku 5.5**
- **Haiku 4.5**
- **Sonnet 5.5** ($0.10 cache reads)
- **Sonnet 5.5** ($0.20 cache reads)

010203040506070Score (pass@1, %)0.5012510Cost per attempt (USD, log scale)LowMedHighXhighMax

Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface.

This chart illustrates an important difference between Haiku 5.5 and our larger models. Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0. By contrast, Haiku 5.5 is best suited to more narrowly scoped tasks that might otherwise have been cost-prohibitive with previous versions of Claude—like compaction, summarization, or subagent work.

Second, this week, we’ll roll out **a new monthly API credit to all Max and Team subscribers for use on the Claude Platform**. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, [see our Help Center article](https://support.claude.com/en/articles/17154008).

For developers, we’re also **updating our Claude Python and TypeScript SDKs to add support for computer use and browser use** in beta. Haiku 5.5 is especially well-suited to these tasks, given its combination of speed, capability, and price. You can read more about this [in our Claude Platform docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-sdk).

## Footnotes

1 Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.

2 Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. This calculation also accounts for changes between Haiku 4.5 and Haiku 5.5 in how many tokens are used to complete a given piece of work: Haiku 5.5 has an updated tokenizer (similar to Sonnet 5.5’s and Opus 5.5’s), which means it uses slightly more tokens per task.
AMBNNJ602113
🟠 redditIntroducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released
ClaudeAI
ClaudeOfficial2364334
🟧 hnClaude Haiku 5.5
Retrieved article excerpt

Open article · Retrieved 2026-10-07T18:33:09.235870+00:00

[@claudeai](https://x.com/claudeai)

[Claude](https://x.com/claudeai)

[Anthropic](https://x.com/AnthropicAI)

[@claudeai](https://x.com/claudeai)

Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.

![](https://pbs.twimg.com/media/HUC-jfVW8AAz_Ic?format=webp&name=medium)

00:00

[6:01 PM · Oct 7, 2026](https://x.com/claudeai/status/2107894039626277339)·[336.8K

Views](https://x.com/claudeai/status/2107894039626277339)

[514](https://x.com/compose/post?in_reply_to=2107894039626277339)

761

9.3K

972
denysvitali60
🟧 hnClaude Haiku 5.5
Retrieved article excerpt

Open article · Retrieved 2026-10-07T18:33:10.287194+00:00

[Models & pricing](https://platform.claude.com/docs/en/models/overview)Models

# Claude Haiku 5.5Latest

For high-volume, latency-sensitive tasks such as classification, extraction, and routing

Copy page



claude-haiku-5-5[Try in playground](https://platform.claude.com/playground?model=claude-haiku-5-5)

Copy page



Context window
:   1Mtokens

Max output
:   128Ktokens

Input pricing
:   From $0.10/ MTok

Output pricing
:   From $0.50/ MTok

## Overview

Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.

For code changes, see the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide). For model IDs, pricing, and limits, see the [Claude Haiku 5.5 overview](https://platform.claude.com/docs/en/models/haiku-5-5/overview). For prompting guidance, see [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5).

[What's new in Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/whats-new-haiku-5-5)

## How it compares

| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
| --- | --- | --- | --- | --- | --- | --- | --- |
| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
| Claude Haiku 5.5This model | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |

## Specifications

### Model IDs

Claude API
:   claude-haiku-5-5

[Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
:   anthropic.claude-haiku-5-5

[Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
:   claude-haiku-5-5

[Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
:   claude-haiku-5-5

[Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
:   claude-haiku-5-5

### Pricing

Input
:   $0.10 / MTok for prompts up to 100,000 tokens$0.50 / MTok for prompts over 100,000 tokens

Output
:   $0.50 / MTok for prompts up to 100,000 tokens$2.50 / MTok for prompts over 100,000 tokens

[5m cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
:   $0.125 / MTok for prompts up to 100,000 tokens$0.625 / MTok for prompts over 100,000 tokens

[1h cache write](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
:   $0.20 / MTok for prompts up to 100,000 tokens$1 / MTok for prompts over 100,000 tokens

[Cache read](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
:   $0.01 / MTok for prompts up to 100,000 tokens$0.05 / MTok for prompts over 100,000 tokens

[Batch API](https://platform.claude.com/docs/en/build-with-claude/batch-processing)
:   50% discount on input and output

[Full price list](https://platform.claude.com/docs/en/about-claude/pricing)

### Capabilities

[Context window](https://platform.claude.com/docs/en/build-with-claude/context-windows)
:   1M tokens

Max output
:   128K tokens

[Max output (Batch API, beta)](https://platform.claude.com/docs/en/build-with-claude/batch-processing#extended-output-beta)
:   300K tokens

[Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking)
:   Adaptive

[Default effort](https://platform.claude.com/docs/en/build-with-claude/effort)
:   `medium`

Comparative latency
:   Fastest

Input → output
:   Text and images → text

Reliable knowledge cutoff
:   Jun 2026

Training data cutoff
:   Jun 2026

### Availability

[Status](https://platform.claude.com/docs/en/about-claude/model-deprecations)
:   Active (latest)

Released
:   October 7, 2026

Retirement
:   Not sooner than October 7, 2027

Platforms
:   Claude API[Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)[Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)[Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)[Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)

## Good to know

- Adaptive thinking is on by default. Control thinking depth with the [effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort).
- Omit `temperature`, `top_p`, and `top_k`, since a non-default value for any of them returns a 400 error.
- On the [Message Batches API](https://platform.claude.com/docs/en/build-with-claude/batch-processing#extended-output-beta), Claude Haiku 5.5 supports up to 300k output tokens with the `output-300k-2026-03-24` beta header.
- Query limits and capabilities programmatically with the [Models API](https://platform.claude.com/docs/en/api/models/list).

## Resources



[Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5)

Behavioral differences and prompting patterns specific to Claude Haiku 5.5.



[Reduce latency](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency)

Choose a model and effort level, shape prompts, and stream output for faster responses.



[Adaptive thinking](https://platform.claude.com/docs/en/build-with-claude/thinking)

Claude Haiku 5.5 decides when and how much to think. Steer depth with `effort`.



[Context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows)

1M tokens. How the window is counted and managed.

## Reference



[System prompt](https://platform.claude.com/docs/en/release-notes/system-prompts/claude-haiku-5-5)

The system prompt Claude Haiku 5.5 uses on claude.ai and the Claude apps.



[System card](https://www.anthropic.com/document/claude-haiku-5-5-system-card)

Safety evaluations and deployment decisions for Claude Haiku 5.5.

[Pricing](https://platform.claude.com/docs/en/about-claude/pricing)

Full price list, including batch discounts and prompt caching rates.

[Model IDs and versioning](https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions)

How model IDs, aliases, and pinned snapshots work.



[Model deprecations](https://platform.claude.com/docs/en/about-claude/model-deprecations)

Lifecycle status and retirement commitments for every Claude model.

Was this page helpful?


denysvitali131
🟧 hnClaude Haiku 5.5sfkgtbor1040484
🟧 echo.blog ⭐"Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we've ever released" — on average around 75% less to run Anthropic——
🟠 redditClaude Haiku 5.5 official prompting guide
ClaudeAI
BuffaloConscious791944
🟧 hnClaude Haiku 5.5: Intelligence, Performance and Price Analysistheanonymousone10
🟠 redditI Tested Haiku 5.5 vs Luna vs DeepSeek Flash vs Gemini 3.8 Flash on real-ish work stuff. Basically a tie on quality, big differences in speed
ClaudeAI
NiagaraPeloton15220
🟠 redditHaiku 5.5 token usage
ClaudeAI
ceramgcf2620
🟠 redditHaiku 5.5 is a game-changer (at least for me). Some real benchmarks
ClaudeAI
shadow_nik2125
🟠 redditHaiku 5.5 price to performance if there was no long context upcharge
singularity
NoFaithlessness951358
🟠 redditClaude Haiku 5.5 cost 12x more than GPT-6 Luna for the same voxel pagoda
ClaudeAI
Fun-Meaning-64741183132
🟠 redditDeepSeek V4.1 Flash beat Haiku 5.5 as my research sub-agent and title model (small eval)
LocalLLaMA
dergachoff13
🟠 redditHaiku 5.5 vs DeepSeek V4.1 Flash on my 2 boring agent jobs: DeepSeek kept both
ClaudeAI
dergachoff83
🟠 redditHaiku 5.5 vs DeepSeek V4.1 Flash on a real agent pipeline: ~3x cheaper per video, 2.6x faster, and it's not because of the per-token price
ClaudeAI
zzJoeyyy21
🟠 redditClaude Haiku 5.5 is out and it is 75% cheaper than the last one
artificial
Alone-Dragonfruit60257
🟠 redditSame prompt, same 3D voxel scene: Claude Haiku 5.5 vs GPT-6 Luna — cost & time compared
ClaudeAI
KeyoAPI02
🟠 redditIs Haiku really 75% cheaper ?
ClaudeAI
steph_pop01
🟠 redditHaiku 5.5 solved 50/55 CTF flags for $1.87: speed and cost of 15 models at every effort level
ClaudeAI
krauq_com33
🟠 redditTesting every Claude effort level vs GPT effort level (cost vs speed vs accuracy)
ClaudeAI
krauq_com113
🟠 redditHaiku is really fast, about 242 tokens/s
ClaudeAI
localaisimple4817
🟧 hnOpus 5.5 Took Fable's Job. Haiku 5.5 Proved Me Wrongjoozio20
🟠 redditHaiku 5.5 Crushes Opus 5.5 when using skills at ~15x lower cost
ClaudeAI
NoScene7932418
🟠 redditI benchmarked Haiku 5.5 against models costing 20x more on math problems
ClaudeAI
Hoddmachine11
🟠 redditI tested Haiku 5.5, Sonnet 5.5 and Opus 5.5 as subagents on 6 real tasks
ClaudeAI
emarkosov46
🟠 redditWilson's Survival Guide for October 2-9, 2026 now available!
ClaudeAI
ClaudeAI-mod-bot53
🟠 redditIs Haiku 5.5 smarter than Opus 4.6?
ClaudeAI
Voiston441820
🟠 redditIs Haiku 5.5 actually worth using? I didn’t think I’d care, but the under-100K pricing changed my mind for API work
ClaudeAI
tjrobertson-seo015
🟠 redditPSA: Claude Code 2.1.289's cost readout prices Haiku 5.5 at Opus rates (17x too high in my runs)
ClaudeAI
Sufficient-Storage8752

Interpretation history

Decision trace