Anthropic Releases Haiku 5.5 With 90% Lower Token Prices for Shorter Prompts
The biggest discount applies to prompts up to 100,000 tokens. Higher token counts and adjustable reasoning make the rate card only part of the savings calculation.
Loading page…
The biggest discount applies to prompts up to 100,000 tokens. Higher token counts and adjustable reasoning make the rate card only part of the savings calculation.
Listen to this story
Haiku 5.5 is designed for frequent, bounded tasks and delegated work, not as Anthropic’s choice for the hardest coding jobs. Although its short-prompt rates are 90% below Haiku 4.5’s, Anthropic estimates average operating costs fall about 75%; a new tokenizer can count identical text as roughly 30% more tokens, so developers should remeasure workloads. The release also expands context and output limits, but migration requires API changes and capacity planning because Haiku 5.5 lacks Priority Tier.
Haiku 5.5 supports a one-million-token context window, up from 200,000, and up to 128,000 output tokens.
Anthropic reports 39.2% on Terminal-Bench 4.0 at maximum effort; VentureBeat notes its medium-effort score is roughly 20%, versus Sonnet 5.5’s 70.6%.
The model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.
Anthropic released Claude Haiku 5.5 on October 7, 2026, lowering the price of repeated AI work such as summaries, classification and smaller assignments inside larger projects. Its API rates start at $0.10 per million input tokens and $0.50 per million output tokens—a 90% cut from Haiku 4.5 for prompts up to 100,000 tokens.
The discount shrinks for longer prompts. Above 100,000 tokens, Haiku 5.5 charges $0.50 per million input tokens and $2.50 per million output tokens, half Haiku 4.5’s rates. Anthropic says roughly 90% of requests to the previous model fell into the shorter category.
Those rate cuts are not the same as savings on completed work. Anthropic estimates that Haiku 5.5 costs around 75% less to run on average, accounting for request sizes and changed token consumption. Tokens are the billable units into which the model divides content.
The model’s updated tokenizer counts the same input text as approximately 30% more tokens than Haiku 4.5, depending on the content. Anthropic’s documentation therefore tells developers to recount prompts and revisit output limits and cost estimates rather than reuse their old measurements.
Published input rates fall from $1.00 for Haiku 4.5 to $0.10 for Haiku 5.5.
Published output rates fall from $5.00 for Haiku 4.5 to $0.50 for Haiku 5.5.
Anthropic positions Haiku for narrowly scoped, high-volume jobs rather than the hardest coding assignments. Financial AI company Rogo supplied one example: a larger model builds a presentation while a Haiku subagent—a model assigned one piece of the work—retrieves the revenue figure needed for a slide.
Haiku 5.5 can hold one million tokens of context, up from 200,000 in Haiku 4.5, and produce up to 128,000 output tokens, twice its predecessor’s limit. It also introduces adjustable effort to the Haiku family. Adaptive thinking is enabled by default, with medium effort as the starting setting.
The effort setting matters when reading performance claims. Anthropic’s table gives Haiku 5.5 a 39.2% score on Terminal-Bench 4.0, a test of multistep command-line work, against Sonnet 5.5’s 70.6%. VentureBeat notes that Haiku’s chart places the roughly 39% result at maximum effort, while medium scores roughly 20%. These are vendor-reported results.
Anthropic still recommends Sonnet and Opus for complex coding tasks. Its speed claim has a separate boundary: Haiku 5.5 is its fastest model at standard speeds, but Opus models running in Fast Mode are faster.
Haiku 5.5 is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Its Claude API identifier is claude-haiku-5-5. For existing integrations, the migration guide details changes beyond replacing the model name:
There is also a capacity-planning distinction: Haiku 5.5 does not support Priority Tier. Organizations with a Haiku 4.5 commitment under that service must plan capacity separately.
The coordinated announcement also halves Sonnet 5.5 cache-read rates to $0.10 per million tokens. Cache reads reuse stored context; Anthropic estimates around 20% savings on most agent workloads. Monthly API credits roll out this week: $100 for Max 5x, $200 for Max 20x and up to $500 pooled across Team users, usable on any model.
Loading discussion...
Join the conversation
Explain which would help you make a fairer comparison.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.