Anthropic Releases Claude Opus 5.5, Cutting Prices From Opus 5
The new flagship model is Anthropic’s bet that agentic work will be won as much by lower serving costs and faster output as by benchmark gains.
Loading page…
The new flagship model is Anthropic’s bet that agentic work will be won as much by lower serving costs and faster output as by benchmark gains.
Listen to this story
Anthropic has made Claude Opus 5.5 its flagship for coding, computer use, and knowledge work, pairing claimed capability gains with lower operating costs. Default workloads are said to cost 40% less than Opus 5, while a faster mode costs twice as much as standard pricing. The practical test will be whether savings persist on long-running agents, where cache reads dominate, and whether benchmark results translate to production.
Standard pricing is $4 per million input tokens, $20 output, and $0.20 for cache reads—down from $0.50 on Opus 5.
Fast mode reaches 2.5× speed but doubles prices to $8 per million input tokens and $40 output.
Anthropic reports 66.4% on Terminal-Bench 4.0 and 81.8% partial on OSWorld 2.0, using maximum adaptive thinking.
Anthropic has released Claude Opus 5.5, its first model in the Claude 5.5 family, through Claude Code and the Claude Platform. The company is positioning the new flagship around a practical promise: stronger performance for long-running agent tasks, at a lower cost than Opus 5.
The release makes Opus 5.5 Anthropic’s leading model for agentic coding, computer use and knowledge work, according to the company. It arrives after Anthropic’s prior Opus 5 model, with the company saying the new version needs less compute to serve and generates output more than 30% faster.
Anthropic says Opus 5.5 costs 40% less than Opus 5 on typical workloads at default settings. Its listed price is $4 per million input tokens and $20 per million output tokens, while cache reads cost $0.20 per million tokens. Cache reads matter especially for coding and agent tasks because the company says they make up most of those workloads’ costs.
There is also a faster, higher-priced option. Fast mode is available in Claude Code and the Claude Platform with up to 2.5 times the speed, at $8 per million input tokens and $40 per million output tokens. Anthropic is also increasing five-hour usage limits on its Pro, Max and Team plans.
Anthropic reported leading results across its cited tests for coding, computer use and knowledge work. The company says Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0 and 81.8% partial on OSWorld 2.0.
Anthropic also offers an important qualification to its own numbers: at this level of capability, it says benchmark margins have become a less reliable guide to real-world differences. Unless noted otherwise, its published Opus 5.5 results use adaptive thinking at maximum effort, and some safety interventions can route sensitive benchmark tasks to earlier Claude models.
Alongside the performance pitch, Anthropic says Opus 5.5 achieved its strongest result to date on the company’s automated behavioral audit, an alignment suite of simulated scenarios. It says the model is less likely than recent models to take hard-to-reverse actions or exceed assigned boundaries, and is more resistant to prompt injection than Opus 5.
For biology and cybersecurity work, Anthropic is retaining controlled access. Vetted organizations can apply to use Opus 5.5 through its Life Sciences Verification Program, while the company says it will expand its Cyber Verification Program in coming weeks. Sonnet 5.5 and Haiku 5.5 are also planned for the coming weeks, leaving the next question whether the flagship’s claimed cost and safety gains will carry across the family.
Story updates
A September 25 Claude Developers breakdown puts a number on the gap between a lower rate card and a cheaper finished job. With identical token counts, its illustrative coding session costs about 31% less on Opus 5.5 than on Opus 5. Anthropic’s 40% estimate for typical workloads assumes the newer model also uses fewer tokens per task.
The distinction follows Anthropic’s September 24 explanation of why Opus 5.5 is priced for longer coding sessions. That post gave the roughly 40% estimate and pointed to pricing, model behavior and changes to Claude Code, its coding tool. The new breakdown turns the pricing part into a worked receipt, then asks readers to measure whether their own tasks produce the additional savings.
The receipt assigns each model the same 2 million cached input tokens, 200,000 fresh input tokens and 60,000 output tokens. At API list prices, Opus 5 costs $3.50 for that mix and Opus 5.5 costs $2.40. Holding token use fixed isolates the price change; it does not show how either model would perform on a real coding task.
API list-price total for the example’s fixed mix of cached input, fresh input and output tokens.
API list-price total using exactly the same illustrative token counts as Opus 5.
The rates behind that result are $4 per million fresh input tokens and $20 per million output tokens on Opus 5.5, each 20% below Opus 5. Reading tokens already stored in the cache costs $0.20 per million, a 60% reduction. A cache lets a coding session reuse previously processed conversation text rather than pay the fresh-input rate each time the model reads it again.
The 40% figure is Anthropic’s estimate for typical token-billed workloads at default settings, not a discount on every API token. It depends on Opus 5.5 taking less work to complete a task. The breakdown says a well-scoped job may take about the same number of turns on both models, leaving only the rate-card savings. An open-ended job has more room for savings if the newer model avoids a false start.
There is another wrinkle in a default-settings comparison: Opus 5.5 defaults to medium effort in Claude Code, while Opus 5 defaults to high. Effort affects how much thinking, writing and tool use a model does on each turn. The post also cautions that the same named effort level need not mean the same amount of thinking across models. A lower bill may therefore reflect both model behavior and the settings used to run it.
For someone paying by the token, the next question is not which percentage to put in a budget. It is how many turns and output tokens each model needs on the same work. Claude Developers recommends running a task on both models and checking the session’s input, output and cached-token figures with the Claude Code command /usage. Thinking is billed as output, including when Claude Code displays only a summary of it.
The comparison also has to count completion, not just a cheap first attempt. The breakdown warns that reducing effort, context or model size can save tokens but trigger a retry that costs more. Its worked numbers establish what the lower rates do for one token mix; whether the larger workload estimate holds depends on how reliably each model finishes the reader’s actual tasks.
claude.comClaude Opus 5.5 is now available in Cursor! It's the new top model on CursorBench at 57.8% (Max) and costs 40% less per task than Opus 5.
Claude Opus 5.5 is now available as the Standard effort level in Perplexity Computer. On our WANDR benchmark, it scored 0.610 at $4.13 per task, slightly outperforming Claude Fable 5.1 while costing 67.6% less per task.
Editorial analysis
Anthropic is presenting Opus 5.5 as a correction to a familiar frontier-model tradeoff: stronger systems have often demanded more time and tokens. The notable claim is not simply that the model scores well, but that it can handle comparable work at a substantially lower typical cost than Opus 5. The important evidence to watch next is whether that efficiency holds in ordinary deployments, where Anthropic itself says benchmark margins are becoming a less reliable guide to real-world differences. The coming Sonnet 5.5 and Haiku 5.5 releases will also show how much of this efficiency-and-safety package reaches beyond the flagship tier.
Loading discussion...
Join the conversation
Explain where you would keep human review.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.