Skip to content
Review

Claude Opus 5.5 Review: The 40% Cost Claim, the Real Math, and Who Should Migrate

Published September 23, 2026 · by Pondero Reviews

4.5

The short version

Claude Opus 5.5 launched September 22, 2026 at $4/$20 per million tokens, a 20% cut from Opus 5 and the first Opus-tier model Anthropic has priced below its predecessor. Here is what the 40% savings claim actually means, the four breaking API changes, and the buy call for API teams and subscribers, with pricing dated 2026-09-22.

Pros

  • ✓Per-token prices dropped 20% to $4/$20 input/output, the first Opus-tier model Anthropic has launched below its predecessor's price (per digitalapplied, September 22, 2026)
  • ✓Cache reads fell 60% to $0.20 per million, and cache reads are the bulk of agentic and coding spend (per cellcog)
  • ✓On Anthropic's own table it beats Opus 5 on agentic coding: Terminal-Bench 4.0 of 66.4% against 52.3% (per digitalapplied)
  • ✓At the new medium default it matches or exceeds Opus 5 at high on coding and knowledge-work evals, so the cheaper setting becomes the baseline (per digitalapplied)
  • ✓1M-token context, 128K max output, June 2026 knowledge cutoff (per llm-stats)

Cons

  • ✕The headline 40% bundles the 20% price cut with a medium-instead-of-high default; pin high effort and you keep roughly the 20% (per digitalapplied)
  • ✕Four breaking API changes reject Opus 5 code with a 400, and one of them needs a session-architecture change, not a parameter swap (per digitalapplied)
  • ✕Vals AI measured regressions on legal, medical-coding, and tax high-effort workloads, so those need benchmarking before you migrate (per digitalapplied)
  • ✕Trails GPT-6 Astra on AutomationBench (40.0% against 41.4%) and on Terminal-Bench-Science (per digitalapplied)
  • ✕Every launch benchmark is Anthropic's own or a partner's, none independently replicated at launch (per digitalapplied)

Claude Opus 5.5 Review: The 40% Cost Claim, the Real Math, and Who Should Migrate

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, a 20% cut from Opus 5 and the first time Anthropic has launched an Opus-tier model below the price of the one it replaces, per digitalapplied. Anthropic markets a bigger number, 40% cheaper, and it scrapped the old five-hour usage caps on the way, per Benzinga. The verdict up front: if you run Opus 5 in production, migrate, but budget an afternoon for four breaking API changes and benchmark your legal, medical-coding, and tax workloads before you move them.

Here is the decision before the detail. API team on Opus 5: migrate, test effort levels on your own traffic, benchmark the high-risk categories first. API team on Fable 5.1: Opus 5.5 lists at 40% of Fable's per-token price, per tech-insider (fetched 2026-09-23), so run it against your Fable baseline on your hardest jobs. Pro, Max, or Team subscriber: nothing to do, the higher session limits apply automatically, per cellcog. The rest is why.

Quick specs

SpecClaude Opus 5.5Source
Input price$4 / M tokensllm-stats
Output price$20 / M tokensllm-stats
Cache read$0.20 / M tokensdigitalapplied
Context window1M tokensllm-stats
Max output128K tokensllm-stats
Knowledge cutoffJune 2026llm-stats
Default effortMedium (adaptive thinking always on)digitalapplied
Fast mode$8 / $40, 2.5x throughput, Claude API research preview onlycellcog
LatencyP95 TTFT 9.46s, 176 chars/sec sustained floorllm-stats

Pricing breakdown

Every per-token line moved down. Input and output each fell 20%, from $5/$25 on Opus 5 to $4/$20. Cache reads took the biggest cut, dropping 60% from $0.50 to $0.20 per million, and that one matters more than the headline suggests, because cache reads are the majority of what agentic and coding sessions actually spend, per cellcog. Five-minute cache writes fell to $5, one-hour writes to $8, and the Batch API to $2/$10, all per digitalapplied.

Line (per M tokens)Opus 5.5Opus 5Fable 5.1Source
Input$4$5$10digitalapplied
Output$20$25$50digitalapplied
Cache read$0.20$0.50$0.25digitalapplied
Cache write (5 min)$5$6.25$12.50digitalapplied
Batch (in / out)$2 / $10$2.50 / $12.50$5 / $25digitalapplied

Anthropic pricing via digitalapplied, September 22, 2026.

Now the "40%" claim, because it is conditional and the condition costs money if you miss it. Anthropic's own wording is that at default settings Opus 5.5 costs 40% less than Opus 5 on typical workloads, per digitalapplied. Default settings do a lot of work in that sentence. Opus 5.5 defaults to medium effort; Opus 5 defaults to high. Anthropic's prompting guide says Opus 5.5 at medium matches or exceeds Opus 5 at high on coding and knowledge-work evals, so the 40% is a 20% price cut stacked on a cheaper default that Anthropic claims is good enough. A team that inherited effort: "high" from its Opus 5 config and only swaps the model ID keeps the 20% price cut and loses most of the effort savings, because Opus 5.5 thinks more per turn than Opus 5 at the same named level, per digitalapplied.

Work a concrete case off the published rates, per digitalapplied. A team sending 10M output tokens a month pays $250 on Opus 5 at $25 per million. Move to Opus 5.5 at $20 and the same volume is $200, a flat $50 saved from the price cut, before you touch the effort dial. If that team was running high effort on Opus 5 and medium turns out to clear its quality bar, the token count per task drops too and the monthly saving climbs toward the $100 mark. That second $50 is the part you have to earn by testing, and the figures here are an illustrative example at a single volume, not a quote. The honest read: bank the 20%, treat the rest as a hypothesis to validate on your own traffic.

Performance

On Anthropic's launch table, Opus 5.5 leads Opus 5 on coding by a wide margin. Terminal-Bench 4.0 comes in at 66.4% against Opus 5's 52.3%, per digitalapplied. That is the clean win, and it is the number to keep if you only keep one. The gap over the field is smaller: the best competing score on Terminal-Bench 4.0 was 57.9% from GPT-6 Astra, so Opus 5.5's lead over Opus 5 is larger than its lead over the frontier.

Two cautions belong next to those scores. Opus 5.5 does not lead everywhere. On AutomationBench, a business-workflow suite Zapier ran, it scored 40.0% against GPT-6 Astra's 41.4%, so the newest Claude trails on that axis, per digitalapplied. And every figure on the launch table is Anthropic's own or a partner's, none replicated by a neutral third party at release, with several margins sitting inside the reported standard error. Read the coding lead as real and the field-wide claims as vendor-published until someone reruns them.

The category that should stop a migration cold is regressions. Vals AI measured drops on legal, medical-coding, and tax high-effort workloads, despite the better Terminal-Bench score, per digitalapplied. If your production traffic lives in one of those three, benchmark it against your Opus 5 baseline before you route a single real request. A higher agentic-coding score does not carry over to extraction accuracy in a regulated domain, and here the source says it demonstrably did not.

Migration: the four breaking changes

Opus 5.5 rejects four request patterns that run cleanly on Opus 5 today, each returning a 400 rather than degrading quietly, per digitalapplied. Three are parameter-level fixes. One is not.

Breaking changeWhat Opus 5.5 doesThe fixEffort
Thinking cannot be disabledthinking: {"type": "disabled"} or a manual budget returns a 400Omit thinking or send {"type": "adaptive"}; control depth with effort, starting low where you used to disable itParam swap
No forced tool usetool_choice of any or tool returns a 400Keep auto and add strict: true, or use structured outputs; say in the prompt when a tool appliesParam swap
Thinking blocks bound to model and conversationReplaying a thinking block after the prompt, tools, or an earlier message changed returns a 400 (accounts created on or after August 31, 2026)Keep conversations append-only; re-steer with mid-conversation system messages, not edits to prior turnsSession-architecture change
Old computer-use tool rejectedDeclaring computer_20251124 returns a 400 on the Claude API and Google CloudMove to computer_toolset_20260801 (Amazon Bedrock still accepts the old tool)Param swap

The third row is the one that turns a config edit into an engineering task. If your agent framework edits earlier turns to change instructions mid-run, that pattern now fails on Opus 5.5, and the fix is to make conversation history append-only and steer through system messages instead. That is a change to how your session state is built, not a value you flip in a config file, so scope it as a small refactor rather than a find-and-replace.

One more change fails silently and is worth a check. Short notes the model writes between tool calls now arrive as thinking blocks instead of text blocks, and at the default display setting those blocks are empty, per digitalapplied. An app that streams those notes as progress updates will show a blank between tool calls with no error raised. Set a thinking.display value that returns the text if your UI depends on it.

For subscription holders

The consumer-facing change is more session room, not a price cut. Anthropic raised the five-hour usage limits on Pro, Max, and Team plans and gave subscribers a one-off rate-limit reset they can save and spend whenever they choose, per cellcog. In Claude Code the five-hour session limits rose 20% from September 22, and because Opus 5.5 is priced lower, Anthropic says it goes about 25% further inside those limits, per digitalapplied.

The two levers help different daily patterns. The 20% bump is a steady, automatic lift you feel every session, best if your work spreads evenly across the day. The reset is a single get-out-of-jail credit: you burn through a window, hit the wall, and clear it once at any moment you pick. If your usage is bursty, one heavy afternoon a week, the reset is worth hoarding for that afternoon rather than reflexively cashing on launch day.

Verdict

API teams on Opus 5: migrate, and the reason is the price you can bank plus a coding lift you can measure. Change the model ID in a branch, clear the four breaking changes (budget real time for the append-only conversation refactor), then run two or three effort levels against your Opus 5 baseline on live traffic before you cut over. If medium clears your bar you approach the full 40%; if you need xhigh you keep nearer the 20% price cut, per digitalapplied. Benchmark legal, medical-coding, and tax agents first, because that is exactly where Vals AI found regressions.

API teams on Fable 5.1: test before you switch, because the money argument is strong but not automatic. Opus 5.5 lists at 40% of Fable 5.1's per-token price for teams running API workloads, per tech-insider (fetched 2026-09-23), and the pricing tables back that out: $4/$20 against Fable's $10/$50. The first three breaking changes already apply to Fable, so you have done most of the migration work. Run Opus 5.5 at high effort against your Fable baseline on your hardest tasks, and if you are wiring it into agentic document processing or web-context retrieval, Firecrawl is the retrieval layer that pairs cleanly with a long-context coding model. If quality holds, the cost case is decisive.

Pro, Max, and Team subscribers: no action needed, the session limits already apply. The only real decision is the one-off rate-limit reset. Save it for a long agentic run that would otherwise stall mid-session, rather than spending it the day you read this. That is the whole move.

The score is 4.5 of 5. Anthropic shipped a better coding model below the price of the one it replaces, which is the rarest kind of upgrade, and cut the cache-read line that dominates agent bills. The half-point it gives up is the friction around the win: four breaking changes with one genuine refactor, measured regressions in three regulated categories, and a launch table nobody outside Anthropic has yet reproduced. None of that changes the call. For serious agentic and coding work, Opus 5.5 is the Anthropic model to run after September 22.