Skip to content
Guideintermediate

Claude Sonnet 5.5 vs Opus 5.5: which one to run for coding (September 2026)

Published September 29, 2026 · by Pondero Reviews

The short version

Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0 at half the per-token price, but Opus still catches more bugs on hard code-review cases. The routing rule, the cache-read math, and the pick for solo devs, review pipelines, security teams, and Copilot users.

Table of Contents

Claude Sonnet 5.5 vs Opus 5.5: which one to run for coding (September 2026)

Anthropic shipped Claude Sonnet 5.5 on September 28, six days after Opus 5.5, and on Anthropic's own launch table the cheaper model beats the flagship on Terminal-Bench 4.0, 70.6% to 66.4%, at half the per-token price (Anthropic). Make Sonnet 5.5 your default for coding work. Opus 5.5 keeps a narrower job: the hard, ambiguous cases where a missed bug costs more than the tokens. CodeRabbit's same-day code-review run puts a number on that job. On its 13 hardest known-bug cases, Sonnet 5.5 caught 6 and Opus 5.5 caught 8 to 10 depending on effort (CodeRabbit).

The routing rule below sends well-scoped work to Sonnet 5.5, escalates judgment calls to Opus 5.5, and comes with the config lines to do it. Our Opus 5.5 review still stands. You will just run Opus less often now. The launch-day facts are in our Sonnet 5.5 news report; this is the decision layer on top.

Quick specs

SpecClaude Sonnet 5.5Claude Opus 5.5Source
ReleasedSeptember 28, 2026September 22, 2026llm-stats
Input / output (per M tokens)$2 / $10$4 / $20Anthropic
Cache read / cache write (per M tokens)$0.20 / $2.50$0.20 / $5Anthropic
Context window / max output1M / 128K tokens1M / 128K tokensllm-stats
Terminal-Bench 4.070.6%66.4% (at Xhigh effort)Anthropic
CursorBench 4.055.5%57.8%Anthropic
CodeRabbit hard cases caught (of 13)6 (thinking on)8 (Standard), 10 (Max)CodeRabbit
API model IDclaude-sonnet-5-5claude-opus-5-5Anthropic, Anthropic docs

Sonnet 5.5 kept Sonnet 5's price and cut the tokens

Nothing on the rate card moved. Sonnet 5.5 is $2 per million input tokens, $10 per million output and $0.20 per million cache reads, same as Sonnet 5, per Anthropic's launch page. What changed is how many tokens a task takes. Anthropic says the model "typically needs far fewer tokens to do the same work," costs up to 30% less per task than its predecessor, and generates output more than 30% faster (Anthropic).

The customer numbers on that page are larger than the headline. Balyasny Asset Management reported about 121k tokens per answer on a 2,441-task finance suite, against 497k for Sonnet 5, per the quote Anthropic published (Anthropic). CodeRabbit priced its review runs at list rates and got $0.47 per review for Sonnet 5.5 against $1.16 for Sonnet 5, roughly 60% less, because each core review call read 110.7k input tokens instead of 247.5k and wrote 5.8k output tokens instead of 21.6k (CodeRabbit).

Every one of those figures compares Sonnet 5.5 with Sonnet 5, not with Opus. Several launch-day write-ups blur the two, and the difference matters when you are deciding whether to leave Opus.

The 2x price gap shrinks in a cached agent loop

On the rate card, Opus 5.5 costs exactly twice what Sonnet 5.5 does for input and output, and llm-stats puts the blended gap at 2.0x on a 3:1 input-to-output mix (llm-stats). One line does not double: cache reads are $0.20 per million on both models (Anthropic). In a long coding session, most of each turn's input is the repo context you already cached, so the effective gap depends on your mix.

Here is the arithmetic for one illustrative agent turn, using list prices only:

# One agent turn: 100k cached context, 5k fresh input, 5k output.
# List prices per million tokens, Anthropic launch page, 2026-09-28.
PRICES = {
    "claude-sonnet-5-5": {"cache_read": 0.20, "input": 2.00, "output": 10.00},
    "claude-opus-5-5":   {"cache_read": 0.20, "input": 4.00, "output": 20.00},
}
TURN = {"cache_read": 100_000, "input": 5_000, "output": 5_000}

for model, p in PRICES.items():
    cost = sum(TURN[k] * p[k] / 1_000_000 for k in TURN)
    print(f"{model}: ${cost:.3f}")

# claude-sonnet-5-5: $0.080
# claude-opus-5-5: $0.140

Computed from Anthropic's list prices, that turn costs 1.75x more on Opus, not 2x. Output-heavy work, like generating a large file from a short prompt, sits at the full 2x. Token counts per task also differ between the models, which this sketch ignores, so treat it as a way to estimate your own mix rather than a quote.

The other half of the math is effort. Opus 5.5's 66.4% Terminal-Bench score was recorded at Xhigh, its highest-scoring setting, and Anthropic's cost charts show Sonnet 5.5 complements Opus best at lower effort; at higher settings it "can perform comparably at a similar cost" (Anthropic). Cranking Sonnet 5.5 to Max to dodge Opus pricing buys you little. On FrontierCode 1.1, Sonnet 5.5 scored 46.2% at Max and 52.1% at Xhigh, because at Max it more often ran Claude Code's code-review skill across many subagents, which in two examined cases led to a timeout or to edits outside the task's scope (Anthropic).

Where Sonnet 5.5 matches Opus 5.5

On most of Anthropic's rows the gap is a couple of points. GDPval-AA, a knowledge-work test across 44 occupations, came in at 1844 for Sonnet 5.5 and 1846 for Opus 5.5. OSWorld 2.1 computer use landed at 80.1% against 81.8%, and CursorBench 4.0, built from real multi-file Cursor sessions, at 55.5% against 57.8% (Anthropic). Terminal-Bench 4.0 is the one agentic-coding row where Sonnet leads outright.

llm-stats' composite reads the same way. Its LLM Stats Score has Opus 5.5 at 60.3 and Sonnet 5.5 at 58.1, and Opus wins 18 of the 29 benchmarks both models report, mostly by small margins (llm-stats).

Speed is where Sonnet pulls ahead in practice. CodeRabbit gave both models the same long build prompt in side-by-side Claude Code sessions: Sonnet 5.5 finished in 29 minutes 27 seconds, Opus 5.5 in 44 minutes 50 seconds, with near-identical results and slightly higher fidelity from Opus (CodeRabbit). CodeRabbit calls that one run, not a benchmark, and so do we.

Where Opus 5.5 still earns its price

Anthropic says it plainly on its own launch page: Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment" (Anthropic). CodeRabbit's Signal set shows what that looks like in a review queue. Each of its 13 cases is a real pull request from projects including Elasticsearch, vLLM, Cilium, axios and Next.js, with one verified bug the reviewer should catch.

ConfigurationKnown issues caught (of 13)Actionable precisionSource
Opus 5.5 Max10 (76.9%)52.0%CodeRabbit
Opus 5.5 Standard8 (61.5%)66.7%CodeRabbit
Sonnet 5.5, thinking on6 (46.2%)41.2%CodeRabbit
Sonnet 5.5, thinking off5 (38.5%)38.5%CodeRabbit
Sonnet 54 (30.8%)40.0%CodeRabbit

That is a 15 to 31 point coverage gap on the hardest cases, and Opus also wastes fewer of a reviewer's minutes: in CodeRabbit's Standard configuration, two-thirds of its actionable comments hit the target bug, against about two in five for Sonnet (CodeRabbit). Four of the 13 cases defeated every Sonnet configuration CodeRabbit ran. Thirteen cases is a small set, and CodeRabbit says so, but the direction matches Anthropic's own framing.

One more flag for security teams. Sonnet 5.5 is the first Sonnet to ship Opus-style cyber safeguards: routine bug-finding in your own code is unaffected, but higher-risk security requests visibly fall back to Sonnet 5 (Anthropic). On the Claude API, the optional server-side fallback (in beta) retries cyber declines on Sonnet 5, per the migration guide. If your review prompts read like offensive security work, log the stop_reason and model fields on every response so you know which model actually answered.

A routing rule you can ship this week

WorkloadSend it toEffort and settingWhy
Bug fixes and feature edits in a codebase you knowSonnet 5.5Medium (the Claude Code default)Anthropic positions it for well-scoped everyday tasks (Anthropic)
AI review on every pull requestSonnet 5.5Thinking onMore catches than thinking off at a small cost premium (CodeRabbit)
Review of auth, payments, concurrency, migrationsOpus 5.5Max if the budget allowsMax caught 10 of 13 hard cases (CodeRabbit)
Architecture and open-ended designOpus 5.5 plans, Sonnet 5.5 implementsHigh for OpusSustained-judgment work stays on Opus (Anthropic)
Docs, slides, spreadsheetsSonnet 5.5Low or MediumNamed use case at launch (Anthropic)
Legal, medical or tax extractionBenchmark bothYour production settingSonnet wins the Legal Agent Benchmark on llm-stats (llm-stats), and our Opus review flagged measured regressions in those categories

Rows two and three carry the savings. Put one model on all review and you either overpay on routine diffs or under-review the risky ones, so split review by path. A few lines of routing does it:

import anthropic

client = anthropic.Anthropic()

HIGH_RISK = ("auth/", "billing/", "payments/", "migrations/", "crypto/")

def review(diff: str, changed_paths: list[str]) -> anthropic.types.Message:
    risky = any(p.startswith(HIGH_RISK) for p in changed_paths)
    return client.messages.create(
        model="claude-opus-5-5" if risky else "claude-sonnet-5-5",
        max_tokens=16000,
        output_config={"effort": "high" if risky else "medium"},
        messages=[{"role": "user", "content": f"Review this diff for bugs:\n\n{diff}"}],
    )

If you ran Sonnet 5 with thinking disabled, that exact request now fails. On Sonnet 5.5, thinking: {"type": "disabled"} returns a 400, and the replacement is between_tools, which Anthropic accepts only at low, medium and high effort (migration guide):

client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=16000,
    thinking={"type": "between_tools"},   # was {"type": "disabled"} on claude-sonnet-5
    output_config={"effort": "high"},     # xhigh or max returns a 400 with between_tools
    messages=[{"role": "user", "content": "..."}],
)

For the code-review case specifically, leave thinking on. In CodeRabbit's runs it caught one more case, added about 15% to the bill, and was not slower (CodeRabbit). In Claude Code, the bundled Claude API skill handles the ID swap and the breaking parameters for you:

/claude-api migrate this project to claude-sonnet-5-5

Where you can run Sonnet 5.5 today

GitHub Copilot added Sonnet 5.5 the same day it launched, generally available on Pro, Pro+, Max, Business and Enterprise (GitHub changelog). It sits in the model picker in VS Code, Visual Studio, the Copilot CLI, the coding agent, JetBrains IDEs, Xcode, Eclipse and github.com. The rollout is gradual, and on Business and Enterprise it is on by default unless an admin has turned off default model enablement. GitHub bills it at provider list pricing under usage-based billing, so Sonnet 5.5's lower token use shows up directly in your Copilot usage charges.

Cursor's model documentation lists Claude Sonnet 5.5 under the ID claude-sonnet-5-5, with a 200k default context and a 1M maximum (Cursor docs, fetched 2026-09-29). If you work in Cursor, it is a model-menu switch.

On the API, Sonnet 5.5 runs on the Claude Platform, AWS, Google Cloud and Microsoft Azure, with zero data retention available (Anthropic; The Decoder).

The pick

Solo developer or small team on Claude Code or the API: run Sonnet 5.5 at the Medium default for everything, and switch to Opus 5.5 only when an output misses twice on the same task. At $2 and $10 per million tokens against $4 and $20 (Anthropic), you pay Opus rates only on the tasks that need Opus.

Team running AI review on every pull request: Sonnet 5.5 with thinking on for the default pass, and Opus 5.5 on the paths where a missed bug is an incident. That split follows CodeRabbit's own plan, which is moving simple and moderate reviews to Sonnet 5.5 now (CodeRabbit).

Security-heavy or compliance-bound team: keep Opus 5.5 on the final review pass, use Sonnet 5.5 for first drafts and routine fixes, and log fallbacks so a Sonnet 5 answer never passes for a Sonnet 5.5 one. Benchmark any legal, medical or tax extraction on both models before you commit.

Copilot subscriber: pick Sonnet 5.5 in the model picker as soon as it appears for your account. GitHub bills it at Anthropic's list price under usage-based billing (GitHub changelog), so every task you move off an Opus-tier model costs half as much per token.

The pick flips in one case. If you find yourself running Sonnet 5.5 at Xhigh or Max to match Opus on a task, stop and use Opus 5.5: at those settings Anthropic's own charts put the two at a similar cost per task (Anthropic), and Opus is the stronger model on hard judgment calls.