AI News

Claude Sonnet 5.5 vs Opus 5.5: Prices, Token Use and Benchmarks

Our take: Start with Claude Sonnet 5.5 for routine coding, document work, and high-volume agents. Use Claude Opus 5.5 when a task is ambiguous, long-running, or costly to get wrong. Sonnet launched September 28 at half Opus’s uncached input and output token prices, but the cheaper token does not always produce the cheaper completed task. Anthropic’s own cost-versus-score chart makes that distinction visible.

Sonnet 5.5 arrived six days after Opus 5.5. Both offer a 1 million token context window and up to 128,000 output tokens. Their differences are speed, price, default thinking effort, and how much judgment they can sustain across difficult work. This comparison uses Anthropic’s Sonnet launch data, its Opus launch data, model documentation, Vals’ launch-day evaluation, and Artificial Analysis’s independent index. We have not run a matched Kingy.ai hands-on test of the two models. For a model-by-model reference, see our Sonnet 5.5 launch guide.

Sonnet 5.5 vs Opus 5.5 at a glance

Specification Claude Sonnet 5.5 Claude Opus 5.5
Released September 28, 2026 September 22, 2026
Claude API model ID claude-sonnet-5-5 claude-opus-5-5
Input / output, per million tokens $2 / $10 $4 / $20
5-minute cache write / cache read, per million $2.50 / $0.20 $5 / $0.20
1-hour cache write, per million $4 $8
Context / maximum output 1M / 128K tokens 1M / 128K tokens
Thinking / API default effort Adaptive / High Adaptive, always on / Medium
Comparative latency / reliable knowledge cutoff Fast / June 2026 Moderate / June 2026

These are Anthropic’s published Claude API specifications for Sonnet and Opus. Both take text and image input and produce text. Both are available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Anthropic calls Sonnet “fast” and Opus “moderate” in its comparative latency labels; those labels are not a measured end-to-end speed guarantee for every workflow.

What a task costs, not just what a token costs

At list price, a request using 100,000 fresh input tokens and producing 20,000 output tokens costs about $0.40 on Sonnet or $0.80 on Opus. If the 100,000 input tokens are instead billed as cache reads, the same output costs about $0.22 on Sonnet or $0.42 on Opus. Those examples exclude the initial cache write, tools, any other billable input, and retries. Anthropic offers a 50% Batch API discount on input and output tokens for both models.

The two models have the same $0.20 per million cache-read price. For a long agent session that repeatedly reuses cached context, output, new input, and cache writes determine the gap. A model that solves a task in fewer steps can overcome a higher per-token rate. Conversely, asking Sonnet to reason at Max effort can consume enough tokens to erase its sticker-price advantage.

Independent update, September 28: On the Artificial Analysis Intelligence Index v4.3.2, Sonnet 5.5 at Max effort with the default safeguard fallback scored 56 and cost a weighted average of $7.60 per index task. Opus 5.5 under the same effort and fallback setting scored 58 at $5.98 per task. Artificial Analysis reports 410 million output tokens across Sonnet’s index run and 260 million across Opus’s. In this ten-evaluation mix, Sonnet’s half-price input and output tokens did not make it the cheaper model per task. The two-point index gap has no uncertainty interval on these model pages, and these totals are specific to the evaluator’s workload and run; they do not predict token use or cost on your own tasks. The index includes the two knowledge-work components subject to Anthropic’s prerelease bug caveat below.

Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task than Sonnet 5 in its testing, despite identical base token rates. That is a comparison with the previous Sonnet, not a promise of a 30% saving against Opus. Partners in the launch material report narrower, workload-specific token savings: Slack saw about 14% fewer output tokens than Sonnet 5 on its offline Slackbot evaluations, while Balyasny reported roughly 121,000 versus 497,000 tokens per answer on its private finance suite. Neither is a matched Sonnet 5.5-versus-Opus 5.5 token count.

Benchmarks: Sonnet’s headline lead has an effort caveat

Anthropic’s launch table gives Sonnet 5.5 a 70.6% score on Terminal-Bench 4.0, above Opus 5.5’s 66.4%. The footnote and cost chart matter: Sonnet’s 70.6% uses Max effort; Opus’s 66.4% uses Xhigh. At the same Xhigh setting in that chart, Sonnet scores 61.5% and Opus 66.4%. At Medium, Sonnet scores 28.8% for $0.83 per attempt, while Opus scores 57.6% for $2.94. Sonnet’s Max run costs $12.54 per attempt; Opus’s Xhigh run costs $7.35. These are Anthropic’s benchmark configurations and costs, not a universal price or an independent Kingy test.

Anthropic launch evaluation Sonnet 5.5 Opus 5.5 What to read into it
Terminal-Bench 4.0, agentic command-line tasks 70.6% (Max) 66.4% (Xhigh) Different effort settings; at matched Xhigh, Opus leads in Anthropic’s cost chart.
FrontierCode 1.1 Main 52.1% (Xhigh); 46.2% (Max) 54.4% Opus leads; Sonnet’s Max score drops because extra review steps could time out or exceed scope.
CursorBench 4.0 55.5% 57.8% Close in this coding test.
GDPval-AA v2.1, occupational work, Elo 1844 1846 Two Elo points apart; see the prerelease bug caveat below.
AA-Briefcase v1.1, knowledge work, Elo 1811 1822 Opus leads by 11 Elo points in the reported run; see the caveat below.
Humanity’s Last Exam, with tools 64.5% 67.7% Opus leads in the reported configuration.
OSWorld 2.1, partial credit 80.1% 81.8% Close on this computer-use measure.
Chartography, no tools 61.6% 64.4% Opus leads on chart recognition without tools.

All rows above are from Anthropic’s September 28 launch table and its footnotes. Anthropic says Artificial Analysis ran the Sonnet GDPval-AA and AA-Briefcase tests on a prerelease deployment with a since-fixed structured-output bug. It expects any effect on those scores to be small and to understate Sonnet’s performance. Treat those two close comparisons with that caveat until the evaluator reruns them. The rows are reported evaluations under stated settings, not a leaderboard of ordinary app use. Anthropic also says Opus remains clearly stronger on complex, open-ended work that needs sustained judgment, even where Sonnet’s scores are close.

Vals reports similar overall scores for Sonnet and Opus, but a different cost-per-test result. Vals scored Sonnet 5.5 at 69.22% ±0.96 on its Vals Index, compared with 69.69% ±0.94 for Opus 5.5. The 0.47-point gap is smaller than the stated uncertainty. Vals also reports $20.80 per Vals Index test for Sonnet and $32.77 for Opus, making Sonnet cheaper on this evaluator’s task mix. On Vals’ Terminal-Bench 4.0 run, Sonnet scored 53.03% and Opus 61.62%. Those scores differ from Anthropic’s because the evaluator, harness, and settings differ; they should not be mixed into a single ranking. Artificial Analysis measures the opposite cost result on its separate ten-evaluation index: $7.60 per task for Sonnet and $5.98 for Opus. These dollar amounts are evaluator-specific and not directly comparable across test suites. The contrast shows why input and output prices alone cannot predict a model’s cost for your workload. Vals ran both at Max effort for most tests and allowed server-side fallbacks when safeguards intervened. Counting fallback-assisted tasks as failures lowers Sonnet’s Terminal-Bench score to 50.51%; Vals reports the corresponding Opus result at 53.54%.

Which model should you use?

  • Use Sonnet 5.5 for well-scoped bug fixes, frequent coding turns, first drafts, document and slide creation, and high-volume workflows where latency and cost per attempt matter. It is the sensible first run when a clear acceptance test lets you catch a miss.
  • Use Opus 5.5 for architecture choices, messy debugging, multi-repository changes, difficult research, and work that requires interpreting conflicting evidence or making consequential judgments. Its higher list price can pay for itself if it avoids retries or human repair.
  • Measure the whole job before routing every request to one model. Record completion rate, elapsed time, input and output tokens, cache reads and writes, tool calls, and human correction time. Test each at the effort setting you intend to deploy.

For developers migrating from Sonnet 5, Anthropic lists five breaking API changes. Among them, forced tool use can error, older computer-use tooling is not accepted on the Claude API and Google Cloud, and applications that display text between tool calls may need to change how they handle thinking blocks. Opus 5.5 also has migration changes. Budget a regression pass before swapping model IDs in a production agent.

Sonnet 5.5 is Anthropic’s first Sonnet launch with the newer cyber safeguards and fallbacks. High-risk cybersecurity requests may visibly fall back to Sonnet 5; ordinary software development should remain available. This matters when interpreting benchmark scores and when testing a security workflow. Anthropic’s Sonnet system card documents the evaluation and safeguards; its Opus system card covers the stronger model’s deployment decisions.

Editorial method, updated September 28, 2026: Prices and specifications were checked against Anthropic’s model documentation. Benchmark figures are labeled by publisher and effort level. Vals and Artificial Analysis supply independent checks. Kingy.ai did not run either model on a matched task set for this article; we will not treat vendor and independent harnesses as interchangeable.