Our take: Start with Claude Sonnet 5.5 for routine coding, document work, and high-volume agents. Use Claude Opus 5.5 when a task is ambiguous, long-running, or costly to get wrong. Sonnet launched September 28 at half Opus’s uncached input and output token prices, but the cheaper token does not always produce the cheaper completed task. Anthropic’s own cost-versus-score chart makes that distinction visible.
Sonnet 5.5 arrived six days after Opus 5.5. Both offer a 1 million token context window and up to 128,000 output tokens. Their differences are speed, price, default thinking effort, and how much judgment they can sustain across difficult work. This comparison uses Anthropic’s Sonnet launch data, its Opus launch data, model documentation, Vals’ launch-day evaluation, and Artificial Analysis’s independent index. We have not run a matched Kingy.ai hands-on test of the two models. For a model-by-model reference, see our Sonnet 5.5 launch guide.
Sonnet 5.5 vs Opus 5.5 at a glance
| Specification | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Released | September 28, 2026 | September 22, 2026 |
| Claude API model ID | claude-sonnet-5-5 |
claude-opus-5-5 |
| Input / output, per million tokens | $2 / $10 | $4 / $20 |
| 5-minute cache write / cache read, per million | $2.50 / $0.20 | $5 / $0.20 |
| 1-hour cache write, per million | $4 | $8 |
| Context / maximum output | 1M / 128K tokens | 1M / 128K tokens |
| Thinking / API default effort | Adaptive / High | Adaptive, always on / Medium |
| Comparative latency / reliable knowledge cutoff | Fast / June 2026 | Moderate / June 2026 |
These are Anthropic’s published Claude API specifications for Sonnet and Opus. Both take text and image input and produce text. Both are available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Anthropic calls Sonnet “fast” and Opus “moderate” in its comparative latency labels; those labels are not a measured end-to-end speed guarantee for every workflow.
What a task costs, not just what a token costs
At list price, a request using 100,000 fresh input tokens and producing 20,000 output tokens costs about $0.40 on Sonnet or $0.80 on Opus. If the 100,000 input tokens are instead billed as cache reads, the same output costs about $0.22 on Sonnet or $0.42 on Opus. Those examples exclude the initial cache write, tools, any other billable input, and retries. Anthropic offers a 50% Batch API discount on input and output tokens for both models.
The two models have the same $0.20 per million cache-read price. For a long agent session that repeatedly reuses cached context, output, new input, and cache writes determine the gap. A model that solves a task in fewer steps can overcome a higher per-token rate. Conversely, asking Sonnet to reason at Max effort can consume enough tokens to erase its sticker-price advantage.
Independent update, September 28: On the Artificial Analysis Intelligence Index v4.3.2, Sonnet 5.5 at Max effort with the default safeguard fallback scored 56 and cost a weighted average of $7.60 per index task. Opus 5.5 under the same effort and fallback setting scored 58 at $5.98 per task. Artificial Analysis reports 410 million output tokens across Sonnet’s index run and 260 million across Opus’s. In this ten-evaluation mix, Sonnet’s half-price input and output tokens did not make it the cheaper model per task. The two-point index gap has no uncertainty interval on these model pages, and these totals are specific to the evaluator’s workload and run; they do not predict token use or cost on your own tasks. The index includes the two knowledge-work components subject to Anthropic’s prerelease bug caveat below.
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task than Sonnet 5 in its testing, despite identical base token rates. That is a comparison with the previous Sonnet, not a promise of a 30% saving against Opus. Partners in the launch material report narrower, workload-specific token savings: Slack saw about 14% fewer output tokens than Sonnet 5 on its offline Slackbot evaluations, while Balyasny reported roughly 121,000 versus 497,000 tokens per answer on its private finance suite. Neither is a matched Sonnet 5.5-versus-Opus 5.5 token count.
Benchmarks: Sonnet’s headline lead has an effort caveat
Anthropic’s launch table gives Sonnet 5.5 a 70.6% score on Terminal-Bench 4.0, above Opus 5.5’s 66.4%. The footnote and cost chart matter: Sonnet’s 70.6% uses Max effort; Opus’s 66.4% uses Xhigh. At the same Xhigh setting in that chart, Sonnet scores 61.5% and Opus 66.4%. At Medium, Sonnet scores 28.8% for $0.83 per attempt, while Opus scores 57.6% for $2.94. Sonnet’s Max run costs $12.54 per attempt; Opus’s Xhigh run costs $7.35. These are Anthropic’s benchmark configurations and costs, not a universal price or an independent Kingy test.
| Anthropic launch evaluation | Sonnet 5.5 | Opus 5.5 | What to read into it |
|---|---|---|---|
| Terminal-Bench 4.0, agentic command-line tasks | 70.6% (Max) | 66.4% (Xhigh) | Different effort settings; at matched Xhigh, Opus leads in Anthropic’s cost chart. |
| FrontierCode 1.1 Main | 52.1% (Xhigh); 46.2% (Max) | 54.4% | Opus leads; Sonnet’s Max score drops because extra review steps could time out or exceed scope. |
| CursorBench 4.0 | 55.5% | 57.8% | Close in this coding test. |
| GDPval-AA v2.1, occupational work, Elo | 1844 | 1846 | Two Elo points apart; see the prerelease bug caveat below. |
| AA-Briefcase v1.1, knowledge work, Elo | 1811 | 1822 | Opus leads by 11 Elo points in the reported run; see the caveat below. |
| Humanity’s Last Exam, with tools | 64.5% | 67.7% | Opus leads in the reported configuration. |
| OSWorld 2.1, partial credit | 80.1% | 81.8% | Close on this computer-use measure. |
| Chartography, no tools | 61.6% | 64.4% | Opus leads on chart recognition without tools. |
All rows above are from Anthropic’s September 28 launch table and its footnotes. Anthropic says Artificial Analysis ran the Sonnet GDPval-AA and AA-Briefcase tests on a prerelease deployment with a since-fixed structured-output bug. It expects any effect on those scores to be small and to understate Sonnet’s performance. Treat those two close comparisons with that caveat until the evaluator reruns them. The rows are reported evaluations under stated settings, not a leaderboard of ordinary app use. Anthropic also says Opus remains clearly stronger on complex, open-ended work that needs sustained judgment, even where Sonnet’s scores are close.
Vals reports similar overall scores for Sonnet and Opus, but a different cost-per-test result. Vals scored Sonnet 5.5 at 69.22% ±0.96 on its Vals Index, compared with 69.69% ±0.94 for Opus 5.5. The 0.47-point gap is smaller than the stated uncertainty. Vals also reports $20.80 per Vals Index test for Sonnet and $32.77 for Opus, making Sonnet cheaper on this evaluator’s task mix. On Vals’ Terminal-Bench 4.0 run, Sonnet scored 53.03% and Opus 61.62%. Those scores differ from Anthropic’s because the evaluator, harness, and settings differ; they should not be mixed into a single ranking. Artificial Analysis measures the opposite cost result on its separate ten-evaluation index: $7.60 per task for Sonnet and $5.98 for Opus. These dollar amounts are evaluator-specific and not directly comparable across test suites. The contrast shows why input and output prices alone cannot predict a model’s cost for your workload. Vals ran both at Max effort for most tests and allowed server-side fallbacks when safeguards intervened. Counting fallback-assisted tasks as failures lowers Sonnet’s Terminal-Bench score to 50.51%; Vals reports the corresponding Opus result at 53.54%.
Which model should you use?
- Use Sonnet 5.5 for well-scoped bug fixes, frequent coding turns, first drafts, document and slide creation, and high-volume workflows where latency and cost per attempt matter. It is the sensible first run when a clear acceptance test lets you catch a miss.
- Use Opus 5.5 for architecture choices, messy debugging, multi-repository changes, difficult research, and work that requires interpreting conflicting evidence or making consequential judgments. Its higher list price can pay for itself if it avoids retries or human repair.
- Measure the whole job before routing every request to one model. Record completion rate, elapsed time, input and output tokens, cache reads and writes, tool calls, and human correction time. Test each at the effort setting you intend to deploy.
For developers migrating from Sonnet 5, Anthropic lists five breaking API changes. Among them, forced tool use can error, older computer-use tooling is not accepted on the Claude API and Google Cloud, and applications that display text between tool calls may need to change how they handle thinking blocks. Opus 5.5 also has migration changes. Budget a regression pass before swapping model IDs in a production agent.
Sonnet 5.5 is Anthropic’s first Sonnet launch with the newer cyber safeguards and fallbacks. High-risk cybersecurity requests may visibly fall back to Sonnet 5; ordinary software development should remain available. This matters when interpreting benchmark scores and when testing a security workflow. Anthropic’s Sonnet system card documents the evaluation and safeguards; its Opus system card covers the stronger model’s deployment decisions.
Editorial method, updated September 28, 2026: Prices and specifications were checked against Anthropic’s model documentation. Benchmark figures are labeled by publisher and effort level. Vals and Artificial Analysis supply independent checks. Kingy.ai did not run either model on a matched task set for this article; we will not treat vendor and independent harnesses as interchangeable.
Trending on Kingy
Keep reading with the stories getting the most attention now.
