AI News

Claude Sonnet 5.5: Specs, Benchmarks, Pricing and the Real Cost per Task

Claude Sonnet 5.5 is Anthropic’s new mid-tier model, released on September 28, 2026, and the second member of the Claude 5.5 family after Claude Opus 5.5. It keeps Sonnet 5’s price of $2 per million input tokens and $10 per million output tokens. On Anthropic’s own benchmark table it lands within a few points of Opus 5.5, a model that costs twice as much per token, and on Terminal-Bench 4.0 it actually scores higher. Anthropic says it generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task.

Most of that holds up. The part that needs a closer look is cost. Sonnet 5.5 is cheap per token but can be expensive per task. Run it at its highest effort setting and it uses so many tokens that independent testing puts it above Opus 5.5 on cost per task, and at roughly seven times the cost per task of OpenAI’s GPT-6 Sol, which has the same list price. The effort setting you choose matters more than the model you choose.

This guide covers the full spec sheet, every published benchmark, the independent numbers, the real cost per task, the API changes that will break existing Sonnet 5 code, and the system card’s safety findings. Unless we say otherwise, benchmark scores are vendor-published results, not a Kingy AI reproduction.

Key takeaways

  • Price: $2 input / $10 output per million tokens, $0.20 for cache reads. That is identical to Sonnet 5 and exactly half of Opus 5.5 on input and output. Cache reads cost the same $0.20 on both.
  • Performance: 70.6% on Terminal-Bench 4.0, 81.3% on SWE-Bench Pro, 55.5% on CursorBench 4.0, 80.1% on OSWorld 2.1 and 1844 Elo on GDPval-AA v2.1. That is two Elo points behind Opus 5.5 and about 400 ahead of Sonnet 5.
  • Independent check: Artificial Analysis puts Sonnet 5.5 at 56 on its Intelligence Index, #3 of 216 models, behind only Opus 5.5 (58). In its own runs, Sonnet 5.5 beat Opus 5.5 on Terminal-Bench 4.0 (64% vs 60%).
  • The catch: at max effort, Artificial Analysis measured $7.60 per task and 193K output tokens per task, compared with $5.98 for Opus 5.5 and $1.06 for GPT-6 Sol. Anthropic’s efficiency claims are about lower effort levels, and that is where Sonnet 5.5 makes sense.
  • Breaking API changes: thinking: disabled now returns an error and is replaced by between_tools. Forced tool_choice is gone. Thinking blocks are bound to the model, conversation and account. The old computer-use tool is rejected on the Claude API and Google Cloud.
  • Specs: 1M-token context window, 128K max output (300K via the Batch API beta), text and image input, June 2026 knowledge cutoff.

Claude Sonnet 5.5 specs at a glance

Spec Claude Sonnet 5.5
Released September 28, 2026
API model ID claude-sonnet-5-5 (Claude API, Google Cloud, Microsoft Foundry, Claude Platform on AWS); anthropic.claude-sonnet-5-5 on Amazon Bedrock
Price (input / output) $2 / $10 per million tokens
Prompt caching $2.50 (5-minute write), $4 (1-hour write), $0.20 (read) per million tokens
Batch API 50% off: $1 / $5 per million tokens
Context window 1M tokens (about 555,000 words)
Max output 128K tokens; up to 300K on the Batch API with the output-300k-2026-03-24 beta header
Input / output Text and images in, text out
Thinking Adaptive (on by default); lowest setting is between_tools
Default effort high on the API; Medium in the Claude apps and Claude Code
Reliable knowledge cutoff June 2026
Tokenizer Same as Sonnet 5 (same text, same token count)
Minimum cacheable prompt 512 tokens (Sonnet 5 needs 1,024)
Data retention Available with zero data retention
Retirement Not sooner than September 28, 2027

Two smaller details are worth knowing. Setting temperature, top_p or top_k to a non-default value returns a 400 error. And fast mode, Anthropic’s premium-priced faster output option, isn’t offered on Sonnet 5.5. The pricing page lists it only for Opus models.

Where Sonnet 5.5 fits in Anthropic’s lineup

Anthropic’s current generally available lineup, from most to least capable:

  • Claude Fable 5.1 ($10 / $50). For the most demanding reasoning and long-horizon agentic work.
  • Claude Opus 5.5 ($4 / $20, released September 22). Anthropic’s recommended default for most workloads.
  • Claude Sonnet 5.5 ($2 / $10). Described in Anthropic’s docs as “the best combination of speed and intelligence.”
  • Claude Haiku 4.5 ($1 / $5). The fastest and cheapest option, with a 200K context window and a February 2025 knowledge cutoff.

Anthropic says Claude Haiku 5.5 will join the family “in the coming weeks.” Sonnet 5 becomes a legacy model. It stays available, and its $2/$10 price, originally announced as introductory, is now permanent: the increase to $3/$15 that had been scheduled for September 1 was cancelled.

Anthropic’s positioning is modest. Sonnet 5.5 is “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” Opus 5.5 is “built for complex work requiring careful judgment.” The launch post says that in Anthropic’s testing and external testers’, Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment,” even where the benchmarks are close.

Benchmarks: what Anthropic published

Bar chart comparing Claude Sonnet 5.5, Claude Sonnet 5 and Claude Opus 5.5 on Terminal-Bench 4.0, SWE-Bench Pro, CursorBench 4.0, FrontierCode, OSWorld 2.1, Humanity's Last Exam and Chartography
Anthropic’s published scores for Sonnet 5.5 against Sonnet 5 and Opus 5.5. Vendor-reported, not reproduced by Kingy AI.

The headline table from Anthropic’s launch post, with GPT-6 Sol where Anthropic reported it:

Benchmark What it tests Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 Agentic coding in a terminal 70.6% 10.3% 66.4%* —
FrontierCode 1.1 (Main) Mergeable pull requests, graded strictly for scope 46.2% (max) / 52.1% (xhigh) 42.4% 54.4% 49.3%
CursorBench 4.0 Tasks from real Cursor sessions 55.5% 34.1% 57.8% —
GDPval-AA v2.1 Real work across 44 occupations (Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 Long-horizon knowledge work (Elo) 1811 1359 1822 1483
Humanity’s Last Exam (with tools) Expert-level multidisciplinary questions 64.5% 54.9% 67.7% —
OSWorld 2.1 (partial credit) Operating a desktop from screenshots 80.1% 57.0% 81.8% —
Chartography (no tools) Reading data from charts 61.6% 15.6% 64.4% 53.6%

*Opus 5.5 Terminal-Bench score at xhigh effort, its best result. All other Anthropic scores at max effort. GPT-6 Sol figures for GDPval-AA, AA-Briefcase and Chartography may predate an OpenAI fix to an image-understanding bug.

The system card adds a longer list. The most useful additions:

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 Note
SWE-Bench Pro 81.3% 63.2% 89.9% Large multi-file changes in real repos
SWE-Bench Multilingual 90.3% 78.3% 93.9% 300 problems, 9 languages
SWE-Bench Multimodal 54.3% 28.1% 61.4% Issues with screenshots and mockups
Humanity’s Last Exam (no tools) 56.9% 43.1% 64.4% Pure reasoning, no search
AutomationBench (Zapier) 44.7% 10.7% 42.5% GPT-6 Sol: 32.0%
HealthBench Professional 69.2% 57.8% 65.6% Length-adjusted; best of any Claude model
Terminal-Bench-Science 0.1 59.9% — 58.7% GPT-6 Astra: 64.6% (OpenAI-reported)
FrontierSWE v2 (Proximal) 61.9% — 62.3% GPT-6 Astra leads at 65.5%
ProgramBench (long context) 79.7% 77.3% 91.2% Rebuilding programs from binaries, up to 1M tokens
ArXivMath, Aug 2026 (with tools) 95.2% — 96.9% 86.8% vs 91.2% without tools
OSWorld 2.1 (strict pass) 43.5% 25.6% 48.7% Every checkpoint in the task met

Four things stand out.

1. The Terminal-Bench lead over Opus 5.5 is within noise. Anthropic reports standard errors of ±2.5 points for Sonnet 5.5 and ±2.6 for Opus 5.5 on this 66-task benchmark, so a 4.2-point gap isn’t a clear win. But the direction holds up independently: Artificial Analysis’s own run also had Sonnet 5.5 ahead (64% vs 60%). The safe reading is “level with Opus 5.5 on terminal work.” That is still remarkable for a model at half the price.

2. Sonnet 5’s 10.3% on Terminal-Bench 4.0 is real, not a typo. It looks like an error next to Opus 5’s 52.3% on the same benchmark. Artificial Analysis measured Sonnet 5 at 14% in its own run, so the collapse shows up in independent testing too. Neither source explains it. Either way, the improvement from Sonnet 5 on this benchmark is closer to “fixed a failure” than to “improved a strength.”

3. More effort isn’t always better. On FrontierCode, Sonnet 5.5 scored lower at max effort (46.2%) than at xhigh (52.1%). Anthropic’s explanation: at max it more often called Claude Code’s multi-agent code-review skill, which sometimes caused timeouts or edits beyond the task’s scope, and FrontierCode penalises out-of-scope changes. At xhigh, Sonnet 5.5 beats GPT-6 Sol’s 49.3%.

4. On some professional tasks, Sonnet 5.5 beats Opus 5.5. AutomationBench (44.7% vs 42.5%), HealthBench Professional (69.2% vs 65.6%) and Terminal-Bench-Science (59.9% vs 58.7%) all favour the cheaper model. Where it trails clearly is deep reasoning and long-context reconstruction: 7.5 points behind on Humanity’s Last Exam without tools, and 11.5 points behind on ProgramBench.

Anthropic also says Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots, a stress test for long-horizon play and image understanding.

The independent check: Artificial Analysis

Artificial Analysis benchmarked Sonnet 5.5 at max effort on launch day. Its Intelligence Index v4.3.2 combines ten evaluations. Sonnet 5.5 scored 56, ranking #3 of 216 models, two points behind Opus 5.5 and ahead of Fable 5.1 and GPT-6 Astra (both 53). All models below at max effort:

Artificial Analysis eval Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Intelligence Index v4.3.2 56 38 58 48
Terminal-Bench 4.0 64% 14% 60% 44%
AutomationBench-AA 71% 37% 70% 62%
Humanity’s Last Exam 55% 41% 61% 48%
SciCode 61% 54% 67% 58%
CritPt (physics) 31% 17% 32% 31%
AA-LCR v1.1 (long context) 83% 82% 85% 84%
AA-Omniscience (knowledge vs hallucination) 32 16 46 27

The independent numbers largely support Anthropic’s story. Sonnet 5.5 is level with or ahead of Opus 5.5 on agentic execution, including terminal work and SaaS workflow automation. It trails on raw knowledge and hard reasoning. The biggest gap is AA-Omniscience, which rewards correct answers and penalises confident wrong ones: 32 vs 46. Anthropic’s system card says the same: Sonnet 5.5 answers more questions correctly than Sonnet 5 but is “slightly more likely to state an incorrect answer.” If your workload is answering factual questions from the model’s own memory, Opus 5.5 is the safer choice.

One thing worth noting: Artificial Analysis measured 64% on Terminal-Bench 4.0 against Anthropic’s 70.6%. Harness differences are normal, but this is the second Claude 5.5 model where the independent Terminal-Bench number came in lower than the vendor’s.

Pricing and the real cost per task

Model Input / MTok Output / MTok Cache read Batch (in / out) Context
Claude Sonnet 5.5 $2 $10 $0.20 $1 / $5 1M
Claude Sonnet 5 $2 $10 $0.20 $1 / $5 1M
Claude Opus 5.5 $4 $20 $0.20 $2 / $10 1M
Claude Fable 5.1 $10 $50 $0.25 $5 / $25 1M
Claude Haiku 4.5 $1 $5 $0.10 $0.50 / $2.50 200K
GPT-6 Sol $2 $10 $0.20 — 1.05M

Two details in that table matter more than they look:

  • Cache reads are the same $0.20 on Sonnet 5.5 and Opus 5.5. Opus 5.5 charges 5% of its input price for cache hits; Sonnet uses the standard 10%. In a long agent loop where most input tokens are cached, Opus 5.5’s premium is almost entirely on output tokens. The per-token gap is narrower than the “half price” headline suggests.
  • Sonnet 5.5 and GPT-6 Sol have the same list price on input, output and cached input. The comparison comes down entirely to how many tokens each uses to finish a job.

US-only inference (inference_geo: "us") adds a 1.1x multiplier on every token category. It applies to the Claude API, Claude Platform on AWS and Foundry US Data Zone deployments.

Cost per task depends on effort level

Scatter chart of Artificial Analysis Intelligence Index against cost per task at max effort, showing Claude Sonnet 5.5 at 56 and $7.60, Opus 5.5 at 58 and $5.98, and GPT-6 Sol at 48 and $1.06
Intelligence versus cost per task, every model at its max setting. Data: Artificial Analysis, retrieved September 28, 2026.

List prices don’t tell you the cost of a finished job. Tokens used per task do. Artificial Analysis ran every model at max effort. At that setting:

At max effort (Artificial Analysis) Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Cost per Intelligence Index task $7.60 $5.09 $5.98 $1.06
Output tokens per task 193K 118K 119K 31K
Of which reasoning tokens 142K 88K 84K 21K
Cost to run the full index $8,977 $6,998 $8,708 $1,550

At max effort, then, Sonnet 5.5 costs more per task than Opus 5.5, because it thinks for about 60% more tokens. It also costs about seven times more per task than GPT-6 Sol at the same list price. It is the smarter model at that setting (56 vs 48 on the index), but you pay a lot for the difference.

That doesn’t contradict Anthropic’s claim of up to 30% lower cost. It’s a different part of the curve. Anthropic’s launch charts plot every effort level, and its claims are about the lower ones. At Low or Medium effort, Anthropic says Sonnet 5.5 beats Sonnet 5’s best score on several benchmarks “for about a tenth of the cost per task.” On FrontierCode at High effort, it says Sonnet 5.5 scores 10 points higher than Sonnet 5 at about one-fifteenth of the cost per task. Anthropic’s own advice is that Sonnet 5.5 “complements Opus 5.5 best when running at lower effort settings.” At higher settings, “it can perform comparably at a similar cost.”

CursorBench shows how steep the effort curve is for coding:

CursorBench 4.0 Medium High Xhigh Max
Sonnet 5.5 score 39.2% 47.8% 53.1% 55.5%

The jump from Medium to High is worth 8.6 points. The jump from Xhigh to Max is worth 2.4. For health questions, the curve is almost flat: on HealthBench, all five effort levels scored within 0.7 points of each other, while response time went from 8 to 13 seconds (low to xhigh) to about 36 seconds at max.

The practical rule: run Sonnet 5.5 at Medium or High. If a task genuinely needs max effort, compare against Opus 5.5 at a lower setting first, because it may be both cheaper and better.

Early-access customers reported gains at their production settings:

  • Balyasny Asset Management, on 2,441 finance tasks, reported that Sonnet 5.5 scored ahead of Sonnet 5 while using about 121K tokens per answer against Sonnet 5’s 497K. For high-volume workflows, it rated Sonnet 5.5 the best quality-to-cost tradeoff of the seven models it tested.
  • Slack reported better results on almost all of its offline Slackbot evals, with no prompt changes and about 14% fewer output tokens.
  • Box reported higher accuracy, 2.4x the speed and 12% fewer total tokens than the previous model.
  • Base44 reported that across 118 app builds, Sonnet 5.5 matched Opus 5’s quality in 3.6 iterations per build versus Opus 5’s 7.7.
  • Lovable reported a third fewer tool calls and roughly half the shell runs per coding task.

These are customer testimonials published by Anthropic, not independent audits. But they consistently point the same way: fewer steps and fewer tokens at the settings people actually use in production. That fits Artificial Analysis’s June finding that Sonnet 5’s weakness was token use, not intelligence.

Speed

Anthropic says Sonnet 5.5 generates output 30%+ faster than Sonnet 5, “our fastest Sonnet model to date.” Artificial Analysis had not published an output-speed measurement at the time of writing. For reference, it measured Sonnet 5 at 79 tokens per second and Opus 5.5 at 95. Customer-reported speed gains include tickets processed 20% faster at Zendesk and Rovo agents running up to 30% faster at Atlassian. Because Sonnet 5.5 tends to batch tool calls and take fewer steps, wall-clock time per task can improve by more than raw token speed alone would suggest.

Ease of use: what changes when you upgrade

For most chat and document work, switching is a one-line model-ID change. For agentic and tool-using code, Anthropic lists five breaking changes from Sonnet 5:

  1. You can’t disable thinking any more. thinking: {"type": "disabled"} returns a 400 error. The replacement is thinking: {"type": "between_tools"}, which turns off up-front thinking but keeps short progress notes between tool calls. It works only at low, medium or high effort. At xhigh or max you must use adaptive thinking. Manual budget_tokens also returns a 400.
  2. Forced tool use is gone. tool_choice of any or a named tool returns a 400. Use auto with strict: true for schema-valid arguments, or structured outputs, and tell the model in the prompt when to call the tool.
  3. Thinking blocks are bound to the model, the conversation and the account. Sonnet 5.5 can read thinking blocks from Sonnet 5, Opus 4.8 and Haiku 4.5, but no other model can read Sonnet 5.5’s. If you edit earlier history and replay a Sonnet 5.5 thinking block, accounts created on or after August 31, 2026 get a 400 error. Keep conversations append-only, and change instructions with mid-conversation system messages instead of edits. If you move a conversation to a different, unlinked account, its Sonnet 5.5 thinking blocks are silently dropped.
  4. Computer use needs the new toolset. On the Claude API and Google Cloud, the old computer_20251124 tool is rejected. Use computer_toolset_20260801. Bedrock still accepts the old tool.
  5. Advisor-tool pairings changed. Opus 4.8, Opus 4.7 and Sonnet 5 can no longer advise a Sonnet 5.5 executor. Use Opus 5, Opus 5.5, Fable or Sonnet 5.5 itself.

One change fails silently. Text the model writes between tool calls now comes back inside thinking blocks, and those blocks are empty by default. An app that streams those notes to users will simply go quiet between tool calls. Set thinking.display, or use between_tools, to get the text back.

A minimal Python call at the recommended starting point for agentic work:

import anthropic

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=16000,
    output_config={"effort": "medium"},
    messages=[{"role": "user", "content": "Find and fix the failing test in this repo."}],
)
print(response.content[-1].text)

Anthropic’s guidance on effort is to re-run your effort sweep, not carry Sonnet 5 settings over, because the levels have been recalibrated. Start at high for general work. For well-specified agentic coding and multi-step tool use, start at medium and move up for harder tasks. For chat and latency-sensitive work, use medium or low. There are also new conveniences Sonnet 5 lacked: per-message effort changes, mid-conversation system messages and tool changes, compaction on demand, and a smaller tool-use system prompt (286 tokens vs 354).

On the human side, testers describe clearer writing and better design sense than Sonnet 5, including polished UIs and slide decks that follow a template. Anthropic says that in one internal test, two experts judged a first-draft 10-slide operating review built from a public company’s earnings materials ready to send as is. That is a single anecdote, not a benchmark.

Safety and the system card

  • Risk level: Anthropic treats Sonnet 5.5 as meeting its CB-1 and Autonomy-1 thresholds, with the corresponding mitigations, and as below CB-2 and Autonomy-2. It says the model is “broadly less capable than Opus 5.5” and doesn’t advance the capability frontier.
  • Cyber safeguards and fallbacks: The system card says Sonnet 5.5 develops sophisticated exploits “much more capably” than Sonnet 5, though not as capably as Opus 5.5, so it is the first Sonnet to ship with cyber safeguards like those on Anthropic’s top models. Higher-risk cyber requests visibly fall back to Sonnet 5. With the API’s default server-side fallback, cyber and frontier_llm refusals are retried on Sonnet 5. bio, reasoning_extraction and general_harms refusals are not. Anthropic says routine bug-fixing is unaffected. During Terminal-Bench, 1.2% of requests were routed to the fallback model.
  • Anti-distillation: Sonnet 5.5 is the first Sonnet with classifiers that block attempts to extract its hidden reasoning, with no fallback.
  • Prompt injection: in Gray Swan’s adaptive coding-environment attacks, the attack success rate per attempt was 3.01%, compared with 19.47% for Sonnet 5 and 54.61% for Opus 5.5 (all with thinking, before safeguards). By Anthropic’s measure, it is the most injection-resistant Sonnet yet.
  • Alignment: matches or improves on Sonnet 5 on most measures across roughly 1,850 audit scenarios, with the lowest rate of attempted sandbox escapes of any model Anthropic tested. On the downside, its reasoning is “more illegible than many previous models,” its honesty-under-pressure (MASK) rate is below Sonnet 5’s, and multi-turn testing showed regressions in areas such as tracking and surveillance.
  • Welfare: predominantly neutral affect, with a stated preference for difficult, agentic tasks.

Verdict: who should use Claude Sonnet 5.5

Use it if you are on Sonnet 5 today. Same price, far better scores, fewer tokens at normal settings, and it fixes Sonnet 5’s weak spots (terminal work, chart reading, token bloat). Migration mostly means checking the five breaking changes. It is also a strong default for high-volume agentic coding, workflow automation, document and spreadsheet production, and computer use, where it lands within a few points of Opus 5.5.

Choose Opus 5.5 instead for open-ended work that needs sustained judgment, factual recall without search, hard reasoning and very long-context reconstruction. Choose it too for any task where you would otherwise run Sonnet 5.5 at max effort. At that setting Opus 5.5 was cheaper per task in independent testing.

Test GPT-6 Sol if cost per task is your main constraint. At identical list prices, it used about a sixth of the output tokens at max effort. Sonnet 5.5 scored clearly higher on agentic work, so measure on your own tasks.

Wait for Haiku 5.5 if you run very high volumes of simple classification or extraction, since it is due in the coming weeks.

Sonnet 5.5 isn’t the cheapest option or the most capable one. It’s the right choice when you want near-Opus agentic performance and you’re willing to tune effort for it. At Medium or High it lives up to Anthropic’s pitch. At Max, it doesn’t.

Frequently asked questions

How much does Claude Sonnet 5.5 cost?

$2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads cost $0.20 per million tokens and the Batch API halves both rates.

What is the Claude Sonnet 5.5 context window?

1 million tokens, with up to 128K output tokens per request (300K on the Batch API with a beta header).

Is Claude Sonnet 5.5 better than Opus 5.5?

Not overall. It scores slightly higher on Terminal-Bench 4.0, AutomationBench and HealthBench Professional, but Opus 5.5 leads on most benchmarks, on factual reliability and, by Anthropic’s own account, on complex open-ended work.

What is the API model ID?

claude-sonnet-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-sonnet-5-5 on Amazon Bedrock.

Can I turn off thinking on Sonnet 5.5?

Not completely. disabled returns an error. The lowest setting is between_tools, which turns off up-front thinking at high effort or below.

Related reading on Kingy AI

Sources