Evidence checked September 28, 2026. Prices are first-party standard API list rates in USD. Benchmark figures come from Anthropic’s and OpenAI’s launch materials and from Artificial Analysis. This is a source-led comparison, not a Kingy head-to-head test, and each vendor chose which evaluations to publish.
Anthropic released Claude Sonnet 5.5 today, six days after OpenAI shipped GPT-6 Sol. The two are aimed at the same buyer: teams that want near-flagship coding and agent performance without paying flagship rates. On the rate card they cost the same, $2 per million input tokens and $10 per million output tokens, and both charge $0.20 for a cached read.
That makes the rate card the least useful part of the comparison. The real differences are in what each model scores, how many tokens it burns to get there, what happens to your bill past 272,000 tokens of input, and which benchmarks each vendor chose not to publish.
The short version: Sonnet 5.5 has the higher ceiling. On Artificial Analysis’ independent Intelligence Index it reaches 56 at max effort, against 48 for GPT-6 Sol. It also leads clearly on long-horizon knowledge work. GPT-6 Sol is much cheaper per finished task at every effort level the index measures, and it posts the better FrontierCode score in the only coding benchmark both vendors report. The two lines cross at about $1 per task. Below that, Sol gives you more for your money. Above it, Sonnet is the only one of the two that keeps climbing.
Claude Sonnet 5.5 vs GPT-6 Sol at a glance
| Item | Claude Sonnet 5.5 | GPT-6 Sol |
|---|---|---|
| Released | September 28, 2026 | September 22, 2026 |
| API model ID | claude-sonnet-5-5 |
gpt-6-sol |
| Position in family | Middle tier, below Opus 5.5 and Fable 5.1 | Middle tier, below GPT-6 Astra, above Luna |
| Input / output price | $2 / $10 per million tokens | $2 / $10 per million tokens |
| Context window | 1M tokens (Artificial Analysis listing)* | 1,050,000 tokens |
| Max output | Not yet published for 5.5 (Sonnet 5: 128K)* | 128,000 tokens |
| Knowledge cutoff | Not yet published* | April 20, 2026 |
| Modalities | Text and image in; text out | Text and image in; text out (no audio or video) |
| Reasoning effort settings | Low, Medium, High, Xhigh, Max | none, low, medium, high, xhigh, max |
| Default effort | Medium in Claude apps and Claude Code; High on the Claude Platform | medium |
| Where to get it | Claude apps, Claude Platform, AWS, Google Cloud, Microsoft Azure | ChatGPT Work and Codex (paid plans), OpenAI API, Azure |
*Anthropic’s launch post does not list the context window, output limit or cutoff, and its Sonnet 5.5 documentation page was not yet live when this was published. Artificial Analysis lists a 1M-token window. Sonnet 5, the model it replaces, has a 1M window, 128K max output (300K via the Batch API beta) and a January 2026 cutoff according to Anthropic’s Sonnet 5 page. Confirm the 5.5 limits in Anthropic’s docs before you build around them. GPT-6 Sol’s specs come from OpenAI’s model page.
Pricing: identical rate cards, different bills
Anthropic kept Sonnet 5.5 at Sonnet 5’s price. OpenAI cut Sol to half of GPT-5.6 Sol’s promotional rate, which was $4 input and $20 output. The result is a near-perfect match on the base rates:
| Per million tokens | Claude Sonnet 5.5 | GPT-6 Sol |
|---|---|---|
| Input | $2.00 | $2.00 |
| Output | $10.00 | $10.00 |
| Cached input read | $0.20 | $0.20 |
| Cache write | $2.50 (5-minute); 1-hour writes bill at 2× input under Anthropic’s standard multiplier | $2.50 (1.25× input) |
| Batch discount | 50% off input and output | 50% off (Batch and Flex) |
| Very long prompts | Standard rate across the full window | Above 272K input tokens: 2× input and cache rates, 1.5× output, for the whole request |
| Fast mode | Not offered on Sonnet (Opus only) | Available at 2× the applicable rates |
| Data residency | 1.1× for US-only inference | +10% for regional processing |
Sources: Anthropic’s Sonnet 5.5 announcement, Anthropic’s pricing docs, and OpenAI’s GPT-6 Sol model page. Anthropic’s pricing docs say Claude models from 4.6 onward bill the full context window at the standard rate. We are applying that rule to Sonnet 5.5, but its own pricing row had not been published yet.
Where the bills actually diverge
1. Long prompts. This is the biggest structural difference. Once a Sol request carries more than 272,000 input tokens, the entire request reprices. Take a 500,000-token document review that produces 10,000 tokens of output:
- Sonnet 5.5: 500K × $2 = $1.00, plus 10K × $10 = $0.10. Total $1.10.
- GPT-6 Sol: 500K × $4 = $2.00, plus 10K × $15 = $0.15. Total $2.15.
That is nearly double for the same job, before either model has answered a question. If your workload regularly loads whole repositories, long contracts or large transcripts into one call, this line matters more than any benchmark below.
2. Short, cached agent turns. Here the two are the same. A typical coding-agent turn with 50K cached tokens, 10K fresh input and 4K output costs about $0.07 on either model ($0.01 cached + $0.02 fresh + $0.04 output).
3. Tokens per task. This is the difference you cannot read off a price sheet. Both vendors now pitch cost per task rather than cost per token. Anthropic says Sonnet 5.5 costs “up to 30% less per task” than Sonnet 5 because it needs fewer tokens. OpenAI prices Sol’s AutomationBench run at $0.27 per task at xhigh effort. The independent data in the next section shows that at matched effort labels, Sonnet 5.5 spends considerably more per task than Sol.
The independent view: Artificial Analysis Intelligence Index
Vendor benchmark tables are curated. The most useful neutral cross-check available on launch day is Artificial Analysis, which runs both models through the same ten-evaluation suite: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. It also records what each run cost at list price.
| Effort | Sonnet 5.5 score | Sonnet 5.5 cost/task | GPT-6 Sol score | GPT-6 Sol cost/task |
|---|---|---|---|---|
| Low | 36 | $0.41 | 34 | $0.13 |
| Medium | 41 | $0.59 | 40 | $0.25 |
| High | 47 | $1.08 | 43 | $0.37 |
| Xhigh | 52 | $2.74 | 44 | $0.53 |
| Max | 56 | $7.60 | 48 | $1.06 |
Source: Artificial Analysis release pages for Claude Sonnet 5.5 and GPT-6 Sol, Intelligence Index v4.3.2, read September 28, 2026. AA also lists a non-reasoning Sol setting that scores 28.
Three things stand out.
Sonnet scores higher at every matched setting, and the gap widens with effort. At low and medium the two are within two points. At max, Sonnet 5.5 leads by eight. On this index its max score of 56 is also above the highest GPT-6 Astra setting AA lists (53), and above Claude Fable 5.1 (53). Only Opus 5.5 (58) is higher. For context, Sonnet 5 topped out at 38.
Sol is cheaper per task at every matched setting, and that gap also widens. Sonnet costs 2.4× more per task at medium, 2.9× at high, 5.2× at xhigh and 7.2× at max. With identical rate cards, the difference comes almost entirely from how many tokens each model spends. The effort labels share names but are not calibrated to each other, so “high” on one model is not the same amount of work as “high” on the other.
At the same spend, the two are close. The fairer comparison is by cost, not by label. At roughly $1 per task, Sonnet 5.5 at high (47, $1.08) and GPT-6 Sol at max (48, $1.06) are effectively tied. Below $1, Sol delivers more per dollar: Sol at medium scores 40 for $0.25, while Sonnet’s cheapest setting costs $0.41 for 36. Above $1, Sol has nowhere left to go, and Sonnet keeps climbing to 52 and then 56.
One label worth knowing: AA lists every Sonnet 5.5 configuration as “Default Fallback.” That refers to Anthropic’s new cyber safeguard, under which higher-risk cybersecurity requests visibly fall back to Sonnet 5. AA ran the model the way customers will actually receive it.
Head-to-head: the benchmarks both sides report
Anthropic’s launch table includes a GPT-6 Sol column, which is unusual and useful. OpenAI’s Sol launch compares against Claude Opus 5 and Fable 5.1 but not Sonnet 5.5, which didn’t exist yet. These are the only evaluations with published scores for both models:
| Benchmark | What it tests | Sonnet 5.5 | GPT-6 Sol | Leader |
|---|---|---|---|---|
| FrontierCode 1.1 (Main) | Code changes that are correct and mergeable without human edits | 46.2% (Max) | 49.3%; 52.1% at xhigh | GPT-6 Sol |
| GDPval-AA v2.1 | Real-world work tasks across 44 occupations (Elo) | 1844 | 1487 | Sonnet 5.5, by 357 |
| AA-Briefcase v1.1 | Long-horizon knowledge work (Elo) | 1811 | 1483 | Sonnet 5.5, by 328 |
| Chartography (no tools) | Reading data off charts | 61.6% | 53.6% | Sonnet 5.5 |
Source: Anthropic’s Sonnet 5.5 announcement. GDPval-AA and AA-Briefcase were run by Artificial Analysis. Chartography scores come from Surge AI.
Read the footnotes before quoting these, because both sides carry a bug disclosure:
- Sonnet 5.5: Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a structured-outputs bug. Anthropic says the effect, if any, should be small and would understate Sonnet’s scores.
- GPT-6 Sol: OpenAI recently fixed a bug that degraded Sol’s image understanding, and the third-party scores may predate the fix. Artificial Analysis does not expect major changes to the Elo scores, and Anthropic says its internal Chartography testing showed no impact.
- FrontierCode effort: Sonnet 5.5 actually scores lower at Max than at Xhigh. Anthropic says that at Max it more often launched Claude Code’s multi-agent code-review skill, which caused a timeout or out-of-scope edits in cases Cognition examined. Out-of-scope edits count against mergeability. Anthropic did not publish the higher Xhigh figure.
Even with those caveats, a 300-plus Elo gap on two separate knowledge-work suites is too large to explain away. On FrontierCode, Sol’s lead is modest but real in the numbers Anthropic itself published.
Coding: each vendor’s best numbers
Beyond FrontierCode, each vendor reports coding benchmarks the other doesn’t:
| Benchmark | Sonnet 5.5 | GPT-6 Sol | Notes |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | Not reported | Sonnet 5 scored 10.3%; Opus 5.5 scored 66.4% at Xhigh |
| CursorBench 4.0 | 55.5% | Not reported | Opus 5.5: 57.8% |
| DeepSWE v1.1 | Not reported | 68.8% (max) | OpenAI: within 1.1 points of Fable 5’s 69.9% |
The Terminal-Bench 4.0 result is Sonnet 5.5’s headline. It went from 10.3% to 70.6% in one generation and now scores above Opus 5.5. Anthropic notes that neither Terminal-Bench nor OpenAI had published a GPT-6 Sol score, so its cost chart uses GPT-5.6 Sol instead. Terminal-Bench 4.0 is one of the ten evaluations inside the Artificial Analysis index, so it is part of the head-to-head above, but AA’s launch pages do not break it out.
OpenAI’s coding case rests on DeepSWE and on cost. It says Sol gets within 1.1 points of Claude Fable 5’s best DeepSWE score at about 80% lower cost per task. The practical reading: if your coding work is terminal-driven and agentic, Sonnet 5.5 has the stronger published evidence. If you care most about clean, mergeable pull requests at a low per-task cost, Sol’s FrontierCode and DeepSWE results make the better case.
Anthropic’s early-access quotes add useful signals about efficiency. Take them as vendor-selected. Lovable reported about a third fewer tool calls and roughly half the shell runs per task compared with Sonnet 5. Base44 reported 3.6 iterations per app build against Opus 5’s 7.7, at a similar quality score.
Computer use: two OSWorld numbers that aren’t comparable
You will see these two figures side by side in launch coverage:
- Sonnet 5.5: 80.1% on OSWorld 2.1 (partial reward), per Anthropic.
- GPT-6 Sol: 60.5% on OSWorld 2.0 offline (partial reward, v2026.08.08 release) at xhigh effort, per OpenAI.
Do not compare them. They are different versions of the benchmark, and each vendor ran its own model in its own harness. The honest comparisons are within each vendor. Sonnet 5.5 is 1.7 points behind Opus 5.5 (81.8%) and 23 points ahead of Sonnet 5 (57.0%). OpenAI says Sol at xhigh roughly matches Claude Opus 5 at medium effort (60.3%) for about 80% less per task, and that GPT-6 Astra remains its best computer-use model.
Anthropic’s Pokémon Red test is anecdotal but informative. Sonnet 5.5 is the first Sonnet to finish the game working only from screenshots, which demonstrates sustained visual control over a very long session.
Knowledge work, reasoning and factuality
Knowledge work is where Sonnet 5.5’s case is strongest. Beyond the GDPval-AA and AA-Briefcase leads above, Anthropic reports:
- Humanity’s Last Exam (with tools): 64.5%, up from Sonnet 5’s 54.9% and close to Opus 5.5’s 67.7%. OpenAI did not report a Sol figure.
- GDPval-AA 1844 vs Opus 5.5’s 1846. On that suite, Sonnet 5.5 is effectively level with Anthropic’s more expensive model, which costs $4/$20.
OpenAI’s knowledge-work evidence for Sol uses different benchmarks:
- AutomationBench 1.0.6: 33.2% at xhigh for $0.27 per task. That beats Claude Opus 5 at max (26.9%, at 11.1× Sol’s cost) and low-effort GPT-6 Astra (30.3%).
- Agents’ Last Exam V1: 56.4% at max, which OpenAI says beats Opus 5’s best score at 60% lower cost per task.
- Factuality: on OpenAI’s internal evaluation, built from conversations where users flagged factual errors, Sol makes about half as many mistakes as GPT-5.6 Sol. The test set is deliberately error-prone and internal, so no outside number can be checked against it.
Sonnet 5.5 has not been reported on AutomationBench or Agents’ Last Exam, and Sol has not been reported on HLE. None of these results can be compared directly. The only cross-vendor knowledge-work numbers are the Artificial Analysis Elo scores, and there Sonnet leads by a wide margin.
Speed and latency
Anthropic says Sonnet 5.5 generates output 30%+ faster than Sonnet 5, its fastest Sonnet yet. Customer quotes point the same way: Box reported 2.4× faster results and Zendesk reported tickets processed 20% faster. Artificial Analysis had not published measured output speed for Sonnet 5.5 at the time of writing.
For GPT-6 Sol, Artificial Analysis measures 79–90 output tokens per second depending on effort (90 t/s at max, 83 t/s at low, high and xhigh) and 0.95 seconds to first answer token in non-reasoning mode. OpenAI also offers a fast mode at double the rate. Anthropic reserves fast mode for Opus.
Raw tokens per second is only part of the picture, because the model that uses fewer tokens finishes first. On the Artificial Analysis cost data, Sol uses far fewer tokens per task at every matched effort, which should translate into shorter wall-clock time as well. We will update this section once measured Sonnet 5.5 throughput is available.
Safety, safeguards and behaviour changes
Sonnet 5.5 is the first Sonnet to launch with the cyber safeguards Anthropic built for its top models, because its cyber capability is comparable to Opus 5’s. Routine bug-finding and fixing are unaffected, but higher-risk security tasks fall back to Sonnet 5. Verified security teams can apply to an expanded Cyber Verification Program. It is also the first Sonnet with classifiers that block reasoning extraction (distillation), and it expands “preserved thinking,” which ties the model’s reasoning to the account that created it. Two migration details for developers:
- If you run Sonnet with thinking off, you need to switch to the new
between_toolssetting before moving to 5.5. - Moving conversations between accounts, including switching accounts mid-session in Claude Code, now behaves differently because of preserved thinking.
On Anthropic’s automated behavioural audit of roughly 1,850 scenarios, Sonnet 5.5 matches or improves on Sonnet 5 on most measures. It also comes close to Opus 5.5 on sandbox-escape containment tests. Zero data retention is available.
GPT-6 Sol inherits the alignment training OpenAI introduced with Astra. In OpenAI’s alignment evaluations it improves on GPT-5.6 Sol, including lower rates of misleading claims about its own coding work. OpenAI also changed Sol’s writing style to use less jargon and give slightly shorter answers, which will be noticeable in coding conversations. Full results are in OpenAI’s system card.
Availability and access
- Sonnet 5.5 is available across Anthropic’s platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, as
claude-sonnet-5-5. It defaults to Medium effort in Claude apps and Claude Code and to High on the Claude Platform. Claude Haiku 5.5 is due in the coming weeks. - GPT-6 Sol is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, but not yet in standard ChatGPT chat. In the API it is
gpt-6-sol. OpenAI’s published rate limits run from 500 RPM and 500K TPM at Tier 1 to 15,000 RPM and 40M TPM at Tier 5. Sol supports OpenAI’s full Responses API toolset (web search, file search, code interpreter, hosted shell, apply patch, computer use, MCP and tool search) plus structured outputs. It does not support fine-tuning.
What the benchmarks don’t tell you
- Each vendor picked its own benchmarks. Anthropic highlights Terminal-Bench, CursorBench and HLE. OpenAI highlights AutomationBench, Agents’ Last Exam and DeepSWE. Neither covers the other’s choices.
- Effort labels don’t match across vendors. Compare by cost per task, not by setting name.
- Both models just launched. Both have disclosed deployment bugs affecting third-party scores, and Sonnet 5.5 is only a few hours old. Expect independent numbers to move.
- Your prompts are not a benchmark. Anthropic itself says Opus 5.5 remains clearly stronger at open-ended work that needs sustained judgment, even where Sonnet 5.5 matches it on benchmarks.
Which should you choose?
Choose Claude Sonnet 5.5 if:
- You need the highest ceiling in this price tier. Its max-effort Intelligence Index score is 8 points above Sol’s and above GPT-6 Astra’s listed best.
- Your work is long-horizon knowledge work such as documents, analysis, spreadsheets or slide decks, where it leads Sol by 300+ Elo on both GDPval-AA and AA-Briefcase.
- You run terminal-driven coding agents. Its 70.6% on Terminal-Bench 4.0 is the strongest published agentic-coding number in this matchup.
- You routinely send prompts above 272K tokens. You avoid Sol’s surcharge, which can nearly double a long-document bill.
Choose GPT-6 Sol if:
- You run high volumes where cost per task is what matters. Artificial Analysis measured it 2.4–7.2× cheaper per task than Sonnet 5.5 at matched effort, at scores within 1–4 points up to high effort.
- Your definition of done is a mergeable pull request. It leads Sonnet 5.5 on FrontierCode in Anthropic’s own table.
- You need a cheap floor. Sol at low effort costs $0.13 per task, about a third of Sonnet’s cheapest setting.
- You are already on OpenAI’s Responses API toolset or Codex.
For many teams the answer is both. The cost curves cross near $1 per task. A reasonable default is to route routine, high-volume work to Sol at low or medium effort and escalate the hard long-horizon jobs to Sonnet 5.5 at high or above. Measure successful tasks per dollar on your own workload, not tokens per dollar on a rate card. On the rate card, these two models are identical.
More context from Kingy AI: GPT-6 Sol vs Claude Sonnet 5 (the model Sonnet 5.5 replaces), GPT-6 Sol vs Claude Opus 5.5, GPT-6 Sol and Luna specs, Claude Opus 5.5 specs, and OpenAI vs Anthropic API pricing in 2026.
Sources
- Anthropic: Introducing Claude Sonnet 5.5 (September 28, 2026): benchmarks, footnotes, pricing, safeguards
- Anthropic: Claude pricing documentation: cache, batch, long-context and data-residency rules
- Anthropic: Claude Sonnet 5 model page: predecessor specs
- OpenAI: Introducing GPT-6 Sol and Luna (September 22, 2026): benchmarks, pricing, availability
- OpenAI: GPT-6 Sol model documentation: context, output, cutoff, pricing modifiers, rate limits
- Artificial Analysis: Claude Sonnet 5.5 and GPT-6 Sol: Intelligence Index v4.3.2, cost per task, speed
Trending on Kingy
Keep reading with the stories getting the most attention now.
