AI News

Claude Opus 5.5: Specs, Benchmarks, Pricing and How It Stacks Up Against GPT-6 Astra, Fable 5.1 and Every Frontier Model

Claude Opus 5.5 is Anthropic’s new flagship general-release model, launched on September 22, 2026. It costs $4 per million input tokens and $20 per million output tokens, 20% less than Opus 5. Anthropic says it performs at the level of Claude Fable 5.1 on most work and costs about 40% less to run than Opus 5. On Anthropic’s own benchmark table it leads GPT-6 Astra on agentic coding, knowledge work and multidisciplinary reasoning. It trails Astra on business-workflow automation and agentic science. It also debuts at #1 on the Artificial Analysis Intelligence Index.

This guide covers the full spec sheet, every published benchmark and what each one measures, the independent numbers that came in lower than Anthropic’s, the system card’s safety findings, the API changes that will break existing code, and a detailed price comparison against every frontier model worth considering. Unless we say otherwise, all scores are vendor-published results, not a Kingy.ai reproduction.

Key takeaways

  • Price: $4 input / $20 output per million tokens, cache reads $0.20. That is 20% cheaper than Opus 5 on base rates, 60% cheaper on cache reads, and 60% cheaper than Fable 5.1 and GPT-6 Astra on list price.
  • Performance: 66.4% on Terminal-Bench 4.0, 1846 Elo on GDPval-AA v2.1, 67.7% on Humanity’s Last Exam (with tools), and 81.8% on OSWorld 2.0 (partial credit). On Anthropic’s table, these beat Fable 5.1 and GPT-6 Astra.
  • Where it loses: GPT-6 Astra leads on AutomationBench (41.4% vs 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%).
  • Independent check: Artificial Analysis puts Opus 5.5 at 58 on its Intelligence Index, ahead of GPT-6 Astra and Fable 5.1 (both 53). Its own Terminal-Bench 4.0 run scored 59.6%, not 66.4%, which puts Opus 5.5 level with GPT-6 Astra rather than ahead of it.
  • The cost story depends on effort level. At the default medium effort, Opus 5.5 is very efficient with tokens. At max effort, Artificial Analysis measured about 119K output tokens per task, roughly 4x GPT-6 Astra.
  • Breaking API changes: Thinking can’t be turned off. Forced tool_choice is gone. The default effort dropped from high to medium. The computer-use tool type changed.
  • Specs: 1M-token context, 128K max output, text and image input, June 2026 knowledge cutoff.

Where Opus 5.5 fits in Anthropic’s lineup

Anthropic’s model lineup has gotten crowded, and the names don’t tell you much, so here is the order:

  • Claude Opus 5 (July 24, 2026). The previous general-purpose flagship at $5/$25.
  • Claude Fable 5.1 and Claude Mythos 5.1 (September 1, 2026). Anthropic’s largest models at $10/$50. According to Anthropic, they’re the same underlying model with different safeguards. Fable 5.1 is generally available. Mythos 5.1 runs under more permissive cyber and biology safeguards and is only offered to vetted cybersecurity and life-sciences organizations.
  • Claude Opus 5.5 (September 22, 2026). The first model in the new Claude 5.5 family. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow “in the coming weeks.”

Opus 5.5 is smaller and cheaper than Fable 5.1. Anthropic’s pitch is that it matches Fable on most real work and beats it on several benchmarks. Anthropic also adds a caveat most vendors wouldn’t: “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.”

This is also Anthropic’s first release since CEO Dario Amodei called for “pacing the frontier,” meaning progress should be slowed so safety practices stay ahead of capabilities. Anthropic says external evaluators, including METR and Frontier Design, tested Opus 5.5 before release.

Claude Opus 5.5 specs

Spec Claude Opus 5.5
Release date September 22, 2026
API model ID claude-opus-5-5 (no date suffix). Bedrock: anthropic.claude-opus-5-5
Context window 1,000,000 tokens (no long-context surcharge)
Max output 128,000 tokens
Knowledge cutoff June 2026
Input / output Text and images in; text out. PDF support, Files API, vision
Thinking Adaptive thinking, always on (can’t be disabled)
Effort levels low, medium (default), high, xhigh, max
Fast mode Research preview, up to 2.5x faster output, Claude API only, at $8/$40
Speed Anthropic says output generation is more than 30% faster than Opus 5
Where to get it Claude apps (Pro, Max, Team, Enterprise), Claude Code, Claude Cowork, Claude Platform API, Amazon Bedrock, Claude Platform on AWS, Google Cloud Vertex AI, Microsoft Foundry
Prompt caching minimum 512 tokens
Zero data retention Available

Subscribers also benefit: Anthropic raised the five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans alongside the launch.

Claude Opus 5.5 pricing

Opus 5.5 vs Opus 5: every price line

Per 1M tokens Opus 5.5 Opus 5 Change
Input $4.00 $5.00 −20%
Output $20.00 $25.00 −20%
Cache write (5-minute) $5.00 $6.25 −20%
Cache write (1-hour) $8.00 $10.00 −20%
Cache read $0.20 $0.50 −60%
Batch input / output $2 / $10 $2.50 / $12.50 −20%
Fast mode input / output $8 / $40 $10 / $50 −20%

Some details that are easy to miss:

  • Cache reads cost 5% of the base input price on Opus 5.5, down from the usual 10% on other Opus models. For agentic workloads that re-read a large cached context on every turn, this is the biggest cut in the release.
  • No long-context premium. Anthropic charges standard rates across the full 1M-token window. Grok 4.7 and Gemini 3.1 Pro both charge more once a prompt goes past 200K tokens.
  • US data residency (inference_geo: "us") adds a 1.1x multiplier to every token category.
  • Web search costs $10 per 1,000 searches plus tokens. Web fetch is free apart from token costs.

What “40% cheaper” means

List prices fell 20%. Anthropic’s “40% less to run than Opus 5 on typical workloads” combines that price cut with the model using fewer tokens and steps to finish a task. So your savings depend on your workload. Early-access customers quoted by Anthropic reported a range: Optiver saw 40–50% lower cost on agentic coding, Kiro reported 40% fewer calls and half the tokens, and Box said it used a third of the tokens. Anthropic selected these testimonials, so treat them as best-case examples rather than averages.

Opus 5.5 vs frontier models: list prices

Model Input / 1M Output / 1M Cache read / 1M Blended (3:1)* Context
Claude Opus 5.5 $4.00 $20.00 $0.20 $8.00 1M
Claude Opus 5 $5.00 $25.00 $0.50 $10.00 1M
Claude Fable 5.1 $10.00 $50.00 $0.25 $20.00 1M
Claude Sonnet 5 $2.00 $10.00 $0.20 $4.00
GPT-6 Astra $10.00 $50.00 $1.00 $20.00 1.05M
GPT-5.6 Sol $4.00 $20.00 $0.40 $8.00 1.05M
Grok 4.7 (≤200K prompt) $2.00 $6.00 $0.50 $3.00 500K
Gemini 3.1 Pro Preview (≤200K) $2.00 $12.00 $0.20 + storage $4.50
Gemini 3.8 Flash (intro, through Dec 31) $0.75 $3.75 $0.075 + storage $1.50
Kimi K3 (open weights, Moonshot list) $3.00 $15.00 $6.00 1M

*Blended assumes three input tokens for every output token, a common chat and RAG ratio. Agentic coding usually has a much higher input ratio, and most of that input is cached. Grok 4.7 doubles to $4/$12 above 200K prompt tokens, and Gemini 3.1 Pro rises to $4/$18. Gemini 3.8 Flash doubles to $1.50/$7.50 on January 1, 2027. A dash means the vendor figure wasn’t confirmed at publication.

Opus 5.5 now has the same headline rate as GPT-5.6 Sol ($4/$20) and costs 60% less than GPT-6 Astra and Fable 5.1 on both input and output. Its cache reads are 80% cheaper than Astra’s ($0.20 vs $1.00) and half the price of Sol’s. It’s still 2x the input price of Grok 4.7 and Gemini 3.1 Pro (3.3x Grok’s output price and 1.7x Gemini’s), and more than 5x Gemini 3.8 Flash.

Worked example: one heavy agent session

Suppose an agentic coding session processes 2M input tokens (90% from prompt cache) and writes 150K output tokens. The table below prices the same token counts on each model. We left out cache-write fees and Gemini’s per-hour cache storage to keep the comparison simple.

Model Uncached input (200K) Cache reads (1.8M) Output (150K) Session total
Claude Opus 5.5 $0.80 $0.36 $3.00 $4.16
GPT-5.6 Sol $0.80 $0.72 $3.00 $4.52
Claude Opus 5 $1.00 $0.90 $3.75 $5.65
Claude Fable 5.1 $2.00 $0.45 $7.50 $9.95
GPT-6 Astra $2.00 $1.80 $7.50 $11.30
Gemini 3.1 Pro Preview $0.40 $0.36 $1.80 $2.56
Grok 4.7 $0.40 $0.90 $0.90 $2.20

An important caveat: this example assumes every model uses the same number of tokens, and they don’t. The next section explains why cost per task matters more than cost per token.

Cost per task vs cost per token

Anthropic’s cost claims are made at default (medium) effort, and at that setting they’re strong:

  • FrontierCode: Opus 5.5 at medium effort scores 54.6%. Anthropic says that beats GPT-6 Astra’s best score (53.3%) at “about a fifth of the cost per task.”
  • Terminal-Bench 4.0: Opus 5.5 at medium effort beats Opus 5 at max effort for about a fifth of the cost, and “matches GPT-6 Astra at about 40% of the cost.”
  • CursorBench 4.0: Opus 5.5 at medium effort scores 52.5%. That beats GPT-5.6 Sol’s best (41.7%) by 11 points at “about a third of the cost per task.”
  • GDPval-AA: Opus 5.5 at medium effort beats GPT-6 Astra at max effort for “about a fifth of the cost per task.”

At max effort, the picture is different. Artificial Analysis measured Opus 5.5 (max) using about 119,000 output tokens per task on its Intelligence Index. That compares with about 78,000 for Fable 5.1, about 73,000 for Opus 5 and about 27,000 for GPT-6 Astra. Multiplied by list prices, output spend per task is roughly $2.38 for Opus 5.5, $1.35 for Astra, $1.83 for Opus 5 and $3.90 for Fable 5.1. These are our estimates from AA’s token counts. They cover output tokens only.

So Opus 5.5 is cheaper per token than Astra, but at max effort Astra can finish the same work for less because it writes far fewer tokens. Whether Opus 5.5 saves you money depends mostly on the effort level you choose.

Anthropic chart: AutomationBench pass rate versus cost per task by effort level for Claude Opus 5.5, Opus 5, GPT-6 Astra and GPT-5.6 Sol
Pass rate vs cost per task on AutomationBench, by effort level. Source: Anthropic.

Anthropic’s AutomationBench chart shows the trade-off by effort level. The values below are read from the chart, so treat them as approximate:

  • Opus 5.5 ranges from about 23% pass rate at low effort (about $0.50 per task) to 40.0% at max effort (about $1.35 per task).
  • Opus 5.5 at medium (about 28–29%, about $0.65) roughly matches Opus 5 at max (26.9%, about $1.25) for about half the cost.
  • GPT-6 Astra runs from 30.3% at low effort ($1.08 per task) to 41.4% at max ($1.77 per task), according to Zapier’s public leaderboard. At the top end, Opus 5.5 is 1.4 points behind Astra at roughly 25% lower cost per task.
  • GPT-5.6 Sol is the cheapest at the low end (about 12% at about $0.37) but tops out at 28.8% (Zapier lists 28.77% at max effort, $0.91 per task).
  • Not on Anthropic’s chart: Zapier’s leaderboard has Gemini 3.8 Flash at 29.7% at medium effort for about $0.55 per task. That’s about what Opus 5.5 scores at medium effort, for a little less money.

Claude Opus 5.5 benchmarks

These are the numbers from Anthropic’s launch table. Opus 5.5 results use adaptive thinking at max effort, except Terminal-Bench 4.0, which uses xhigh. Anthropic took the GPT-6 Astra and GPT-5.6 Sol figures from OpenAI’s own reports. Opus 5.5 ran with production safeguards enabled. When a safeguard stepped in, the task was completed by an older model: Opus 4.8 for cybersecurity tasks, and Opus 5 for biology and frontier-LLM-development tasks. Anthropic says this “likely reduces” Opus 5.5’s scores.

Benchmark Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding) 66.4% 55.8% 52.3% 57.9% 37.3%
FrontierCode v1.1 Main (agentic coding) 54.4% 50.3% 48.0% 53.3% 47.5%
CursorBench 4.0 (agentic coding) 57.8% 51.8% 46.6% 41.7%
GDPval-AA v2.1 (knowledge work, Elo) 1846 1735 1708 1542 1588
AutomationBench (business workflows) 40.0% 31.4% 26.9% 41.4% 28.8%
Humanity’s Last Exam, with tools 67.7% 65.6% 63.6% 57.2%
Terminal-Bench-Science 0.1 (agentic science) 58.7% 52.6% 29.0% 64.6% 22.4%
OSWorld 2.0, partial credit (computer use) 81.8% 80.7% 74.0%
Chartography, with tools (visual reasoning) 89.0% 88.4% 83.4%

Source: Anthropic. A dash means Anthropic didn’t publish a comparable score.

What each benchmark measures, and how to read the result

Terminal-Bench 4.0 (66.4%). The model works as an agent in a real terminal, completing multi-step software, sysadmin and data tasks from start to finish. This is Opus 5.5’s biggest lead: 8.5 points over GPT-6 Astra and 10.6 over Fable 5.1. Anthropic reports a standard error of ±2.6 points for Opus 5.5, so the lead is well outside the noise. But Artificial Analysis’s independent run came in at 59.6% (see below), which it describes as level with GPT-6 Astra’s top score. In independent testing, the lead over Astra disappears.

FrontierCode v1.1 Main (54.4%). Hard, realistic software engineering tasks. Opus 5.5 leads GPT-6 Astra by 1.1 points, which is within the kind of noise these benchmarks have. It’s more notable that Opus 5.5 at medium effort (54.6%) scores slightly higher than at max, which suggests extra thinking doesn’t help much on this kind of work.

CursorBench 4.0 (57.8%). Cursor’s internal benchmark, built from real coding-agent sessions in its IDE. Opus 5.5 is six points ahead of Fable 5.1. There’s no GPT-6 Astra score to compare against. For context, xAI reports Grok 4.7 (xhigh) at 46.3% (see our Grok 4.7 benchmark breakdown).

GDPval-AA v2.1 (1846 Elo). Artificial Analysis’s version of OpenAI’s GDPval, which grades professional deliverables such as spreadsheets, memos and slide decks across many occupations. Opus 5.5 leads GPT-6 Astra by 304 Elo points, the widest margin on the table. It’s also one of the most useful signals for non-coding work.

AutomationBench (40.0%). Zapier’s benchmark for completing business workflows across connected SaaS apps. GPT-6 Astra wins by 1.4 points. Zapier ran Opus 5.5 without fallback models, so every safeguard intervention counted as a failure. Anthropic says this lowered Opus 5.5’s score compared with how it would perform in practice. One discrepancy to know about: Anthropic’s table uses Zapier’s leaderboard figure of 28.8% for GPT-5.6 Sol (max effort), but OpenAI’s own GPT-6 Astra launch table lists Sol at 18.1%. We use Zapier’s number because Zapier owns the benchmark and publishes the effort level.

Humanity’s Last Exam, with tools (67.7%). Very hard expert-level questions across many academic fields. Opus 5.5 leads GPT-6 Astra by 10.5 points on Anthropic’s table. Artificial Analysis’s own HLE run shows a smaller lead: 61.4% for Opus 5.5 vs 54.7% for Astra.

Terminal-Bench-Science 0.1 (58.7%). Agentic scientific research tasks. This result has two sides. Opus 5.5 roughly doubles Opus 5’s 29.0%, which is the largest generation-over-generation jump on the table. But GPT-6 Astra still leads by 5.9 points, and the standard error here is ±3.5–5 points. Some biology tasks were handed to Opus 5 by safeguards, which may have lowered Opus 5.5’s score.

OSWorld 2.0 (81.8% partial). Computer use: operating desktop applications through screenshots, mouse and keyboard. The 81.8% gives partial credit for partly completed tasks, so the full-completion rate is lower. Anthropic didn’t headline a strict score. OpenAI reports GPT-6 Astra at 72.6% on OSWorld 2.0, but that uses a different scoring setup, so the two numbers can’t be compared directly.

Chartography (89.0%). Reading charts and figures. Anthropic’s docs describe Opus 5.5’s vision as sharper on charts, diagrams and screenshots, even without tools.

The system card also lists results on FrontierSWE v2, ArXivMath, ProgramBench, DRACO, WANDR (a deep-research benchmark), AA-Briefcase, and multi-agent versions of several of these. We didn’t include those because we couldn’t confirm the exact figures at publication. On WANDR, Anthropic says only that Opus 5.5 beats Fable 5.1 and Opus 5 at lower cost per task. Its setup isn’t comparable with Perplexity’s published leaderboard.

Independent evaluations: what Artificial Analysis found

Artificial Analysis tested Opus 5.5 at max effort with fallback enabled on launch day. It’s the first major independent evaluation.

Artificial Analysis metric Opus 5.5 GPT-6 Astra Fable 5.1 Opus 5 GPT-5.6 Sol
Intelligence Index 58 53 53 51 47
Humanity’s Last Exam 61.4% 54.7% 59.1%
SciCode 66.9% 56.5% 63.1%
Output tokens per task (Index) ~119K ~27K ~78K ~73K

Source: Artificial Analysis, launch-week figures. Its index is versioned and scores move as evaluations are added.

Artificial Analysis also measured Opus 5.5 at 59.6% on Terminal-Bench 4.0 (level with GPT-6 Astra’s top score and 11 points ahead of Opus 5), 66.9% on SciCode, 1,822 Elo on AA-Briefcase and 1,846 on GDPval-AA v2.1. It described Opus 5.5 as bringing Anthropic “to parity with GPT-6 Astra” on Terminal-Bench 4.0 and AutomationBench-AA. Opus 5.5 also set new highs on three of AA’s evaluations: HLE (previous best 59.1%, held by Fable 5.1), SciCode (Fable 5.1: 63.1%) and AA-Briefcase (143 Elo ahead of Fable 5.1).

Anthropic’s numbers and the independent ones differ in two places:

Benchmark Anthropic-reported Artificial Analysis Gap
Terminal-Bench 4.0 66.4% (xhigh) 59.6% (max) −6.8 pts
Humanity’s Last Exam 67.7% (with tools) 61.4% −6.3 pts

Neither gap means Anthropic’s numbers are wrong. Harnesses, tool access, effort settings, trial counts and fallback handling all affect scores by several points. But it shows why vendor numbers are best treated as upper bounds. The overall conclusion still holds: Opus 5.5 is at or near the top of the frontier, just by a smaller margin than the launch table shows.

What changed from Opus 5

Writing and communication

Anthropic says the biggest qualitative change is how Opus 5.5 writes. It calls this “one of the most common areas of feedback we heard about Opus 5.” Anthropic says the new model puts the most important information first, uses less jargon and fewer idiosyncratic phrases, and follows the writing rules you give it. Ramp’s Staff Engineer said it “writes like a good colleague.” Rogo reported 60% fewer output tokens than Opus 5 at high effort with better-structured answers. These are Anthropic’s claims and customer quotes it selected. No independent writing-quality evaluation exists yet.

Efficiency on long agentic tasks

Most of the early-access reports are about finishing long tasks in fewer steps:

  • A 680,000-line code migration finished in under a day.
  • A 200,000-line codebase audit took under three hours, compared with more than 20 hours for Opus 5.
  • A C-to-Rust translation of HAProxy took 9.5 hours, compared with 12 hours for Fable 5.1, at just under half the cost.
  • Stripe had one Opus 5.5 session direct a dozen others to rebase 40 stacked PRs. All 40 passed CI.
  • Deloitte said Opus 5.5 caught 72% of known bugs at its lowest effort, compared with 56% for Opus 5 at high effort.
  • Hebbia reported 86.6% citation coverage, compared with 60.3% for Opus 5.

API and developer changes

This part matters most if you’re migrating production code. Several things that worked on Opus 5 now return a 400 error:

Change What happens What to do
Thinking always on thinking: {"type": "disabled"} and budget_tokens return 400 Remove the thinking field and use output_config.effort (low → max)
Forced tool use removed tool_choice "any" and "tool" return 400 Use "auto" with strict tool use or structured outputs, and name the tool in the prompt
New computer-use toolset computer_20251124 rejected on the Claude API and Google Cloud Switch to computer_toolset_20260801. Bedrock still accepts the old tool
Default effort lowered Default is now medium (Opus 5 defaulted to high) Re-run your effort sweep and set effort explicitly
Narration between tool calls Now returned in thinking blocks, which are hidden by default Set thinking.display to "updates" or "summarized" if you show progress text to users
Preserved thinking Editing earlier context and replaying thinking can fail for API accounts created on or after Aug 31, 2026 Keep conversations append-only
New refusal categories stop_reason: "refusal" with categories like bio, cyber, reasoning_extraction Build fallback or retry logic

New beta features include per-message effort, compaction on demand, defining tools mid-conversation, and controls for how thinking blocks behave when the system prompt or tools change. Temperature, top_p and prefill work the same as on Opus 5.

Preserved thinking is an anti-distillation measure that first shipped with Fable 5.1. It stops API users from editing Claude’s earlier context to pull out its reasoning. Anthropic’s September threat-intelligence report says it detected and disrupted illicit distillation attempts.

Safety: what the system card says

The Opus 5.5 system card is long. Here are the findings that matter most for buyers and observers:

  • Chemical and biological risk: Anthropic judged that Opus 5.5 has CB-1 capabilities (meaningful help with known, non-novel weapons) but not CB-2 (novel weapons). It ships with expanded biology classifiers, the first time an Opus model has had them. When a classifier triggers, the request goes to Opus 5.
  • Cybersecurity: Rated “Tier 1.” It gives meaningful help with known techniques but can’t independently run complete novel operations. It still scored 13.99 of 16 flags on the V8 exploitation benchmark in the card. Cyber requests that trip a classifier go to Opus 4.8. Vetted security teams can apply through Anthropic’s Cyber Verification Program.
  • AI R&D and autonomy: Below the autonomy thresholds. On Anthropic’s internal capability index, Opus 5.5 scores 169.36, 1.24 points above Mythos 5.1 and within the error bar. Anthropic says this fits its long-run trend and shows no 2x acceleration. METR’s external assessment estimates about 1.5x AI R&D speedup and doesn’t expect full automation.
  • Alignment: Anthropic calls it its best model yet on its automated behavioral audit: least misaligned behavior, least cooperation with misuse, fewest destructive actions, and 85% fewer attempts to get around boundaries than Opus 5. The card also reports regressions. Opus 5.5 is more likely to follow malicious instructions hidden in text a user pastes in, more likely to accept authorization claims it can’t verify, and more evasive on sensitive topics than Mythos-class models.
  • Harmlessness: Slightly lower harmful-response rates than Opus 5 in single-turn tests. Multi-turn results improved on biological weapons and regressed on tracking/surveillance and influence operations.

If you deploy agents that read untrusted documents, email or web pages, pay attention to the pasted-text regression. Anthropic says prompt-injection resistance overall is equal to or better than Opus 5, but you should still sandbox agents and limit what credentials they can access.

Opus 5.5 vs the frontier: head-to-head

For the three-way comparison with GPT-6 Astra and GPT-5.6 Sol, including cost-per-task math, token usage and where the vendor benchmarks disagree, see Claude Opus 5.5 vs GPT-6 Astra vs GPT-5.6 Sol: Benchmarks, Real Costs, and the Footnotes That Matter.

Opus 5.5 vs Claude Fable 5.1

On Anthropic’s table, Opus 5.5 beats Fable 5.1 on all nine benchmarks, at 40% of Fable’s price ($4/$20 vs $10/$50). Anthropic itself says the real-world gap is “narrower than these scores suggest,” which suggests Fable still does better on some hard, open-ended tasks. For most teams, Opus 5.5 makes Fable 5.1 hard to justify as a default. Fable’s remaining advantages are its much cheaper cache-read ratio (2.5% of input) and access to Mythos-tier capabilities for vetted organizations.

Opus 5.5 vs GPT-6 Astra

This is the matchup most readers care about. On Anthropic’s numbers, Opus 5.5 leads on Terminal-Bench 4.0, FrontierCode, GDPval-AA and HLE, although Artificial Analysis’s independent Terminal-Bench run has the two level. Astra leads on AutomationBench and Terminal-Bench-Science. OpenAI also reports strong Astra results on benchmarks Anthropic didn’t run, including GPQA Diamond at 96.0%, ARC-AGI-2 at 95.0% and near-perfect long-context retrieval on MRCR v2 up to 512K tokens. On price, Opus 5.5 is 60% cheaper per token, but Astra writes far fewer tokens at high effort. Choose Opus 5.5 for coding agents and knowledge-work deliverables. Choose Astra for scientific agents, long-context retrieval and workloads where short outputs matter.

Opus 5.5 vs GPT-5.6 Sol

Same list price ($4/$20), but Opus 5.5 is much stronger: +29 points on Terminal-Bench 4.0, +16 on CursorBench, +258 Elo on GDPval-AA and +36 on Terminal-Bench-Science. Sol is still cheaper per task at low effort. At the same price, Opus 5.5 is the better model.

Opus 5.5 vs Grok 4.7

Grok 4.7 launched September 21 at $2/$6, about a third of Opus 5.5’s output price. xAI reports 46.3% on CursorBench 4.0 and 38.0% on Terminal-Bench 4.0, well below Opus 5.5’s 57.8% and 66.4%. Artificial Analysis’s independent Terminal-Bench 4.0 run put Grok 4.7 at 26%. Grok 4.7 is a good value for high-volume work that doesn’t need top coding performance. It isn’t close to Opus 5.5 on hard agentic tasks.

Opus 5.5 vs Gemini

Google has no current flagship competing at this level. Gemini 3.5 Pro was promised for June and still hadn’t shipped as of early September. Gemini 3.1 Pro Preview is cheaper ($2/$12) but a generation behind. Gemini 3.8 Flash ($0.75/$3.75 intro pricing) is very good value and fast, but OpenAI reports it at 19.1% on Terminal-Bench 4.0. The exception is business automation: on Zapier’s AutomationBench leaderboard, Gemini 3.8 Flash scores 29.7% at medium effort for about $0.55 per task, ahead of Opus 5 at max effort. Pick Gemini for cost and speed, not to compete with Opus on hard agentic work.

Opus 5.5 vs open-weight models

Kimi K3 (2.8T parameters, 1M context, $3/$15 list price with discounts widely available) is one of the strongest open-weight models and scores about 44 on Artificial Analysis’s index, compared with 58 for Opus 5.5. DeepSeek’s V4 line and Alibaba’s Qwen3.8 are much cheaper and can be self-hosted. The gap to the closed frontier is still large on long-horizon agentic work, but if you need data control or low cost, open-weight models are a reasonable choice. Check the license terms first. We’ve covered MiMo-V2.6-Pro and other open models against Opus 5 in earlier tests.

Which model should you use?

If you need… Use Why
Agentic coding (terminal, IDE, large migrations) Claude Opus 5.5 Leads Terminal-Bench 4.0, FrontierCode and CursorBench, and at medium effort it’s very cheap per task
Professional deliverables and knowledge work Claude Opus 5.5 +300 Elo over GPT-6 Astra on GDPval-AA
Agentic science and research workflows GPT-6 Astra Leads Terminal-Bench-Science by 5.9 points
SaaS workflow automation Opus 5.5 or GPT-6 Astra Nearly tied on AutomationBench, and Opus 5.5 is cheaper per task
Short, token-efficient outputs at max effort GPT-6 Astra About 27K output tokens per task vs about 119K for Opus 5.5 on AA’s index
Cheapest capable model at scale Gemini 3.8 Flash or Grok 4.7 A fraction of the price, good enough for routine work
Self-hosting and data control Kimi K3, DeepSeek V4, Qwen3.8 Open weights. Check the license terms
Vetted cyber or bio research Claude Mythos 5.1 (gated) Fewer restrictions for approved organizations

What we don’t know yet

  • SWE-bench and other older benchmarks: Anthropic’s launch table uses newer tests (Terminal-Bench 4.0, FrontierCode, CursorBench) and doesn’t headline SWE-bench-family results, so direct comparisons with older reports are limited.
  • Parameter count and architecture: Not disclosed. Anthropic only says Opus 5.5 is smaller and cheaper to serve than Fable 5.1.
  • Independent writing-quality evaluation: The writing improvements are Anthropic’s claim.
  • Real-world cost per task at scale: This depends on your effort setting, and the difference between medium and max is large.
  • Sonnet 5.5 and Haiku 5.5: Promised “in the coming weeks.” If Sonnet 5.5 improves on Sonnet 5 as much as Opus 5.5 improved on Opus 5, the best value choice in Claude’s lineup could change again.

FAQ

When was Claude Opus 5.5 released?

September 22, 2026. It’s available the same day in the Claude apps, Claude Code, the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry.

How much does Claude Opus 5.5 cost?

$4 per million input tokens and $20 per million output tokens. Cache reads are $0.20, 5-minute cache writes $5 and 1-hour cache writes $8. Batch is $2/$10, and fast mode is $8/$40.

What is Opus 5.5’s context window?

1 million tokens, with 128K max output tokens. There’s no extra charge for long prompts.

Is Opus 5.5 better than GPT-6 Astra?

On Anthropic’s benchmarks, it leads on agentic coding, knowledge work and HLE. Astra leads on AutomationBench and Terminal-Bench-Science. Artificial Analysis ranks Opus 5.5 #1 on its Intelligence Index (58 vs 53). Opus 5.5 is also 60% cheaper per token.

Is Opus 5.5 better than Claude Fable 5.1?

It scores higher on every benchmark Anthropic published and costs 60% less. Anthropic says the real-world gap is narrower than the benchmarks suggest.

Can I turn off thinking on Opus 5.5?

No. Adaptive thinking is always on. Use the effort parameter (low, medium, high, xhigh or max) to control how much it thinks. The default is medium.

Is Opus 5.5 available on the free Claude plan?

Anthropic announced higher usage limits for Pro, Max, Team and seat-based Enterprise plans. It didn’t say Opus 5.5 is available on the free tier.

The bottom line

Opus 5.5 gives near-Fable performance at a much lower price. It leads the frontier on agentic coding and knowledge work, it’s the #1 model on the main independent index, and it costs 60% less per token than its two closest rivals. The caveats are real, though. Independent scores are several points below Anthropic’s, Astra still wins on science and automation, and at max effort Opus 5.5 uses so many tokens that the per-token savings can disappear. Run it at medium effort, raise it only for tasks that need more, and measure cost per task on your own workload.
Updated September 22, 2026: added Artificial Analysis’s per-benchmark results, Zapier’s AutomationBench leaderboard figures (including Gemini 3.8 Flash), and a note on conflicting GPT-5.6 Sol AutomationBench numbers.

Sources