Evidence checked September 22, 2026. Prices are direct API rates in US dollars. This comparison uses published results; Kingy did not run paid benchmark tests.
Claude Opus 5.5 at medium or high effort is the first configuration I would evaluate for cost-sensitive coding. GPT-6 Astra deserves a place alongside it for business workflows and difficult work where its particular strengths matter. The published curves support those starting choices. They also show why setting everything to max can waste money.
On Anthropic’s Terminal-Bench 4.0 chart, Opus 5.5 medium scores 57.6% at $2.94 per attempt. Astra high scores 57.9% at $7.21. That is a 0.3-percentage-point gap, with the Opus point costing $2.94 ÷ $7.21 = 40.8% as much. These are published configurations, not a controlled Kingy comparison. Anthropic attributes Astra’s results to OpenAI. Source: Anthropic’s launch charts and methodology.
The useful choice is a model plus an effort setting plus a task. This article follows that choice through coding, professional work and a bill you can calculate yourself. For the wider launch overview, see our Opus 5.5 specs and benchmarks comparison.
What low, medium, high, xhigh and max mean
Astra’s API supports low, medium, high, xhigh and max. Opus 5.5 exposes the same five names, with medium as its default. Neither offers an ordinary thinking-off mode: Astra rejects reasoning effort none, and Opus 5.5 keeps adaptive thinking on. OpenAI’s reasoning guide, Anthropic’s effort documentation, Opus 5.5 changes.
Treat effort as a control over behavior and token use. It is not a fixed number of seconds or a shared compute budget. Medium on one model need not resemble medium on another. Anthropic also says Opus 5.5 tends to think more per turn than Opus 5 at the same setting, especially at xhigh and max. Carrying an old preference into the new model can change the bill. Opus 5.5 behavior changes.
Keep reasoning effort separate from processing speed. Astra Fast mode doubles applicable rates; Opus 5.5 Fast mode lists $8 input and $40 output per million tokens. Those are separate choices from asking either model to reason harder. The calculations below use Standard processing. Astra model reference, Opus 5.5 launch pricing.
Coding across all five effort settings
The following table transcribes the point labels in Anthropic’s Terminal-Bench 4.0 chart. Each cell gives the score and published cost per attempt. Astra’s series is attributed to OpenAI; putting the series on one chart does not make the underlying prompts, tools or evaluation setup identical.
| Effort | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|
| Low | 38.5% / $1.29 | 49.7% / $4.95 |
| Medium | 57.6% / $2.94 | 53.9% / $6.15 |
| High | 64.2% / $3.88 | 57.9% / $7.21 |
| Xhigh | 66.4% / $7.35 | 57.6% / $7.48 |
| Max | 64.8% / $11.24 | 56.7% / $10.35 |
Source: Anthropic, Terminal-Bench 4.0 chart.
At low effort, Astra has the higher score and the higher cost. From medium onward, the displayed Opus scores are higher at every matching label. At max, however, Opus also costs more than Astra. That is enough to reject a blanket claim that either model is always the cheaper or stronger option.
The most useful comparison is inside each model’s curve:
- Moving Opus from medium to high adds 64.2 − 57.6 = 6.6 percentage points, while cost rises from $2.94 to $3.88. The increase is ($3.88 − $2.94) ÷ $2.94 = 32.0%.
- Moving Opus from high to xhigh adds 2.2 points and ($7.35 − $3.88) ÷ $3.88 = 89.4% more cost.
- Opus max costs $11.24 ÷ $7.35 = 1.53 times its xhigh point, but scores 1.6 points lower.
- Astra’s displayed peak is high. Max costs $10.35 ÷ $7.21 = 1.44 times as much and scores 1.2 points lower.
Those calculations describe the published points. They do not prove that extra reasoning causes worse performance. Anthropic reports a ±2.6-point standard error for its headline Opus result, and small differences should not be treated as certain quality changes. The practical implication is narrower: the chart gives no reason to make max the automatic choice for this workload. Benchmark footnotes.
For an initial shortlist, Opus medium is a reasonable budget baseline, high is the next setting to examine, and xhigh becomes interesting when the added accepted work is worth the larger bill. Astra high is a more defensible coding comparison point than assuming its maximum effort must be best.
FrontierCode and knowledge work tell different stories
FrontierCode evaluates changes to real codebases. In Anthropic’s displayed series, Opus medium scores 54.6% at $0.80 per task; its max point is 54.4% at $6.19. Astra max is 53.3% at $4.36. Opus medium therefore costs $0.80 ÷ $4.36 = 18.3% of Astra max’s displayed cost, with a 1.3-point higher score. The tiny medium-versus-max Opus score difference should not be sold as a meaningful win; the cost difference is the useful signal. Anthropic’s FrontierCode chart.
GDPval-AA v2.1, which evaluates professional work, has a different effort curve. Here are the launch page’s displayed Elo and estimated cost per task:
| Effort | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|
| Low | 1224 Elo / $0.21 | 1366 Elo / $0.85 |
| Medium | 1576 Elo / $0.86 | 1468 Elo / $1.82 |
| High | 1692 Elo / $1.54 | 1485 Elo / $2.43 |
| Xhigh | 1820 Elo / $4.21 | 1516 Elo / $3.04 |
| Max | 1846 Elo / $8.92 | 1542 Elo / $4.53 |
Source: Anthropic’s GDPval-AA v2.1 chart. Elo is a relative rating, not a completion percentage.
Opus improves at every step in this series. Its medium point is 34 Elo above Astra max, at $0.86 ÷ $4.53 = 19.0% of the displayed cost. Within Opus, moving from xhigh to max adds 26 Elo while cost increases by ($8.92 − $4.21) ÷ $4.21 = 111.9%. That is a capacity-versus-budget decision, not an automatic upgrade.
Business automation also needs its own answer. Anthropic’s AutomationBench chart puts Opus max at 40.0% and $1.37, versus Astra max at 41.4% and $1.77. The page says these results came from Zapier, with Opus evaluated during early access and safeguard interventions counted as failures rather than completed through fallback models. This gives Astra a higher displayed score on that task set, while Opus has the lower displayed cost. AutomationBench methodology and chart.
Do not average those percentages with GDPval Elo or turn all four charts into one “overall winner.” They measure different work.
What we can say about computer use and research
The Opus launch table reports 81.8% partial score on OSWorld 2.0 and leaves its Astra cell empty. OpenAI separately publishes Astra results on an offline OSWorld set. A missing cell should stay missing until the versions, task sets and scoring are reconciled. We do not use those two headline numbers to award a computer-use win. Anthropic’s table, OpenAI’s computer-use evaluation description.
For research, choose evidence that resembles the deliverable. A terminal-based science benchmark, a web-search task and a sourced company report stress different abilities. The material reviewed here does not provide a complete, matched Astra-versus-Opus effort sweep for all three. Readers choosing a research assistant should weigh source accuracy, the ability to finish the deliverable, and review time alongside the coding tables.
That limitation matters more than another decimal place in a composite ranking. Our cost-per-successful-job guide explains how failures and corrections belong in the denominator.
Work through the API bill
These are Standard rates per million tokens. Cache writes are a separate billing category, not an extra fee to add to the same tokens after already charging them as ordinary fresh input.
| Billing category | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|
| Ordinary input | $10 | $4 |
| Cache read | $1 | $0.20 |
| Cache write | $12.50 | $5 for five minutes; $8 for one hour |
| Output, including billable reasoning | $50 | $20 |
Sources: Astra model reference, Opus model reference, OpenAI reasoning billing, Anthropic’s explanation of thinking costs.
The ordinary input and output ratio is $10 ÷ $4 = $50 ÷ $20 = 2.5. Cache reads have a different ratio: $1 ÷ $0.20 = 5. A cache-heavy session can therefore have a wider price gap than the headline 2.5× figure suggests.
Example 1: one short, uncached request
Assume 20,000 input tokens and 2,000 output tokens, including any billable reasoning:
- Astra: 0.020 × $10 + 0.002 × $50 = $0.20 + $0.10 = $0.30.
- Opus: 0.020 × $4 + 0.002 × $20 = $0.08 + $0.04 = $0.12.
The ratio is $0.30 ÷ $0.12 = 2.5. Equal token counts isolate price; they do not predict either model’s actual usage on the same prompt.
Example 2: an agent session with explicit cache writes
Assume the session totals 200,000 ordinary input tokens, 100,000 cache-write tokens, 1.8 million cache-read tokens and 150,000 output tokens. Each individual Astra request stays at or below 272,000 input tokens. Use Opus’s five-minute write rate.
- Astra: 0.2 × $10 + 0.1 × $12.50 + 1.8 × $1 + 0.15 × $50 = $2 + $1.25 + $1.80 + $7.50 = $12.55.
- Opus: 0.2 × $4 + 0.1 × $5 + 1.8 × $0.20 + 0.15 × $20 = $0.80 + $0.50 + $0.36 + $3 = $4.66.
The ratio is $12.55 ÷ $4.66 = 2.69. With a one-hour Opus write, replace $0.50 with 0.1 × $8 = $0.80, giving $4.96. These examples exclude separately charged tools, regional premiums, taxes and discounts. The token counts already include the stipulated writes; do not add them again.
Astra’s rates rise for a request above 272,000 input tokens: input and cache rates double, and output becomes 1.5 times the standard rate for the whole request. The session example deliberately stays below that threshold. Astra pricing conditions.
Example 3: what extra effort must earn back
Suppose a higher setting adds 20,000 output tokens across a task while everything else stays fixed. The added token cost is 0.020 × $50 = $1 on Astra, or 0.020 × $20 = $0.40 on Opus. This is an illustration, not a measured high-effort surcharge: effort can also change tool calls, input and retries.
A useful break-even rule is: extra effort pays when its added cost is less than the expected rework it prevents. If it costs $0.40 more and prevents a $4 repeat on more than 10% of comparable jobs, then $4 × 0.10 = $0.40 reaches break-even before valuing saved time. Real failure probabilities must come from your workload; the published scores cannot supply them.
Recommendations by workload
| Workload | Starting shortlist | Reason |
|---|---|---|
| Everyday coding with checks you can run | Opus medium, then high | The published coding curves make these useful price-performance candidates. |
| Difficult terminal coding | Opus high/xhigh; Astra high | Compare the points near each model’s displayed peak before paying for max. |
| Professional documents and analysis | Opus medium/high | GDPval gives these configurations a strong published case; review factual accuracy in the actual output. |
| Business workflows across tools | Astra medium/high/max and Opus high/max | AutomationBench shows a different trade-off from coding. Select for accepted completion and budget. |
| Computer use or specialized research | Keep both eligible | The reviewed evidence does not establish a matched effort-by-effort winner for those workloads. |
These are evaluation starting points based on published evidence, not claims of Kingy hands-on results. If your existing Astra workflow is reliable, a cheaper listed rate alone is not a migration plan. If you are starting fresh and cost matters, Opus medium/high is a sensible place to begin.
Sources and comparison method
The numerical effort tables were transcribed from accessible point labels in the publishers’ live charts on September 22. “Med” has been expanded to “medium.” Figures remain attached to their original chart: for example, Anthropic displays Astra’s AutomationBench max cost as $1.77, while OpenAI’s September 22 chart displays $1.73. We do not combine the two cost series or silently reconcile the difference.
Anthropic’s benchmark notes also describe production safeguards and fallback models on some evaluations. That setup is part of the reported result. Chart costs are the publishers’ estimates, not audited Kingy invoices, and small score gaps may reflect evaluation noise.
For a cheaper OpenAI alternative, continue with GPT-6 Sol vs Astra across effort settings. For the broader model family, read our GPT-6 Sol and Luna comparison.
