Updated September 22, 2026 to include GPT-6 Sol and GPT-6 Luna. The original URL is unchanged. The featured chart predates these additions; the tables below cover all five models.
Claude Opus 5.5 and OpenAI’s GPT-6 Sol and GPT-6 Luna arrived on the same day. That changes this comparison. GPT-5.6 Sol is no longer the obvious budget choice: GPT-6 Sol costs half as much per token at standard rates, and Luna costs one-twentieth as much as the new Sol.
Opus 5.5 remains a strong candidate for demanding coding and knowledge work. GPT-6 Astra has strong reported results in computer use, mathematics and science. The new Sol and Luna make the decision more sensitive to workload and budget. A model that solves a task reliably for a few cents can be more useful than a higher-scoring model that costs several dollars.
We compare specs, published benchmarks and worked API costs for all five. Vendor results are attributed, missing measurements stay blank, and Kingy has not yet run its own matched evaluation of this lineup. OpenAI’s launch comparisons against Claude Opus 5 must not be read as comparisons against the newly released Opus 5.5.
For the individual releases, see our Claude Opus 5.5 guide and GPT-6 Sol and Luna guide. This article focuses on the buying decision.
The verdict in 60 seconds
- Demanding coding and knowledge work: shortlist Claude Opus 5.5. Anthropic reports 54.4% on FrontierCode Main and 1,846 Elo on GDPval-AA. The independent-evaluation snapshot below also supports its candidacy. These results do not establish a universal winner across the expanded lineup.
- Everyday coding and agents on a tighter budget: start by testing GPT-6 Sol. Its $2 input / $10 output rate is half Opus 5.5’s and half GPT-5.6 Sol’s promotional rate. OpenAI reports 49.3% on FrontierCode Main at max effort and 33.2% on AutomationBench at xhigh.
- Focused, high-volume work: test GPT-6 Luna first. At $0.10 input / $0.50 output, it is the cheapest of these five. OpenAI reports 42.4% on FrontierCode Main at max effort, but its lower AutomationBench score makes it a less convincing default for difficult workflows across multiple apps.
- Hard computer-use, math and science tasks: keep GPT-6 Astra in the evaluation. Its reported capability remains strong, but $10/$50 makes it five times GPT-6 Sol’s standard token price and 100 times Luna’s.
- Long context: compare the actual bill. Opus 5.5 has no long-context surcharge. Even after OpenAI’s surcharge, our 600K-input / 20K-output example costs $2.70 on GPT-6 Sol, $2.80 on Opus 5.5 and $0.135 on Luna. Those are equal-token calculations, not quality-adjusted results.
- GPT-5.6 Sol is now the migration baseline. Keep it where you have validated behavior or a specific dependency. Its former budget recommendation needs a fresh comparison with the two new models.
The five models at a glance
| Model | API ID | Context / max output | Role in this comparison |
|---|---|---|---|
| Claude Opus 5.5 | claude-opus-5-5 | 1M / 128K | Demanding coding and knowledge work |
| GPT-6 Astra | gpt-6-astra | 1.05M / 128K | Premium capability for difficult tasks |
| GPT-6 Sol | gpt-6-sol | 1.05M / 128K | Lower-cost coding and agent workflows |
| GPT-6 Luna | gpt-6-luna | 1.05M / 128K | Focused, high-volume tasks |
| GPT-5.6 Sol | gpt-5.6-sol | 1.05M / 128K | Previous-generation baseline |
All five accept text and images and produce text. Opus 5.5 and the two new GPT-6 models launched on September 22, 2026. Astra first launched on September 3; GPT-5.6 Sol entered preview on June 26 and general availability on July 9.
The GPT-6 Sol API page lists an April 20, 2026 knowledge cutoff; Luna’s page lists May 18. Both support none, low, medium (default), high, xhigh and max reasoning effort. Opus 5.5 defaults to medium and cannot disable thinking. Product-level modes such as Ultra should be checked separately from the API’s reasoning controls.
At launch, GPT-6 Sol and Luna are available through the API, ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users can access Luna in the desktop app. OpenAI says they are not yet available in Chat, and the rollout is gradual.
Pricing: the rate cards
USD per million tokens at standard API rates, checked September 22, 2026. Cached reads and writes are separate categories.
| Model | Input | Output | Cached read | Cache write |
|---|---|---|---|---|
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | $5.00 (5-min); $8.00 (1-hr) |
| GPT-6 Astra | $10.00 | $50.00 | $1.00 | $12.50 |
| GPT-6 Sol | $2.00 | $10.00 | $0.20 | $2.50 |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | $0.125 |
| GPT-5.6 Sol (promo) | $4.00 | $20.00 | $0.40 | $5.00 |
| GPT-5.6 Sol (list) | $5.00 | $30.00 | $0.50 | $6.25 |
GPT-5.6 Sol’s promotion was announced August 21 and is scheduled to end November 21, 2026. That deadline belongs to the older model. OpenAI presents GPT-6 Sol’s $2/$10 and Luna’s $0.10/$0.50 as their launch API rates; the launch announcement does not give them the same November expiry.
GPT-6 Sol’s input and output rates are 50% below GPT-5.6 Sol’s promotional rates. Luna’s input is 50% below GPT-5.6 Luna’s $0.20 promotional rate, while output falls from $1.20 to $0.50, a 58.3% reduction. The headline “50% cheaper” is therefore a rounded description for Luna.
For GPT-6 Sol and Luna, prompts above 272K input tokens double input and cache rates and multiply output rates by 1.5 for the full request. This is also the long-context pricing structure used for Astra and GPT-5.6 Sol in the original comparison. Opus 5.5 has no surcharge across its 1M-token window.
The new model pages list Batch and Flex at 50% of standard rates and Fast mode at twice the applicable rates. Regional processing adds 10% where available; EU data residency for Sol and Luna requires Standard processing. Opus 5.5’s Fast mode costs $8/$40, and Anthropic applies a 1.1× multiplier for US-only inference.
GPT-6 Sol now matches Opus 5.5’s cached-read price and undercuts its fresh-input, output and cache-write rates. Luna is cheaper in every standard token category. Caching hit rates, reasoning tokens, retries and tool charges still determine the final bill. Sources: OpenAI launch pricing, Sol API pricing, Luna API pricing.
Benchmarks: the vendor-reported numbers
Every number below was published by a vendor or a benchmark owner, and we label which one in each case. Kingy hasn’t reproduced any of them yet. Scores come from different effort settings, harnesses and prompts, so read the footnotes in the next section before drawing conclusions.
| Benchmark | Opus 5.5 | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | GPT-5.6 Sol | Source |
|---|---|---|---|---|---|---|
| Agentic coding | ||||||
| Terminal-Bench 4.0 | 66.4% (xhigh) | 57.9% (high) | — | — | 37.3% | Anthropic; OpenAI figures as reported by OpenAI |
| FrontierCode v1.1 (Main) | 54.4% | 53.3% | 49.3% (max) | 42.4% (max) | 47.5% | Anthropic / OpenAI |
| FrontierCode v1.1 (Extended) | — | 64.5% | — | — | 60.6% | OpenAI |
| CursorBench 4.0 | 57.8% | — | — | — | 41.7% | Anthropic |
| DeepSWE v1.1 | — | 74.1% | 68.8% (max)† | 66.6% (max)† | 72.7% | OpenAI |
| Knowledge work and business automation | ||||||
| GDPval-AA v2.1 (Elo) | 1846 | 1542 | — | — | 1588 | Anthropic, citing Artificial Analysis |
| AutomationBench (Zapier) | 40.0% | 41.4% (max) | 33.2% (xhigh)† | 20.7% (max)† | 28.8% (max) | Zapier for original models; OpenAI launch for GPT-6 Sol and Luna |
| BrowseComp | — | 91.5% | — | — | 90.4% | OpenAI |
| Reasoning, math and science | ||||||
| Humanity’s Last Exam (with tools) | 67.7% | 57.2% | — | — | — | Anthropic |
| Terminal-Bench-Science 0.1 | 58.7% | 64.6% | — | — | 22.4% | Anthropic; Astra as reported by OpenAI |
| GPQA Diamond | — | 96.0% | — | — | 94.6% | OpenAI |
| FrontierMath Tier 4 | — | 97.6% | — | — | 83.0% | OpenAI |
| ARC-AGI-2 | — | 95.0% | — | — | 92.5% | OpenAI |
| ARC-AGI-3 (standard harness) | — | 62.7% | — | — | — | ARC Prize |
| Computer use and vision | ||||||
| OSWorld 2.0 (partial credit) | 81.8%* | 72.6%* | 60.5% (xhigh)* | — | 65.7%* | *Different harnesses; not comparable across vendors (see below) |
| ScreenSpot-Pro | — | 92.7% | — | — | 76.9% | OpenAI |
| Chartography (with tools) | 89.0% | — | — | — | — | Anthropic |
| Long context and security | ||||||
| OpenAI MRCR v2 (512K–1M) | — | 96.3% | — | — | 73.8% | OpenAI |
| ExploitBench | — | 100% | — | — | 78.5% | OpenAI |
A dash means no verified score is included here, not zero performance. New Sol and Luna figures come from OpenAI’s September 22 release and its chart labels. The older columns retain the attributed launch results from the original comparison. † New launch runs are not a controlled rerun of every older result. * OSWorld scores use different effort settings and vendor harnesses; do not rank models by subtracting these cells.
What GPT-6 Sol and Luna change
The strongest case for the new models is cost per attempt at useful levels of performance. On FrontierCode Main, OpenAI reports GPT-6 Sol at 49.3% and $2.14 per task at max effort, versus GPT-5.6 Sol at 47.5% and $5.19. That is a 1.8-point gain with about 59% lower reported cost per task. Astra scores higher at 53.3%, at $4.59 per task in the same launch chart.
| OpenAI launch evaluation | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| FrontierCode Main (max) | 49.3%; $2.14/task | 42.4%; $0.11/task |
| AutomationBench (selected effort) | 33.2%; $0.27/task (xhigh) | 20.7%; $0.037/task (max) |
| Agents’ Last Exam (max) | 56.4%; $2.93/task | 50.9%; $0.15/task |
| DeepSWE v1.1 (max) | 68.8% | 66.6% |
| Factual-error evaluation (max; lower is better) | 4.6%; $0.18/task | 7.6%; $0.012/task |
All results in this table are OpenAI-reported. Its factuality test deliberately selects conversations where an earlier model made an error; the percentages are not everyday hallucination rates. They also cannot be compared directly with Artificial Analysis’s AA-Omniscience metric.
Luna’s $0.11 FrontierCode attempt is about one-twentieth of Sol’s $2.14 attempt, but its score is 6.9 points lower. For a small change with a reliable test, Luna may be worth trying first. For a migration that needs close judgment across a large repository, failed attempts and review time can erase those savings. That is a workload recommendation, not a result Kingy has measured.
More effort does not always improve the reported score. GPT-6 Sol reaches 33.2% on AutomationBench at xhigh for $0.27, then scores 32.0% at max for $0.34. Start with an effort setting you can afford and test whether increasing it helps your tasks.
GPT-6 Sol is not a proven upgrade on every individual benchmark. The original Astra release reported GPT-5.6 Sol at 72.7% on DeepSWE, while the new release reports GPT-6 Sol at 68.8%. These are separate published runs. The difference deserves testing before migration; it does not justify either a blanket improvement claim or a controlled-regression claim.
Where the numbers don’t line up
The original three-model comparison exposed several measurement differences. Those caveats still apply when adding GPT-6 Sol and Luna, especially when a new launch uses a different effort setting or an older Claude model as its reference.
1. Terminal-Bench 4.0: an 8.5-point lead in Anthropic’s table, a tie in Artificial Analysis’s
Anthropic reports Opus 5.5 at 66.4% against Astra’s 57.9%. The footnotes explain the setup. Opus 5.5 ran at xhigh effort. Astra’s number is OpenAI’s own figure at high effort, and Anthropic describes each as “each model’s highest score.” The two scores come from different labs, different effort levels and presumably different harnesses. Anthropic also puts the standard error at ±2.6 points for Opus 5.5.
Artificial Analysis runs every model in one harness. In that run, Opus 5.5 and Astra tie at 59.6%. That’s the more useful number if you’re choosing between them. Anthropic still holds a real claim here: it says Opus 5.5 at default effort matches Astra “at about 40% of the cost.”
2. OSWorld 2.0: the two labs score Claude differently
Both labs report “partial credit” OSWorld 2.0 scores, but they don’t agree on the same model. Anthropic scores Opus 5 at 74.0%. OpenAI’s chart scores Opus 5 at 70.2% (OpenAI version v2026.08.08, offline). That 3.8-point gap between labs is bigger than several of the gaps between models. So don’t read Opus 5.5’s 81.8% against Astra’s 72.6% as a 9-point lead. They aren’t on the same scale.
3. AutomationBench: GPT-5.6 Sol scores 28.8% or 18.1%, depending on who you ask
OpenAI’s Astra post puts GPT-5.6 Sol at 18.1%. Zapier’s public AutomationBench leaderboard puts GPT-5.6 Sol (max) at 28.77%. Zapier created the benchmark, and it’s also the source Anthropic cites. Effort settings or evaluation setup may explain the gap, but the published figures do not establish the cause. Zapier’s leaderboard shows something else worth knowing: Astra’s lead depends on effort. It scores 41.4% at max, 37.1% at high and 30.3% at low, at $1.77, $1.45 and $1.08 per task respectively.
4. FrontierCode: two subsets, two stories
OpenAI reports both FrontierCode v1.1 subsets. On “Main,” Astra scores 53.3%, which is the number Anthropic compares against (Opus 5.5: 54.4%). On “Extended,” Astra scores 64.5%. Anthropic didn’t publish an Extended score. OpenAI also notes that Astra ran “with a developer message” about code-quality conventions. The two labs even report different scores for their shared reference model, Fable 5.1, on Main: 50.3% from Anthropic and 50.9% from OpenAI.
5. ARC-AGI-3: 62.7% or 99.9%, depending on the harness
OpenAI’s headline ARC-AGI-3 score for Astra is 99.9%. ARC Prize’s own write-up shows that number came from a “provider adapter” harness that keeps opaque reasoning state between requests and compacts long conversations. It cost $18,817 at high effort. Under ARC’s standard harness at max effort, Astra scores 62.7%, at a cost of $26,098. That’s still state of the art. But 62.7% is the number to compare against other models.
6. Your harness and effort setting shape the results as much as the model does
Every score above depends on the effort level, the scaffolding, and in Anthropic’s case the safeguards. Anthropic ran Opus 5.5 “with production safeguards enabled.” When they triggered, cybersecurity tasks were finished by Claude Opus 4.8 and biology tasks by Claude Opus 5. Safeguard interventions on AutomationBench were scored as failures. So some Opus 5.5 scores in cyber-adjacent areas are partly an older model’s work. That’s also roughly what you’ll get in production.
Independent evals
Artificial Analysis already lists both new models. This September 22 snapshot uses Intelligence Index v4.3.2 and the max-effort model pages; Opus 5.5 is the variant with default fallback enabled. The scores measure this evaluation suite, not every possible workload.
| Model | AA Index | Output tokens for Index | Weighted cost / Index task | Output speed |
|---|---|---|---|---|
| Claude Opus 5.5 | 58 | 260M | $5.98 | Not listed |
| GPT-6 Astra | 53 | 60M | $3.26 | 57.7 tokens/s |
| GPT-6 Sol | 48 | 77M | Not listed | 104.4 tokens/s |
| GPT-6 Luna | 37 | 150M | Not listed | 157.2 tokens/s |
| GPT-5.6 Sol | 47 | 90M | $1.99 | 72.6 tokens/s |
Sources: Artificial Analysis model pages for Opus 5.5, Astra, GPT-6 Sol, Luna and GPT-5.6 Sol. Missing cost values are not zero. Total output tokens and weighted per-task cost are different aggregates; dividing one by the other does not give a valid task count.
Opus 5.5 leads these five on this index. GPT-6 Sol improves only one index point over GPT-5.6 Sol, while halving standard input and output prices. Luna scores lower but is substantially cheaper and faster. That supports three different buying decisions, rather than a single overall ranking.
Artificial Analysis currently lists an 872K context window for GPT-6 Sol, while OpenAI’s API page lists 1,050,000 tokens. The specifications table uses the provider’s documented limit. We have not independently tested maximum usable context. Different index versions also must not be mixed: Astra’s older launch score of 61 belongs to v4.1.1, not this v4.3.2 comparison.
Token usage: the numbers that drive cost
A low token price does not guarantee a low bill. Total cost includes fresh input, cached input, cache writes, output including billed reasoning, tool use and retries. For agents, the useful comparison is the total bill divided by tasks that pass your acceptance checks.
Artificial Analysis’s current max-effort pages put Opus 5.5 at $5.98 per weighted Index task and Astra at $3.26. On that workload, Astra costs less despite its higher token rates. Opus 5.5 generates 260M output tokens across the Index, versus Astra’s 60M. These are max-effort measurements, not a prediction for ordinary calls at default settings.
Anthropic reports much lower costs at Opus 5.5’s default medium effort: on FrontierCode it says the model beats Astra at about one-fifth of the cost per task. That is a vendor result on a named benchmark. It cannot be generalized into “Opus is always cheaper at medium,” especially now that GPT-6 Sol and Luna are available.
The new models also differ in token use. Artificial Analysis reports 77M total output tokens for GPT-6 Sol and 150M for Luna across its Index. Luna emits nearly twice as many, but its standard output rate is one-twentieth of Sol’s. We have not substituted these totals for the missing all-in per-task costs.
Caching and tokenization
OpenAI says GPT-6 caching improvements preserve earlier cached context when developers change reasoning effort or enable and disable tools, and add explicit cache breakpoints. These features can help an agent move to higher effort for a hard step without losing all the savings from its existing context. Measure actual cache hits in your application.
Token counts also differ between providers. This update removes the original article’s estimated “equivalent token” price, which extrapolated from older Claude models without verifying Opus 5.5 or GPT-6. For a defensible comparison, send the same documents and tasks to each model and record the billed usage.
What it costs: three worked examples
These are calculated API bills using equal token counts, with no tool fees, regional premiums, Batch discounts, retries or human review time. They isolate the rate card; they do not imply equal output quality or token usage.
| Model | A: standard call | B: long-context call | C: cached agent loop |
|---|---|---|---|
| Claude Opus 5.5 | $0.60 | $2.80 | $7.00 |
| GPT-6 Astra | $1.50 | $13.50 | $22.50 |
| GPT-6 Sol | $0.30 | $2.70 | $4.50 |
| GPT-6 Luna | $0.015 | $0.135 | $0.225 |
| GPT-5.6 Sol (promo) | $0.60 | $5.40 | $9.00 |
| GPT-5.6 Sol (list) | $0.80 | $6.90 | $12.00 |
- A: 100K input + 10K output. Below the long-context threshold. GPT-6 Sol costs 100,000 × $2/M + 10,000 × $10/M = $0.30. Luna costs 1.5 cents.
- B: 600K input + 20K output. OpenAI’s long-context rates apply to the whole request. Sol uses $4/M input and $15/M output: $2.40 + $0.30 = $2.70. Luna uses $0.20/$0.75: $0.12 + $0.015 = $0.135. Opus 5.5 stays at $4/$20: $2.40 + $0.40 = $2.80.
- C: 50 turns with a 200K cached prefix, 5K fresh input and 3K output per turn, plus one 200K cache write. That is 10M cached reads, 250K fresh input and 150K output. Sol costs $2 + $0.50 + $1.50 + $0.50 = $4.50. Luna costs $0.10 + $0.025 + $0.075 + $0.025 = $0.225. The prefix is assumed to remain cached without another write; Opus uses its five-minute write rate.
The new models overturn the original article’s claim that Opus 5.5 is the cheapest option for long-context and cached workloads. In these examples, GPT-6 Sol is cheaper in all three cases, and Luna is cheaper again. Opus’s flat long-context rate remains useful, but it no longer wins the bill automatically.
The quality-adjusted choice still needs testing. Count failed attempts, escalations to a stronger model and review time. A cheap call that requires three repairs may be worse value than a more expensive call that passes on the first attempt.
Speed and latency
Artificial Analysis’s current model pages report Luna at 157.2 output tokens per second, GPT-6 Sol at 104.4, GPT-5.6 Sol at 72.6 and Astra at 57.7. These replace the older speed snapshot in this article. Opus 5.5’s AA page does not currently list a throughput measurement; Anthropic says it generates output more than 30% faster than Opus 5.
Throughput measures how quickly output arrives once generation is under way. It does not measure the time spent reasoning before a reply, calling tools, retrying a failed step or finishing the whole job. For an interactive product, measure time to a usable answer as well as tokens per second.
Both new API models support reasoning.effort: "none", so GPT-5.6 Sol is no longer the only model in this comparison with a reasoning-off option. Test quality again when disabling reasoning; a faster response is useful only if it still satisfies the task.
Safety, access and restrictions
GPT-6 Astra is the first model OpenAI rates “Critical” for cybersecurity under its Preparedness Framework. That means it can find new vulnerabilities and develop exploits against well-defended systems with little human direction. Its advanced cyber capabilities are gated behind trusted-access programs. The system card reports gains on several safety metrics. Misaligned outcomes in realistic work environments fell to 3.4% (from GPT-5.6 Sol’s 18.8%). Gray Swan measured indirect prompt-injection success at 8.5% (vs. GPT-5.6 Sol’s 27.0%). The same system card reports a “substantial decrease in chain-of-thought monitorability,” tied to Astra’s recurrent-depth reasoning. Apollo Research measured Astra’s evaluation awareness at 50.6% at max reasoning. Apollo also cautioned that low misbehavior rates “do not provide substantial evidence about the model’s alignment.”
Claude Opus 5.5 ships with the same cyber, biology and frontier-LLM-development safeguards as Fable 5.1. Anthropic says it’s the strongest model it has tested on its ~2,000-scenario automated behavioral audit. It says the model attempted boundary circumvention “85% less often” than Opus 5 and resists prompt injection better. For security teams, the practical point is this: when the cyber classifier fires, Opus 4.8 handles the task instead. Biology work needs enrollment in Anthropic’s Life Sciences Verification Program. If offensive security research is your use case, Astra’s trusted-access tier currently offers more capability.
GPT-5.6 Sol launched as a government-coordinated restricted preview and went GA 13 days later. It sits below the Critical cyber threshold. In Astra’s system card, OpenAI and its third-party evaluators consistently score GPT-5.6 Sol as the less-aligned of the two OpenAI models.
GPT-6 Sol and Luna: OpenAI reports improvements over their GPT-5.6 counterparts in its alignment evaluations, including fewer misleading claims about coding work. The launch post warns that these evaluations deliberately use difficult situations and do not measure everyday failure rates. Do not transfer Astra’s cybersecurity classification or trusted-access permissions to another GPT-6 model merely because they share a family name. Consult the system card linked from the launch for the applicable model and access conditions.
If you’re migrating to Opus 5.5: breaking changes
Opus 5.5 isn’t a drop-in replacement for claude-opus-5. From Anthropic’s migration notes:
- You can’t disable thinking.
thinking: {"type": "disabled"}and manual budgets return a 400 error. Use adaptive thinking and control depth witheffort. - Forced tool use is gone.
tool_choiceofanyor a named tool returns a 400. Useautowith strict tool use, or structured outputs. - Default effort dropped from
hightomedium. Set it explicitly and re-run your evals. - Text between tool calls now arrives as thinking blocks. With the default
display: "omitted", streaming UIs go quiet between tool calls. - Computer use moved to
computer_toolset_20260801on the Claude API and Google Cloud. Bedrock still accepts the older tool. - Budget extra
max_tokensfor thinking, especially atxhighandmax.
If you’re migrating from GPT-5.6 Sol
Use the explicit API IDs gpt-6-sol or gpt-6-luna. Set reasoning effort explicitly, then compare the same tasks against your existing GPT-5.6 Sol configuration. Keep prompts, tools, acceptance checks and retry budgets fixed for the initial comparison.
The new model pages recommend the Responses API for built-in tools and function calling. Chat Completions supports function calling only when reasoning_effort is none. An integration that combines reasoning and tools should account for that before changing model IDs.
For Luna, begin with tasks you can verify cheaply, such as structured extraction, classification or small code changes with meaningful tests. Escalate failures to GPT-6 Sol, Opus 5.5 or Astra based on which model passes that task class most reliably. This is a deployment approach to test, not a measured routing result.
Which model for which job
| Workload | Models to test first | Reason |
|---|---|---|
| Difficult coding, repository migrations | Opus 5.5; Astra | Strong reported coding results. Compare complete, verified changes and review time. |
| Routine coding and general agent workflows | GPT-6 Sol | Half Opus 5.5’s standard token rates; AA Index 48 versus GPT-5.6 Sol’s 47. |
| High-volume extraction and focused tasks | GPT-6 Luna | Lowest rates and highest measured output throughput here. Validate task accuracy. |
| Knowledge work, analysis and decks | Opus 5.5; GPT-6 Sol | Opus has strong GDPval results; Sol offers lower-cost iteration and reported professional-work gains. |
| Long-document work | GPT-6 Sol; Opus 5.5; Luna for simpler tasks | Sol slightly wins the worked long-context bill. Flat pricing alone does not decide value; verify retrieval and reasoning. |
| Demanding GUI agents, math and science | Astra; compare cheaper models on easier steps | Strong reported specialist benchmarks, with a substantial price premium. |
| Security research | Astra or Claude under applicable verified access | Model capability, safeguards and program eligibility all affect what can run. |
| Existing GPT-5.6 Sol deployment | A/B test GPT-6 Sol before migrating | The cost case is strong, but individual benchmark gains and application behavior are not uniform. |
These are starting points for evaluation. The useful winner is the model and effort setting that meets your acceptance criteria at the lowest total cost, including retries and review.
How to choose after the September 22 launches
GPT-6 Sol is the first replacement to test if you were choosing GPT-5.6 Sol for price. It halves the standard token rates, improves the current Artificial Analysis Index score from 47 to 48 and increases measured output throughput. Luna deserves a separate trial for focused work: its much lower price comes with a lower overall index score and weaker results on some complex agent workflows.
Opus 5.5 still leads these five on the current AA Index, while Astra remains a strong candidate for difficult specialist work. Neither is the cheapest option by default. Before committing, run representative tasks at explicit effort settings and record pass rate, elapsed time, billed tokens, cache hits and repair effort. Kingy has not yet published a matched five-model test; the figures here are sourced measurements and transparent calculations.
Sources and methodology
Updated September 22, 2026. New Sol and Luna specifications and pricing were checked against OpenAI’s API pages, launch scores against OpenAI’s announcement and accessible chart labels, and the five-model Index, token-use and speed snapshot against Artificial Analysis’s model pages. Older benchmark rows retain their attributed sources from the original comparison. Kingy has not independently reproduced these evaluations. Worked examples use equal billed token counts and the stated pricing assumptions.
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI API: GPT-6 Sol model page
- OpenAI API: GPT-6 Luna model page
- Artificial Analysis: GPT-6 Sol (max)
- Artificial Analysis: GPT-6 Luna (max)
- Artificial Analysis: Claude Opus 5.5 (max with default fallback)
- Anthropic: Introducing Claude Opus 5.5
- Claude Platform Docs: What’s new in Claude Opus 5.5
- Claude Platform Docs: Pricing
- Claude Platform Docs: Context windows
- OpenAI: GPT-6 Astra
- OpenAI: GPT-6 Astra System Card
- OpenAI API: GPT-6 Astra model page
- OpenAI: Previewing GPT-5.6 Sol
- OpenAI: GPT-5.6 general availability
- OpenAI API: GPT-5.6 Sol model page
- OpenAI Developer Community: GPT-5.6 Sol price reduction
- Artificial Analysis: Benchmarking GPT-6 Astra
- Artificial Analysis: GPT-6 Astra and GPT-5.6 Sol model pages
- OfficeChai: Artificial Analysis results for Claude Opus 5.5
- Zapier: AutomationBench leaderboard
- ARC Prize: GPT-6 Astra on ARC-AGI-3
- Epoch AI: GPT-6 Astra
- CodeRabbit: GPT-6 Astra code review evaluation
- OpenRouter: Claude Opus 5.5
- Simon Willison: Claude token counts and TextKit: tokens per word, measured
Trending on Kingy
Keep reading with the stories getting the most attention now.
-
GPT-6 Sol and GPT-6 Luna: Specs, Benchmarks, Pricing and How They Compare to Claude Opus 5.5, Fable 5.1 and Gemini
Read story -
Claude Opus 5.5: Specs, Benchmarks, Pricing and How It Stacks Up Against GPT-6 Astra, Fable 5.1 and Every Frontier Model
Read story -
Grok 4.7 Benchmarks: Specs, Pricing and Frontier Model Comparisons
Read story
