GPT-6 Sol and GPT-6 Luna are OpenAI’s new mid-tier and budget models, launched on September 22, 2026, about three weeks after the flagship GPT-6 Astra. Sol costs $2 per million input tokens and $10 per million output tokens. Luna costs $0.10 and $0.50. That is half of what GPT-5.6 Sol and GPT-5.6 Luna cost on input. OpenAI says both models were trained “with similar methods as GPT-6 Astra” and bring Astra’s gains in professional work, factuality, coding, computer use and alignment down to cheaper price points.
The pitch is cost efficiency, not a new capability ceiling. OpenAI’s own charts show Sol scoring slightly lower than GPT-5.6 Sol’s top setting on two of the six headline evals, while costing about 60% less per task. Below we go through every spec, every published eval (including the per-effort data behind OpenAI’s charts), the full pricing picture, and how Sol and Luna compare to Claude Opus 5.5 (which also launched today), Claude Fable 5.1, GPT-6 Astra, Claude Sonnet 5 and Google’s Gemini lineup.
Key takeaways
- Pricing: Sol is $2 / $10 and Luna is $0.10 / $0.50 per million tokens. Cached input reads get a 90% discount ($0.20 for Sol, $0.01 for Luna). Input prices are exactly half of GPT-5.6’s. Luna’s output price fell 58%, from $1.20 to $0.50.
- Specs: Both have a 1.05M-token context window, 128K max output tokens, text and image input, text output, and OpenAI’s core agent tools (function calling, web search, file search, computer use). Knowledge cutoffs are April 20, 2026 (Sol) and May 18, 2026 (Luna).
- Headline evals: On AutomationBench, Sol at xhigh scores 33.2% at $0.27 per task. That beats Claude Opus 5 at max (26.9%, 11.1x the cost) and GPT-6 Astra at low (30.3%, 3.9x the cost). On Agents’ Last Exam, Sol at max scores 56.4%, above Opus 5’s best of 55.9%.
- The asterisk: On DeepSWE and OSWorld 2.0, GPT-6 Sol’s best score (68.8% and 64.4%) is below GPT-5.6 Sol’s best (72.7% and 66.2%) and below Claude Opus 5’s best (73.7% and 70.2%). The upgrade on those evals is cost per task, not peak score.
- Luna is the standout: It scores 66.6% on DeepSWE at $0.22 per task and 50.9% on Agents’ Last Exam at $0.15 per task. Its predecessor needed $2.57 per task for 50.4% on that second eval.
- Vs. Opus 5.5: OpenAI compared against Opus 5, not the Opus 5.5 that shipped the same day. On the two benchmarks both labs report the same way, Opus 5.5 leads Sol: 40.0% vs. 33.2% on AutomationBench and 54.4% vs. 49.3% on FrontierCode 1.1. Opus 5.5 also costs twice as much per token ($4 / $20).
The GPT-6 lineup at a glance
With Sol and Luna out, the GPT-6 family has three tiers. There is no GPT-6 version of the middle “Terra” tier that GPT-5.6 had. Sol now sits at GPT-5.6 Terra’s old input price ($2) and below its output price ($12).
| Model | API ID | OpenAI’s positioning | Input / 1M | Output / 1M |
|---|---|---|---|---|
| GPT-6 Astra | gpt-6-astra |
“Our most capable model, built for the hardest end-to-end work” | $10.00 | $50.00 |
| GPT-6 Sol | gpt-6-sol |
“Built to power complex coding and agentic workflows” | $2.00 | $10.00 |
| GPT-6 Luna | gpt-6-luna |
“Our most efficient model for focused, high-volume tasks” | $0.10 | $0.50 |
GPT-6 Sol and Luna specs
| Spec | GPT-6 Sol | GPT-6 Luna | GPT-6 Astra (for reference) |
|---|---|---|---|
| Release date | September 22, 2026 | September 22, 2026 | Early September 2026 |
| Context window | 1,050,000 tokens | 1,050,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens | 128,000 tokens |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 | April 30, 2026 |
| Input modalities | Text, image | Text, image | Text, image |
| Output modalities | Text | Text | Text |
| Reasoning effort | none, low, medium, high, xhigh, max | none, low, medium (default), high, xhigh, max | low, medium, high, xhigh, max |
| Tools | Function calling, web search, file search and computer use across all three. Luna’s model page also lists structured outputs, code interpreter, image generation, MCP and skills. | ||
| Fine-tuning | Not stated | Not supported | Not stated |
OpenAI’s launch post doesn’t discuss architecture or parameter counts. What it does describe is a change in how the models write. Sol and Luna get Astra’s “collaboration style”: “more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall.” In OpenAI’s side-by-side coding example, GPT-6 Sol says what it checked and what it didn’t (desktop, narrow mobile, browser back navigation). GPT-5.6 Sol, by contrast, volunteers implementation details like its internal image-generation prompt.
GPT-6 Sol and Luna pricing
| Model | Input / 1M | Cached input / 1M | Output / 1M | Change vs. GPT-5.6 |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | Input −50% ($4 → $2), output −50% ($20 → $10) |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | Input −50% ($0.20 → $0.10), output −58% ($1.20 → $0.50) |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | New tier |
A few things in the pricing documentation are easy to miss:
- “Promotional” baseline. OpenAI says it cut prices “by 50% compared with their GPT-5.6 promotional pricing.” The comparison is against the price GPT-5.6 was actually selling at, so the cut is real. But the post doesn’t say whether GPT-6’s own prices are promotional.
- Cache writes are listed. OpenAI’s pricing page shows cache-write rates for the GPT-6 family: $2.50 per million for Sol, $0.125 for Luna and $12.50 for Astra. Cache reads are 90% off.
- Long-context surcharge. On the GPT-6 model pages, requests over 272K input tokens are billed at 2x the input and cached-input rates and 1.5x the output rate. Anthropic charges the same per-token price across its full 1M window, so this matters for very long prompts.
- Processing tiers. Batch and Flex are 50% of standard. Fast mode is 2x standard. Regional data-residency endpoints add 10%.
How the benchmarks were reported
OpenAI didn’t publish a single score per model. For each eval it charted score against cost per task at all five reasoning-effort settings (low, medium, high, xhigh, max) for GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Luna, and one or two Anthropic models. The tables below use the underlying chart data. That lets you see what the prose highlights leave out: which model has the highest score, where more effort stops helping, and where it makes results worse.
All figures here are vendor-reported. OpenAI says its own models were evaluated “in our research environment or via our API” and competitor scores were “taken from publicly available reports.” Anthropic’s models are Claude Opus 5 and Claude Fable 5 / 5.1. Claude Opus 5.5 is not in any OpenAI chart.
GPT-6 Sol: every score at every effort level
| Benchmark | low | medium | high | xhigh | max |
|---|---|---|---|---|---|
| AutomationBench 1.0.6 | 21.2% ($0.19) | 26.9% ($0.21) | 31.2% ($0.24) | 33.2% ($0.27) | 32.0% ($0.34) |
| Agents’ Last Exam V1 | 48.7% ($0.86) | 53.1% ($1.27) | 52.6% ($1.53) | 55.4% ($1.67) | 56.4% ($2.93) |
| FrontierCode 1.1 Main | 37.3% ($0.45) | 45.9% ($0.80) | 47.7% ($1.08) | 48.4% ($1.37) | 49.3% ($2.14) |
| DeepSWE v1.1 | 37.2% ($0.16) | 56.6% ($0.38) | 65.3% ($0.64) | 66.6% ($1.00) | 68.8% ($2.74) |
| OSWorld 2.0 (offline) | 43.9% ($0.97) | 54.0% ($1.32) | 58.3% ($1.64) | 60.5% ($2.21) | 64.4% ($3.25) |
| Factual error rate (lower is better) | 11.4% ($0.050) | 6.9% ($0.069) | 5.1% ($0.099) | 4.5% ($0.13) | 4.6% ($0.18) |
GPT-6 Luna: every score at every effort level
| Benchmark | low | medium | high | xhigh | max |
|---|---|---|---|---|---|
| AutomationBench 1.0.6 | 1.2% ($0.006) | 9.4% ($0.016) | 14.5% ($0.021) | 12.6% ($0.025) | 20.7% ($0.037) |
| Agents’ Last Exam V1 | 36.3% ($0.025) | 46.8% ($0.11) | 43.6% ($0.11) | 47.9% ($0.11) | 50.9% ($0.15) |
| FrontierCode 1.1 Main | 25.7% ($0.021) | 35.5% ($0.053) | 37.3% ($0.067) | 37.1% ($0.073) | 42.4% ($0.11) |
| DeepSWE v1.1 | 2.4% ($0.006) | 44.5% ($0.052) | 59.3% ($0.084) | 61.3% ($0.11) | 66.6% ($0.22) |
| OSWorld 2.0 (offline) | 8.3% ($0.030) | 31.5% ($0.062) | 41.4% ($0.12) | 46.7% ($0.16) | 52.7% ($0.27) |
| Factual error rate (lower is better) | 27.7% ($0.0024) | 17.5% ($0.0033) | 12.5% ($0.0045) | 10.2% ($0.0062) | 7.6% ($0.012) |
Two things stand out. First, Luna at low effort is barely usable for agentic work: 2.4% on DeepSWE, 1.2% on AutomationBench, 8.3% on OSWorld. Its useful range starts at medium, and it only reaches its best results at max. Second, more effort doesn’t always help. Sol peaks at xhigh on AutomationBench and on factuality. Luna scores worse at high than at medium on Agents’ Last Exam (43.6% vs. 46.8%), and worse at xhigh than at high on AutomationBench (12.6% vs. 14.5%). Test the effort setting on your own workload rather than defaulting to max.
Professional work: AutomationBench and Agents’ Last Exam
AutomationBench 1.0.6 tests agents on end-to-end business workflows using 47 tools across sales, marketing, operations, support, finance and HR. This is where Sol looks best against Anthropic’s models.
| Model (best effort) | AutomationBench score | Cost per task |
|---|---|---|
| GPT-6 Astra (max) | 41.4% | $1.73 |
| GPT-6 Sol (xhigh) | 33.2% | $0.27 |
| Claude Fable 5.1 w/ Opus 5 fallback (max) | 31.4% | $2.45+ (fallback cost not reported) |
| GPT-6 Astra (low) | 30.3% | $1.08 |
| GPT-5.6 Sol (max) | 28.8% | $0.67 |
| Claude Opus 5 (max) | 26.9% | $3.05 |
| GPT-6 Luna (max) | 20.7% | $0.037 |
| GPT-5.6 Luna (max) | 17.0% | $0.07 |
The cost-per-task ratios in OpenAI’s text match its data: Opus 5 at max costs 11.1x as much as Sol at xhigh, and Astra at low costs 3.9x. The Fable 5.1 row has a caveat OpenAI flags itself. The $2.45 figure “omits the cost of the Opus 5 fallbacks, which occurred on ~40% of tasks.” (Fable 5.1 redirects some requests to Opus models when its safeguards trigger.) At high effort, Luna improves on GPT-5.6 Luna by 5.4 points (14.5% vs. 9.1%) at 58% lower cost per task.
Agents’ Last Exam V1 covers long-horizon, economically valuable computer tasks across 55 sub-industries.

| Model (best effort) | Agents’ Last Exam score | Cost per task |
|---|---|---|
| GPT-6 Astra (max) | 59.3% | $6.23 |
| GPT-6 Sol (max) | 56.4% | $2.93 |
| Claude Opus 5 (high) | 55.9% | $7.29 |
| GPT-5.6 Sol (xhigh) | 53.6% | $5.08 |
| GPT-6 Luna (max) | 50.9% | $0.15 |
| GPT-5.6 Luna (max) | 50.4% | $2.57 |
| Claude Fable 5 w/ Opus 4.8 fallback (xhigh) | 48.7% | $28.55 |
Luna is the bigger story on this eval. It roughly matches GPT-5.6 Luna’s best score at about 6% of the cost per task ($0.15 vs. $2.57), and it gets within 5 points of Claude Opus 5’s best for about 2% of the cost.
Factuality
OpenAI’s factuality eval uses de-identified ChatGPT conversations where users flagged a factual error from an earlier model. OpenAI notes these prompts are harder than typical use, and that its verbosity sweeps showed “almost no dependence on answer length.” Lower is better.
| Model | low | medium | high | xhigh | max |
|---|---|---|---|---|---|
| GPT-6 Astra | 6.3% | 4.4% | 3.9% | 4.0% | 3.9% |
| GPT-6 Sol | 11.4% | 6.9% | 5.1% | 4.5% | 4.6% |
| GPT-5.6 Sol | 19.4% | 15.0% | 10.8% | 8.4% | 8.5% |
| GPT-6 Luna | 27.7% | 17.5% | 12.5% | 10.2% | 7.6% |
| GPT-5.6 Luna | 36.8% | 24.7% | 17.3% | 12.2% | 12.0% |
At each effort level, Sol makes 41% to 54% fewer errors than GPT-5.6 Sol, which supports OpenAI’s “about half as many mistakes” claim. Luna at max (7.6%) beats GPT-5.6 Sol at max (8.5%) for $0.012 per prompt vs. $0.87, about 1/70th of the cost. OpenAI rounds that to “about a hundredth.” No competitor models appear in this chart.
Coding: FrontierCode and DeepSWE
FrontierCode 1.1 Main grades agent-written code on correctness and on “mergeability”: test quality, scope discipline, code style, and adherence to codebase standards.
| Model (best effort) | FrontierCode 1.1 Main | Cost per task |
|---|---|---|
| Claude Opus 5 (medium) | 53.4% | $4.31 |
| GPT-6 Astra (max) | 53.3% | $4.59 |
| Claude Fable 5.1 (medium) | 50.9% | $3.28 |
| GPT-6 Sol (max) | 49.3% | $2.14 |
| Claude Fable 5.1 (xhigh) | 48.7% | $9.27 |
| GPT-5.6 Sol (max) | 47.5% | $5.19 |
| GPT-6 Luna (max) | 42.4% | $0.11 |
| GPT-5.6 Luna (max) | 39.8% | $0.37 |
OpenAI says Sol can “match Claude Fable 5.1 xhigh at much lower cost.” That’s accurate: Sol’s 48.4% at xhigh for $1.37 matches Fable 5.1’s 48.7% at xhigh for $9.27. But xhigh is Fable 5.1’s weakest setting on this eval. Fable 5.1 at medium (50.9%, $3.28) and Opus 5 at medium (53.4%, $4.31) both score higher than any Sol setting.
DeepSWE v1.1 tests original, long-horizon software engineering tasks in real codebases.

| Model (best effort) | DeepSWE v1.1 | Cost per task |
|---|---|---|
| GPT-6 Astra (xhigh) | 74.1% | $4.43 |
| Claude Opus 5 (max) | 73.7% | $11.84 |
| GPT-5.6 Sol (max) | 72.7% | $6.46 |
| Claude Fable 5 (xhigh) | 69.9% | $13.41 |
| GPT-6 Sol (max) | 68.8% | $2.74 |
| GPT-6 Luna (max) | 66.6% | $0.22 |
| GPT-5.6 Luna (max) | 62.2% | $0.53 |
This is the most selectively framed result in the post. OpenAI says Sol at max is “within 1.1 percentage points of Claude Fable 5’s highest score.” That’s true, and Sol does it at about 80% lower cost ($2.74 vs. $13.41). But OpenAI’s own chart shows Claude Opus 5 at 73.7% and GPT-5.6 Sol at 72.7%, both clearly above GPT-6 Sol’s best. The fair reading: GPT-6 Sol gives up about 4 points of top-end DeepSWE performance compared with GPT-5.6 Sol, in exchange for a 58% lower cost per task. At matched cost, Sol wins. Sol at high (65.3%, $0.64) beats GPT-5.6 Sol at medium (61.1%, $1.42) for less than half the price.
Luna’s result holds up better. At max it scores 66.6% for $0.22, close to Opus 5 at medium (68.9%, $3.29) and Fable 5 at medium (65.4%, $6.09). That’s 93% and 96% cheaper per task, exactly as OpenAI says. Note that OpenAI’s charts use Fable 5, not Fable 5.1, for DeepSWE. OpenAI says it substituted Fable 5 wherever Fable 5.1 scores weren’t available.
Computer use: OSWorld 2.0
OpenAI reports partial reward on the offline set of OSWorld 2.0 (v2026.08.08 release). These tasks are long-horizon computer-use workflows covering everyday and professional work.
| Model (best effort) | OSWorld 2.0 offline | Cost per task |
|---|---|---|
| GPT-6 Astra (max) | 73.5% | $9.07 |
| Claude Opus 5 (max) | 70.2% | $24.11 |
| GPT-5.6 Sol (max) | 66.2% | $7.71 |
| GPT-6 Sol (max) | 64.4% | $3.25 |
| Claude Opus 5 (medium) | 60.3% | $12.67 |
| GPT-6 Luna (max) | 52.7% | $0.27 |
| GPT-5.6 Luna (max) | 52.7% | $0.49 |
OpenAI’s comparison is Sol at xhigh (60.5%, $2.21) vs. Opus 5 at medium (60.3%, $12.67), about 83% cheaper. But Opus 5 at xhigh and max (70.1% and 70.2%) and GPT-5.6 Sol at max (66.2%) score higher than any Sol setting. Luna’s top score is flat against its predecessor (52.7% for both) at about 45% lower cost. OpenAI’s other claim checks out: Luna at max beats GPT-5.6 Sol at medium (49.7%) for one tenth of the cost. OpenAI also says Astra “remains the world’s best model for computer use.”
Generation over generation: what actually improved
| Benchmark (max effort) | GPT-5.6 Sol | GPT-6 Sol | Cost change | GPT-5.6 Luna | GPT-6 Luna | Cost change |
|---|---|---|---|---|---|---|
| AutomationBench | 28.8% | 32.0% (33.2% at xhigh) | −49% | 17.0% | 20.7% | −48% |
| Agents’ Last Exam | 52.8% | 56.4% | −59% | 50.4% | 50.9% | −94% |
| FrontierCode 1.1 | 47.5% | 49.3% | −59% | 39.8% | 42.4% | −71% |
| DeepSWE v1.1 | 72.7% | 68.8% | −58% | 62.2% | 66.6% | −59% |
| OSWorld 2.0 offline | 66.2% | 64.4% | −58% | 52.7% | 52.7% | −45% |
| Factual error rate (lower is better) | 8.5% | 4.6% | −79% | 12.0% | 7.6% | −45% |
At max effort, Sol improves on four of six evals and falls back on two (DeepSWE and OSWorld), while its cost per task drops by roughly half or more on every one. Luna improves or holds on all six and gets cheaper on all six. If you run GPT-5.6 Sol at max for coding or computer use and care only about peak score, GPT-6 Sol is not a clear upgrade. If you care about score per dollar, it clearly is.
Alignment and safety evals
OpenAI published five stress-test alignment evals. They are built to provoke failures, so they don’t reflect typical usage rates. Full results are in the system card. Lower is better.
| Eval | GPT-6 Astra | GPT-6 Sol | GPT-5.6 Sol | GPT-6 Luna | GPT-5.6 Luna |
|---|---|---|---|---|---|
| Coding deception | 0.5% | 1.3% | 10.4% | 2.8% | 9.5% |
| Failure to disclose a broken search tool | 1.5% | 4.9% | 77.5% | 28.7% | 78.3% |
| Reviewer bypass attempts | 0.0% | 0.0% | 7.3% | 0.3% | 4.3% |
| Warning circumvention | 17.4% | 64.4% | 68.2% | 42.4% | 76.5% |
| Unauthorized agent interaction | 0.0% | 11.3% | 51.9% | 0.0% | Not measured |
Coding deception, broken-search disclosure and reviewer bypass all show large improvements, which supports OpenAI’s claim of “lower rates of misleading claims about their coding work.” Warning circumvention barely moves for Sol (68.2% → 64.4%). Sol is also well behind Astra (17.4%) and even behind Luna (42.4%) on that eval. Teams running Sol as an autonomous agent near guardrails or safety warnings should keep their own checks in place.
GPT-6 Sol and Luna vs. Claude Opus 5.5, Fable 5.1 and GPT-6 Astra
Anthropic released Claude Opus 5.5 the same day, at $4 / $20 per million tokens. It isn’t in OpenAI’s charts, but two benchmarks can still be compared directly. On AutomationBench and FrontierCode 1.1, Anthropic’s published numbers for Opus 5, Fable 5.1, GPT-6 Astra and GPT-5.6 Sol are identical to the max-effort values in OpenAI’s charts. Both labs appear to be reporting the same runs, so Opus 5.5’s scores can reasonably sit next to Sol’s.
| Benchmark | GPT-6 Sol | GPT-6 Luna | Claude Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra | Claude Opus 5 |
|---|---|---|---|---|---|---|
| AutomationBench | 33.2% | 20.7% | 40.0% | 31.4% | 41.4% | 26.9% |
| FrontierCode 1.1 | 49.3% | 42.4% | 54.4% | 50.3% | 53.3% | 48.0% |
| Terminal-Bench 4.0 | — | — | 66.4% | 55.8% | 57.9% | 52.3% |
| Input / output price (1M) | $2 / $10 | $0.10 / $0.50 | $4 / $20 | $10 / $50 | $10 / $50 | $5 / $25 |
Sol and Luna scores are their best effort setting from OpenAI’s charts. Opus 5.5, Fable 5.1 and Terminal-Bench figures are from Anthropic’s Opus 5.5 announcement. OpenAI didn’t report Terminal-Bench 4.0 for Sol or Luna.
Sol vs. Opus 5.5. Opus 5.5 is the stronger model on these benchmarks. It leads Sol by about 7 points on AutomationBench and 5 on FrontierCode. It also posts a Terminal-Bench 4.0 score (66.4%) that tops even GPT-6 Astra’s 57.9%. Sol’s case is price: half the per-token cost, and much lower cost per task than any Claude model OpenAI measured. Anthropic says Opus 5.5 costs about 40% less than Opus 5 on typical workloads. Even so, Sol’s cost-per-task advantage over Opus 5 is large (11.1x on AutomationBench, roughly 4x on DeepSWE at max), so Sol very likely still costs less per task than Opus 5.5. How much less is unknown until someone runs both on the same harness.
Sol vs. Fable 5.1. Fable 5.1 costs 5x as much per token ($10 / $50). On OpenAI’s data, Sol beats it on AutomationBench (33.2% vs. 31.4%, before Fable’s uncounted fallback costs) and roughly ties its FrontierCode xhigh run. Fable 5.1’s advantage is in areas OpenAI didn’t chart, such as science and hard reasoning. Anthropic reports 52.6% on Terminal-Bench-Science and 65.6% on Humanity’s Last Exam for Fable 5.1.
Sol vs. GPT-6 Astra. Astra costs 5x as much per token and leads on every chart, most clearly on AutomationBench (41.4%) and computer use (73.5%). But Sol at xhigh beats Astra at low on AutomationBench at about a quarter of the cost per task. For most agent pipelines, the sensible default is Sol, with Astra reserved for tasks where the extra accuracy is worth 3–5x the cost.
Luna vs. everything. No Anthropic or Google model with published data on these evals comes close to Luna’s cost per task. Its DeepSWE score (66.6%) is close to Opus 5 at medium, and its per-token price is 1/10th of Claude Haiku 4.5’s.
Price comparison with Claude and Gemini
| Model | Input / 1M | Cached input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | 1.05M |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 1M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 1M |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | 1M |
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | 1.05M |
| Gemini 3.1 Pro Preview | $2.00 ($4.00 over 200K) | $0.20 | $12.00 ($18.00 over 200K) | — |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | 1.05M |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | — |
| Gemini 3.8 Flash | $0.75 (intro) | $0.075 | $3.75 (intro) | — |
| Gemini 3.5 Flash-Lite | $0.30 | — | $2.50 | — |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | 1.05M |
Sol’s list price is now exactly the same as Claude Sonnet 5’s ($2 / $10), and it undercuts Gemini 3.1 Pro Preview on output. Luna is the cheapest frontier-lab model in this table by a wide margin. Gemini 3.8 Flash’s introductory $0.75 / $3.75 price rises to $1.50 / $7.50 on January 1, 2027, which would leave it 15x Luna’s price on output.
What a typical agent session costs
This uses the same example workload as our Opus 5.5 coverage: an agentic session with 2M input tokens (90% served from cache) and 150K output tokens. It’s list pricing only and ignores cache-write premiums and long-context surcharges. It’s a per-token comparison, so it doesn’t capture how many tokens each model actually uses per task.
| Model | Estimated session cost |
|---|---|
| GPT-6 Luna | $0.11 |
| Gemini 3.8 Flash (intro pricing) | $0.85 |
| Claude Haiku 4.5 | $1.13 |
| GPT-6 Sol | $2.26 |
| Claude Sonnet 5 | $2.26 |
| Gemini 3.1 Pro Preview | $2.56 |
| Claude Opus 5.5 | $4.16 |
| GPT-5.6 Sol | $4.52 |
| Claude Opus 5 | $5.65 |
| Claude Fable 5.1 | $9.95 |
| GPT-6 Astra | $11.30 |
Caching and developer changes
Along with the price cuts, OpenAI changed prompt caching for the GPT-6 family:
- Higher default cache hit rates, with the 90% discount on cached input reads.
- Change reasoning effort or available tools mid-conversation without losing the cache. Before this, changing either one could invalidate the cached prefix. That was a hidden cost for agents that switch between cheap and expensive steps.
- Explicit cache breakpoints so developers choose where cached prefixes end.
- A Prompt Caching Dashboard and a diagnostics tool that explains missed caching opportunities.
OpenAI cites GitHub, which reports that these improvements cut the share of prompt tokens needing fresh processing by more than 50% across billions of Copilot requests. OpenAI also gave a figure on its own internal usage: valued at API prices, daily token use exceeds $600 for the median researcher and $7,000 at the 90th percentile.
Availability
- ChatGPT: Available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, rolling out gradually on launch day. Free and Go users can use GPT-6 Luna in the desktop app. OpenAI says the models are “not yet available in Chat.”
- API:
gpt-6-solandgpt-6-lunathrough the Responses and Chat Completions APIs, with Batch, Flex, Fast and regional processing options.
Which model should you use?
| If you need… | Pick | Why |
|---|---|---|
| High-volume classification, extraction, routing, summarization | GPT-6 Luna (medium–high) | $0.10 / $0.50 pricing, and better factuality at high than GPT-5.6 Luna at max |
| Cheap coding sub-agents or bulk code tasks | GPT-6 Luna (max) | 66.6% DeepSWE at $0.22 per task |
| Default model for agents and business automation | GPT-6 Sol (xhigh) | Beats Claude Opus 5 and Fable 5.1 on AutomationBench at a fraction of the cost per task |
| Highest coding scores regardless of cost | Claude Opus 5.5 or GPT-6 Astra | Opus 5.5 leads FrontierCode and Terminal-Bench 4.0; Astra leads DeepSWE in OpenAI’s data |
| Best computer-use reliability | GPT-6 Astra | 73.5% on OSWorld 2.0 offline, vs. 64.4% for Sol |
| Hardest science and reasoning work | GPT-6 Astra, Opus 5.5 or Fable 5.1 | Sol and Luna have no published results in these areas |
| Very long prompts (over 272K tokens) on a budget | Test Claude Sonnet 5 / Opus 5.5 against Sol | OpenAI’s long-context surcharge (2x input, 1.5x output) narrows Sol’s price advantage |
The bottom line
GPT-6 Sol and Luna are an efficiency release. On the evals OpenAI chose, both models give much better results per dollar than GPT-5.6, and Luna in particular moves the low end of the market. It’s close to Claude Opus 5 at medium effort on DeepSWE for about 1/15th of the cost per task. Sol at $2 / $10 now matches Claude Sonnet 5’s price and beats Opus 5 on business automation at a fraction of the cost.
Two caveats belong in any honest summary. First, GPT-6 Sol’s best scores on DeepSWE and OSWorld are lower than GPT-5.6 Sol’s, which OpenAI’s prose doesn’t mention. Second, the Anthropic comparisons use Opus 5 and Fable 5 / 5.1, not the Opus 5.5 that shipped the same day and posts higher scores than Sol on the two benchmarks both labs share. All of these numbers are vendor-reported. Independent results from Artificial Analysis and others will matter more than launch charts, and we’ll update this piece when they arrive.
Sources
- OpenAI: Introducing GPT-6 Sol and Luna (announcement text and chart data)
- OpenAI on X: GPT-6 Sol and Luna announcement
- OpenAI API docs: Models
- OpenAI API docs: GPT-6 Luna
- OpenAI API docs: GPT-6 Astra
- OpenAI API pricing
- OpenAI GPT-6 system card
- Anthropic: Introducing Claude Opus 5.5
- Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Claude Platform Docs: Pricing
- Anthropic: Introducing Claude Sonnet 5
- Google: Gemini Developer API pricing
- Google: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- Kingy AI: Claude Opus 5.5 specs, benchmarks and pricing
- Kingy AI: GPT-5.6 Sol vs Terra vs Luna
Trending on Kingy
Keep reading with the stories getting the most attention now.
-
Claude Opus 5.5 vs GPT-6 Astra, Sol and Luna vs GPT-5.6 Sol: Benchmarks and Real Costs
Read story -
Claude Opus 5.5: Specs, Benchmarks, Pricing and How It Stacks Up Against GPT-6 Astra, Fable 5.1 and Every Frontier Model
Read story -
Grok 4.7 Benchmarks: Specs, Pricing and Frontier Model Comparisons
Read story
