Updated September 29, 2026, 2:44 pm PDT
GPT-6.1 Sol makes a strong case for lower-cost agents. Claude Opus 5.5 retains a measurable capability advantage on the current independent composite evaluation and on some demanding maximum-effort workflows. The available evidence supports testing Sol first when inference cost matters, and paying for Opus where its additional capability produces more accepted work. It does not establish one winner for every job.
OpenAI launched GPT-6.1 Sol on September 29, 2026, one week after Anthropic released Claude Opus 5.5. This comparison uses current API documentation, labeled launch results, and Artificial Analysis’s independent evaluation. Kingy.ai has not run a matched hands-on benchmark of these two models for this article. The task-cost examples below are reproducible calculations, not observed performance.
Checked September 29, 2026. All prices are US dollars. Standard API rates, app subscriptions, benchmark costs, and hypothetical task budgets are distinct. Sol’s half-price input/output advantage applies to prompts at or below 272,000 input tokens; its long-context rates narrow that advantage substantially.
Specifications: similar capacity, different constraints
| Specification | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|
| Direct API model ID | gpt-6.1-sol |
claude-opus-5-5 |
| API context window | 1,050,000 tokens | 1,000,000 tokens |
| Normal maximum output | 128,000 tokens | 128,000 tokens |
| Published knowledge cutoff | April 30, 2026 | June 2026 |
| Native input / output | Text and images / text | Text and images / text |
| Reasoning effort | Low, Medium, High, Xhigh, Max | Low, Medium, High, Xhigh, Max |
| API default effort | Medium | Medium; adaptive thinking always on |
Sources: OpenAI’s Sol model specification, Anthropic’s Opus specification, and Claude effort documentation. Opus additionally supports up to 300,000 output tokens through its Message Batches API with the output-300k-2026-03-24 beta header. That exception does not apply to an ordinary interactive request.
Sol’s extra 50,000 context tokens amount to 5% more advertised capacity. Capacity alone cannot establish how well either model retrieves evidence from a large file, notices conflicting details, or completes a workflow. A long-document evaluation matters more for those questions. Likewise, a later knowledge cutoff gives no guarantee that a particular fact is correct. Current prices, product changes, and events still need retrieval and checking.
These are API specifications. A consumer app can impose different limits or compress earlier conversation history. A tool that generates an image also adds a separate capability to the workflow; text-and-image input does not mean the base model natively generates images, speech, or video. The cited model documentation does not disclose enough architecture, parameter-count, or training-compute information to make a substantiated comparison on those measures.
API pricing: Sol is half price until long context changes the bill
The following are per-million-token rates for direct, standard API processing. A cache write is a separate billing category from an ordinary input token or a later cache read.
| Billable category | Sol: input ≤272K | Sol: input >272K | Opus: full context |
|---|---|---|---|
| Fresh input | $2.00 | $4.00 | $4.00 |
| Cache read | $0.10 | $0.20 | $0.20 |
| Cache write | $2.50 | $5.00 | $5.00 for 5 minutes; $8.00 for 1 hour |
| Output, including billable reasoning | $10.00 | $15.00 | $20.00 |
OpenAI’s price schedule applies the long-context rates to the entire request once input exceeds 272K tokens. Anthropic includes Opus’s full 1M context at standard rates. Above Sol’s threshold, fresh input and cache reads cost the same as Opus, while Sol output remains 25% cheaper.
The threshold can create a sharp change in a growing agent conversation. With 272K fresh input tokens and 20K output tokens, Sol costs $0.744. At 273K fresh input and the same output, it costs $1.392. That is an 87% increase for a 1K-token increase in input. Opus’s corresponding cost moves from $1.488 to $1.492. These are calculations from the rate schedules, with other charges excluded.
This makes compaction, retrieval, and the size of tool responses economically relevant. An agent that keeps appending full documents may cross a pricing threshold even when the final answer stays short. Record the actual billable input on each request rather than estimating from the length of the user’s original prompt.
Caching helps both models, with different retention rules
OpenAI documents a 1,024-visible-token minimum and retention of at least 30 minutes since the latest write or reuse for GPT-6.1 Sol. Reuse refreshes retention without another write fee. Opus’s minimum is 512 tokens, with a five-minute default or a higher-priced one-hour option; cache hits refresh the lifetime. These are different retention contracts, so the two write prices should not be treated as identical products.
Keep stable instructions, tool definitions, and reused source material in a consistent prefix. Changing that prefix or waiting beyond retention can turn an expected cache hit into a new write. The input discount can be large, but a workflow dominated by output or repeated new tool results may save much less overall. Confirm cache-read and cache-write usage fields before projecting savings across thousands of tasks.
Batch, faster processing, and tool fees
Sol’s Batch and Flex rates are 50% below Standard; Fast costs twice Standard. Opus’s Batch input/output rates are also 50% lower. Its Fast mode costs $8 input and $40 output per million tokens. OpenAI’s regional processing and Anthropic’s US-only inference add 10% where applicable. Those terms come from the OpenAI and Claude price lists. Choose a processing tier according to the job’s deadline, then compare both models under that tier.
Both providers list web search at $10 per 1,000 searches or calls, in addition to relevant token charges. Ten searches therefore add $0.10 before result-token costs. Storage, hosted execution, and external services can add other charges. A short $0.04 model request with several searches is not a $0.04 completed research job.
Cost per task: transparent examples you can reproduce
For a request with mutually exclusive token categories, calculate:
Token cost = (fresh input × input rate + cache writes × write rate + cache reads × read rate + billable output × output rate) / 1,000,000
Use the correct rate schedule for each request. The examples assume identical billable token counts, standard processing, no extra tools or runtime, and no retries. Ordinary-input rows assume no cache write. Sol’s implicit caching can instead bill an eligible fresh prefix at the write rate; its caching guide explains explicit-only control. Equal text does not necessarily produce equal token counts across providers. Billable output includes reasoning, even when the user sees only a much shorter answer. See OpenAI’s reasoning guide and Claude’s thinking and cost guide.
| Hypothetical usage | Sol | Opus | Sol saving |
|---|---|---|---|
| 10K fresh input + 2K output | $0.04 | $0.08 | 50% |
| 100K fresh input + 20K output | $0.40 | $0.80 | 50% |
| 100K cache read + 10K fresh input + 10K output; initial write excluded | $0.13 | $0.26 | 50% |
| 20 requests: first writes a 100K prefix; next 19 read it; each has 5K output; no other input | $1.44 | $2.88 | 50% |
| 400K fresh input + 20K output | $1.90 | $2.00 | 5% |
| 900K fresh input + 10K output | $3.75 | $3.80 | 1.3% |
The 20-request example assumes every later request hits the cache, uses Opus’s five-minute cache-write price, and includes 100K total output tokens. Sol’s bill is $0.25 for the write, $0.19 for nineteen reads, and $1.00 for output. Opus’s is $0.50 + $0.38 + $2.00. An expired cache, extra instructions, or another tool result changes the bill.
The two long-context rows explain why a 50% token-price headline can mislead a document-heavy buyer. With large fresh inputs and little output, the models’ inference costs approach each other. With shorter prompts or substantial output, Sol has more room to save money.
Cost per accepted task is the production metric
Cost per accepted task = total cost of all attempts, retries, tools, fallbacks, and infrastructure / tasks meeting the acceptance criteria
Include review and repair separately, or convert that labor into money using an explicit rate. A cheaper model that rarely produces a usable result can cost more than a pricier model that passes consistently.
For an intentionally simplified example, imagine model A costs $0.40 per independent attempt and passes 40% of the time, while model B costs $0.80 and passes 90%. Repeating until success gives expected inference costs of $1.00 and about $0.89, respectively. These are invented inputs to illustrate the arithmetic; they are not Sol or Opus results. Real retries can share the same failure mode, have different costs, or never recover, so use measured task outcomes when budgeting.
At exactly half the attempt cost, the cheaper model needs more than half the expensive model’s success rate to win on that simplified metric. Once tool fees, human repairs, or long-context pricing change the cost ratio, recalculate the break-even point. Avoid dividing a composite intelligence score or partial-credit score into a dollar figure and calling it cost per successful task.
Independent evaluation: Opus leads the index, Sol costs less
Artificial Analysis’s live model pages provide a same-evaluator comparison at maximum effort. The Opus configuration includes its default safeguard fallback. The snapshot below was checked on launch day against the pages for GPT-6.1 Sol and Claude Opus 5.5.
| Artificial Analysis metric | Sol, Max | Opus, Max with default fallback |
|---|---|---|
| Intelligence Index v4.3.2 | 52 | 58 |
| Weighted mean cost per index task | $0.72 | $5.98 |
| Total output tokens across index evaluation | 67 million | 260 million |
| Reported output generation speed | 66.8 tokens/second | 92.5 tokens/second |
Opus leads by six index points. Sol’s reported cost is about 88% lower on this evaluator’s weighted task mix. The index is a composite score, not a percentage of real-world tasks completed. The output totals belong to that evaluation run; they do not predict how many tokens either model will use on your next coding request.
At default Medium effort, the same evaluator reports Sol at 48 and $0.21 per index task, versus Opus with default fallback at 51 and $1.34. Sol is about 84% cheaper in that snapshot, with a three-point capability gap. Moving either model to Max raises both its index result and task cost. A deployment that uses Medium should start with the Medium evidence.
The index methodology combines ten evaluations, with weights of 30% for agents, 20% for coding, 20% for scientific reasoning, and 30% for general capability. It is primarily an English-language, text-based suite; multilingual and other modalities are evaluated separately. Cost reporting uses provider-reported token usage where available and incorporates measured typical cache-hit rates. It should not be read as a quote for a particular deployment.
This evidence favors Opus for the broader capability mix and Sol for inference economy. It also gives an actual counterexample to a blanket claim that Sol must generate tokens faster: Opus’s reported output rate is higher in this snapshot. End-to-end task time still depends on reasoning duration, token volume, tool execution, and retries. A faster stream can take longer to finish if the model produces more output.
Launch benchmarks: the advantage changes with effort and task
Selected OpenAI launch-chart results follow. Costs are per attempted task; Opus includes fallbacks. Competitor results come from public reports, so these are not controlled Kingy.ai tests.
| Evaluation / effort | Sol score / cost | Opus score / cost |
|---|---|---|
| GDP.pdf, Medium | 30.0% / $0.34 | 25.6% / $0.80 |
| GDP.pdf, High | 32.0% / $0.35 | 28.8% / $0.83 |
| AutomationBench 1.0.6, Medium | 31.7% / $0.19 | 29.5% / $0.65 |
| AutomationBench 1.0.6, Max | 36.1% / $0.30 | 42.5% / $1.44 |
| Terminal-Bench Science 0.1, Max | 57.0% / $5.47 | 63.3% / $23.21 |
Sol leads the selected PDF rows and Medium automation point; Opus leads Max automation and science. Sol costs less in every listed configuration. Equal effort labels do not imply equal compute budgets.
These are point estimates. The charts do not establish that every score difference is statistically significant. A benchmark’s success rate also need not transfer to your task set.
GDP.pdf tests answers grounded in complex professional PDFs. AutomationBench tests multi-step business workflows. Terminal-Bench Science tests scientific workflows through terminal tools, including analysis and simulation. Those are different abilities. A model can be economical at extracting facts from a document and still need more help with an open-ended agent task.
Coding and computer-use scores need extra care
The checked DeepSWE v1.1 leaderboard does not provide a matched Sol 6.1 versus Opus 5.5 pair. Anthropic’s Opus launch table reports 66.4% on Terminal-Bench 4.0 at Xhigh and 54.4% on FrontierCode v1.1 Main at Max, with production safeguards enabled. Neither score can be ranked against a DeepSWE result.
OpenAI reports OSWorld 2.0’s offline set from the August 8 release with partial reward. Anthropic reports OSWorld 2.1 partial-credit results. Different versions and subsets prevent a clean computer-use head-to-head. A partial reward also measures something different from a task whose final state either passes or fails. For workflows such as editing a spreadsheet and sending the resulting file, define which final actions must all succeed before giving credit.
Across benchmarks, keep the evaluator, dataset version, agent harness, tool access, effort, token limit, and fallback policy attached to each number. A coding agent with a strong search tool, useful repository instructions, and a verifier is a system. Its score does not isolate the underlying model’s contribution.
Reliability, safeguards, and fallback behavior
OpenAI’s system-card addendum documents Sol’s safety and alignment evaluations. It treats the model as Critical for cybersecurity capability and High for biological/chemical capability under OpenAI’s framework, and applies Astra’s safeguard stack. These are capability classifications under that provider’s framework, not a comparison showing which model is safer for an ordinary company workflow.
Anthropic’s Opus system card is the corresponding source for its safeguards and evaluations. When a benchmark includes a fallback, the measured system may involve a different model on some tasks. Keep that policy explicit, particularly for security work. A fallback-assisted result answers a different deployment question from a result that counts every safeguard intervention as a failed attempt.
Neither provider’s safety tests establish that an autonomous agent can run without checks in your environment. Evaluate ordinary permission handling, prompt injection through files or webpages, source accuracy, and whether the agent reports failed tools honestly. Judge the completed artifact and the actions taken, not merely the fluency of the final response.
Availability and subscriptions: compare the product you will use
Sol is available at launch through the OpenAI API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu. OpenAI says it is not yet available in Chat. OpenAI’s launch resources specify that Sol is off by default in Enterprise and Edu workspaces until an admin enables it. Opus is available on paid Claude plans and through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Sources: the Sol launch announcement and Opus documentation.
ChatGPT Plus is $20 per month. The current Pro tier documentation lists $100, $200, and $500 monthly plans. Claude Pro is $20 monthly or $200 billed annually; Claude Max is $100 or $200 per month. These are app plans with usage conditions, not prepaid quantities of API tokens. API billing is separate.
A subscription’s cost per task can be measured after you know how many acceptable tasks it completed in a billing month. It cannot be inferred reliably by dividing the price by an assumed message allowance. Long coding sessions, short chats, and tool-heavy research do different amounts of work. Count completed jobs and record whether usage limits prevented finishing them.
Migration work also belongs in the cost comparison. Sol’s documentation directs tool-using applications to Responses; Chat Completions is supported without tool calling. Opus’s migration guide documents always-on thinking, errors for forced tool use, model/conversation-bound thinking blocks, and changes to supported computer-use tooling. Test streaming progress, structured output, tool invocation, context handling, and failure recovery before replacing a production model ID.
OpenAI’s September 29 API changelog also confirms Sol’s beta multi-agent support in Responses. The model can coordinate subagents in parallel. Its multi-agent documentation warns that adding subagents can increase token usage. For a delegated workflow, compare total spend and completion time for the entire run; a lower per-token price alone does not establish that parallel delegation saves money or time.
Which model is the better fit?
| Your priority | Evidence-based starting point | What could change the choice |
|---|---|---|
| Frequent tasks under 272K input, with a clear acceptance check | Test Sol first; its base token rates give substantial budget room. | More retries, tool calls, or repair work can consume the saving. |
| Highest capability on the current independent index | Opus leads the reported Max-effort composite. | Your task mix may differ from the index’s weights and grading. |
| Difficult science or maximum-effort business automation | Opus has higher reported scores in the selected launch configurations. | A fixed budget may buy more acceptable work from Sol. |
| Large fresh document inputs with short output | Compare both; long-context inference prices can be close. | Retrieval, caching, and source-grounding quality may dominate. |
| Existing Codex or Claude workflow | Start inside the working integration and measure a candidate switch. | Migration effort and tool differences may outweigh modest inference savings. |
These are starting points inferred from the cited evidence, not hands-on endorsements. For an existing Opus deployment, price an actual Sol replacement using the same task inputs and completion criteria. For a new Sol deployment, keep a small set of tasks that failed and test whether Opus solves enough of them to justify routing those cases to it. Repeating the same unsolved attempt indefinitely is a poor substitute for a defined escalation rule.
A useful local evaluation should retain these items:
- A representative task set, including difficult and failure-prone jobs, with acceptance criteria fixed before the runs.
- Exact model IDs, effort, processing tier, tools, token limits, and fallback policy. Record provider updates when aliases change.
- Every attempt’s fresh input, writes, reads, output, tool charges, elapsed time, and final artifact. Keep failed attempts in the totals.
- Objective checks for numbers, citations, code behavior, and required file outputs, with blinded human review where judgment is needed.
- Accepted-task rate, total cost per accepted task, completion-time distribution, and repair time. Compare at your quality bar and budget.
For related context, see Kingy.ai’s Sonnet 5.5 versus Opus 5.5 comparison and Astra versus Opus 5.5 comparison. This article specifically covers the new GPT-6.1 Sol, whose prices and results should be kept distinct from GPT-6 Sol.
Editorial method: API specifications and rate schedules were checked against first-party documentation. Vendor charts and the independent Artificial Analysis snapshot are labeled separately. Calculated examples expose their assumptions and exclude unmeasured quality claims. No paid inference tests were performed for this article. Launch-day evaluations can be revised; compare future updates within the same benchmark version and configuration.
Updates
September 29, 2026, 2:44 pm PDT: Clarified the Enterprise and Edu admin enablement requirement and added Sol’s beta Responses multi-agent support with its token-use caveat. Sources: OpenAI launch resources, the September 29 API changelog, and multi-agent guidance.
Trending on Kingy
Keep reading with the stories getting the most attention now.
