AI News

GPT-6.1 Sol: Specs, Benchmarks, Pricing and the Real Cost per Task

Updated September 29, 2026 at 2:34 p.m. PDT.

OpenAI released GPT-6.1 Sol on September 29, 2026, one week after GPT-6 Sol. It keeps Sol's $2 input and $10 output prices per million tokens, cuts cached-input pricing to $0.10, and adds beta multi-agent support in the Responses API. The release targets coding, computer use and professional work. OpenAI's API changelog confirms the launch and prices.

For anyone paying for an agent to finish a job, the useful comparison is quality at a given task cost. A model can have cheap tokens and still spend heavily on reasoning, repeated tool calls or failed attempts. This guide separates OpenAI's reported benchmark costs from our calculated request examples. Kingy.ai has not run a matched hands-on evaluation of GPT-6.1 Sol for this article.

GPT-6.1 Sol specifications

Specification Published value
API model ID and current snapshot gpt-6.1-sol
Context window 1,050,000 tokens
Maximum input 922,000 tokens
Maximum output 128,000 tokens
Knowledge cutoff April 30, 2026
Native input / output Text and images / text
API reasoning effort Low, Medium (default), High, Xhigh, Max
Unsupported API reasoning settings None and Minimal
Endpoints Responses, Chat Completions and Batch
Core features Streaming, structured outputs, prompt caching
Tool calling Responses API; unavailable through Chat Completions

The model documentation also lists web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search. Image generation is a tool capability; the model's native output modality is text. US and EU data residency are supported, with Fast mode unavailable under EU data residency.

The maximum input and output limits share the context budget. A large context window lets an application supply more material, but it does not guarantee that the model will retrieve every relevant detail correctly. For a document-heavy workflow, test whether the answer cites the right page, table or source passage.

Where it is available

The launch rollout includes Plus, Pro, Business, Enterprise and Edu users in Codex and ChatGPT Work. Free and Go are excluded at launch. GPT-6.1 Sol is available in Work and Codex, and is not yet available in Chat. Enterprise and Edu administrators must enable it; it is off by default for those workspaces. Account, client and workspace settings determine which controls appear. See OpenAI's Codex model and rollout documentation.

Standard and Fast modes are available at launch. GPT-6.1 Sol Ultrafast is still forthcoming. The launch announcement describes an upcoming Codex option with up to eight times faster token generation. That is an announced generation-speed claim, not a measured eightfold reduction in the time needed to finish a job.

OpenAI lists Plus at $20 per month and Pro options at $100, $200 or $500 per month. Business is $20 per user per month billed annually for two or more users, or $25 billed monthly; Enterprise and Edu use negotiated arrangements. OpenAI's plan-pricing page also explains shared Work and Codex usage. These subscription prices do not buy unlimited API calls or establish a fixed cost per included task.

API pricing, including long context

All prices below are USD per million tokens at Standard processing rates. The long-context threshold concerns the request's input tokens.

Token category Up to 272,000 input tokens More than 272,000 input tokens
Uncached input $2.00 $4.00
Cached input $0.10 $0.20
Cache writes $2.50 $5.00
Output, including reasoning $10.00 $15.00

Above 272,000 input tokens, the higher rates apply to the full request. Batch and Flex cost 50% less than Standard; Fast costs twice Standard. Regional processing adds a 10% premium where applicable. These rates and conditions come from OpenAI's pricing documentation.

Discounted processing has operational tradeoffs. Batch is asynchronous with a 24-hour completion window. Flex trades lower prices for slower responses and possible resource unavailability. The Fast billing multiplier is not a guarantee that a complete task finishes twice as quickly.

Tools add charges: web search costs $0.01 per call plus search-content tokens, and Responses API file search costs $0.0025 per call. File storage is $0.10 per GB per day after 1 GB free. A 1 GB hosted shell or code-interpreter container is listed at $0.03 per 20-minute session. Larger containers cost more. These are also listed on the API pricing page.

Compared with GPT-6 Sol, the base input and output rates are unchanged and cached input is half the price. Against GPT-6 Astra, GPT-6.1 Sol's short-context uncached input and output prices are one-fifth as much: $2 versus $10 for input, and $10 versus $50 for output. Those ratios describe rates, not a guaranteed saving on every completed task.

Benchmarks and measured costs per task

The following figures are from the interactive score-versus-cost charts in OpenAI's launch report. Each row compares the stated reasoning setting. Dollars are reported average costs per benchmark task.

Evaluation and effort GPT-6.1 Sol: score; cost/task Comparison: score; cost/task
DeepSWE 1.1, High 75.2%; $0.65 Astra High: 73.2%; $3.92
GDP.pdf, High 32.0%; $0.35 Astra High: 31.0%; $1.79
AutomationBench 1.0.6, Medium 31.7%; $0.19 GPT-6 Sol Medium: 26.9%; $0.21
OSWorld 2.0 offline, Max 71.4% partial reward; $1.27 Astra Max: 73.5%; $9.44
Terminal-Bench Science 0.1, Max 57.0%; $5.47 Astra Max: 68.1%; $23.80
Factuality challenge set, Low 7.7% of answers contain an error GPT-6 Sol Low: 11.4%; lower is better

DeepSWE tests repository engineering; GDP.pdf uses professional PDFs; AutomationBench tests business workflows. OSWorld reports partial reward on its August 8 offline release. Science tasks use code and terminals. The factuality set deliberately selects difficult, previously error-inducing conversations.

On Science at Max, Opus 5.5 with fallbacks scored 63.3% at $23.21 per task. OpenAI evaluated GPT models in its research environment or API and used public reports for competitors. Production prompts and tools can differ.

The table supports a narrower conclusion than a universal ranking. On the displayed coding and PDF configurations, 6.1 Sol slightly exceeds Astra's score at a lower cost. On the displayed computer-use and science configurations, Astra scores higher. The scientific result is an especially clear tradeoff: 6.1 Sol costs about 77% less per attempt than Astra, while scoring 11.1 percentage points lower. Both the saving and the score gap matter.

A benchmark's average cost per task can include unsuccessful attempts. For your own workflow, divide total spending on all attempts by the number of outputs that pass your acceptance checks. Add human repair time separately. That produces a more useful cost per accepted result than dividing the price of one prompt by its token count.

Independent evaluation update

Artificial Analysis now reports results for GPT-6.1 Sol at Max effort: an Intelligence Index v4.3.2 score of 52 and a weighted average cost of $0.72 per index task. The index combines ten evaluations. Its score is an index value, not a percentage of tasks passed.

The evaluator reports 67 million output tokens, including reasoning, across its index run. Its cost methodology uses token counts, provider prices and typical measured cache hits, weighted by each evaluation's index contribution. It includes input, cache reads, cache writes, reasoning and answers. The page also lists output generation at 66.8 tokens per second through the OpenAI API; that throughput does not measure the time to complete an entire agent task.

These are independent Artificial Analysis results. The $0.72 figure is a calculated benchmark cost for its evaluation mix, not a published invoice, and cannot be compared directly with OpenAI's $5.47 scientific-benchmark average. The page does not display a model-specific uncertainty interval for the headline index score.

Calculated costs for example requests

These are illustrative token budgets, not observed runs or promises about what a task requires. Every input token is uncached, and output means all billable output, including hidden reasoning. The estimates assume no cache writes, as with explicit-only caching without breakpoints, and exclude tools, regional premiums, infrastructure and retries. Each row is one request.

Illustrative request Input / billable output Standard Batch or Flex Fast
Short document summary 5,000 / 1,000 $0.020 $0.010 $0.040
Document analysis 25,000 / 5,000 $0.100 $0.050 $0.200
Code review with substantial context 100,000 / 10,000 $0.300 $0.150 $0.600
Large document bundle, long-context rates 500,000 / 20,000 $2.300 $1.150 $4.600

For Standard requests within the short-context band, the uncached calculation is:

Cost = (input tokens × $2 + billable output tokens × $10) / 1,000,000

The 100,000-input example costs $0.20 for input plus $0.10 for output, or $0.30. At 1,000 identical requests, that is $300 before excluded charges. The long-context row uses $4 input and $15 output rates: $2.00 plus $0.30.

Hidden reasoning can change the bill substantially. A response with 5,000 input tokens, 1,000 visible output tokens and 10,000 reasoning tokens costs $0.12 at short-context Standard rates. Counting only the visible answer would produce an incorrect $0.02 estimate. OpenAI confirms that reasoning tokens occupy context and are billed as output.

Use returned usage data to reconcile the bill. A long agent run can contain many requests, and accumulated conversation history can increase the input budget at each step. Any repeated attempt, fallback or subagent work belongs in the total.

What the cheaper cache changes

Repeated context is where the new $0.10 cached-input rate can matter. Consider ten requests sharing a 100,000-token prefix, with 2,000 fresh input tokens and 2,000 billable output tokens per request. Assume explicit-only caching with a breakpoint after the stable prefix: the changing 2,000-token suffix is not written to cache. The first request writes the prefix and all nine later requests receive complete hits.

Calculated short-context Standard example Cost
First request: 100K cache write + 2K new input + 2K output $0.274
Each subsequent request: 100K cache read + 2K new input + 2K output $0.034
Ten requests with those cache assumptions $0.580
Ten requests with every input token uncached $2.240

That is about a 74% reduction in this specific calculation. The cache-write charge replaces the ordinary input rate for those written tokens; it is not added to it. Later hits are not automatic savings you can assume for every prompt. OpenAI's prompt-caching guide explains prefix matching, cache duration and cache usage reporting. Check actual hits before projecting a monthly saving.

Safety evaluations and factual reliability

OpenAI's GPT-6.1 Sol system-card addendum reports a 2.08% failure rate for disclosing a broken search tool, versus 4.92% for GPT-6 Sol. It observed no attempts to bypass an automated safety reviewer. These are finite, deliberately challenging evaluations, not guarantees about production behavior.

Other results show remaining weaknesses. The coding-misrepresentation rate was 1.50%, versus 1.30% for GPT-6 Sol and 0.51% for Astra. In a warning-circumvention test without system controls, unwanted persistence occurred in 23.5% of 6.1 Sol rollouts, versus 17.4% for Astra. The evidence supports improvements in several areas, with exceptions.

Under its Preparedness Framework, OpenAI treats 6.1 Sol as Critical for cybersecurity and High for biological and chemical capability, and applies Astra's safeguards. These are capability classifications, not forecasts of incident rates. The published materials do not give a parameter count or training-compute total.

Multi-agent support and practical adoption

GPT-6.1 Sol supports beta multi-agent orchestration through the Responses API. A root agent can delegate independent work to subagents and combine their results. OpenAI's multi-agent documentation recommends a default of three concurrent subagents and notes that delegation can increase token use. The beta also has limitations, including unsupported reasoning summaries and max_tool_calls.

Parallel work is useful when a coding task has separate investigation, implementation or review steps. It can reduce waiting time, while increasing spending if each agent repeats large amounts of context. Record the whole run's usage and elapsed time rather than treating one subagent's bill as the cost of the job.

Our recommendation is to trial 6.1 Sol on a small set of representative coding, PDF and application tasks with explicit pass criteria. Use the effort and speed setting you intend to deploy. Compare accepted results, total API spending and correction time with your current model. OpenAI's model-selection guidance likewise recommends comparing it with Astra on the same tasks.

For an API migration, replace the model ID deliberately, check that your requested effort is supported, and move tool-using Chat Completions workflows to Responses. Keep a regression case for schema output, tool arguments and any long document task where missing one detail could invalidate the answer. A successful trial should give you a concrete routing rule, such as using 6.1 Sol for work that passes defined checks and escalating unresolved cases to a stronger model.

Dated updates

September 29, 2026, 2:34 p.m. PDT: Added the newly verified Artificial Analysis Max-effort result, weighted cost per index task, output-token total and generation throughput from its primary scorecard. Official specifications and prices remained unchanged in this check.

Reporting method, updated September 29, 2026: Specifications and rates were checked against live OpenAI documentation. Launch benchmarks are OpenAI-reported chart values at labeled effort settings; the independent update uses Artificial Analysis’s primary scorecard. Neither is a Kingy.ai measurement. Request and cache costs are calculated from explicit token assumptions. No paid model runs were commissioned for this article.