Gemini 4 Argon arrives at a crowded frontier. Google announced it on September 30, one day after OpenAI released GPT-6.1 Sol and two days after Anthropic released Claude Sonnet 5.5. The useful comparison now spans Argon, GPT-6 Astra, GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5 and Fable 5.1, alongside Grok, DeepSeek, Qwen, Kimi, Meta, Xiaomi and Mistral.
Our buying advice is to shortlist by workload. GPT-6.1 Sol and Sonnet 5.5 deserve early tests for frequent professional work. Astra and Opus 5.5 deserve a place on difficult coding, research and agent tasks. Argon has promising launch results, especially in knowledge work, video understanding and long-context reasoning, but its initial access is restricted. Lower-cost models deserve the same acceptance tests before you spend premium-model money on every request.
This comparison covers public specifications, access, token prices, caching, reasoning, benchmark results, independent evaluations, speed, safety and deployment choices. It is a sourced analysis, not a Kingy.ai hands-on benchmark. We have not run these models on a matched private task set, and we do not infer missing scores or undisclosed architecture details.
Research cutoff: September 30, 2026. All prices are USD before tax unless stated otherwise. API rates describe billable usage; subscriptions have separate allowances. Argon is a launch-day announcement, so several API details remain unpublished. “Not reported” means the evidence is missing, not that a model scored zero.
The model names and what you can use today
The official Google name is Gemini 4 Argon. Google says access starts with trusted cyber defenders through Fairwind, followed by a broader rollout beginning with paid API customers and Google AI Ultra subscribers. The announcement gives no firm date for general availability. Treat its prices as announced introductory rates, rather than evidence that every developer can send a production request today.
GPT-6.1 Sol is available through the OpenAI API and is rolling out to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It was not available in ordinary Chat at launch. A subscription, client version and administrator setting can therefore change what you see in the picker. OpenAI’s product documentation also distinguishes app controls from API settings: Ultra uses subagents, while Max increases the reasoning available to a single task.
Anthropic’s current lineup includes Opus 5.5, Sonnet 5.5 and Fable 5.1. Fable remains relevant even though Opus and Sonnet have newer release dates. Invitation-only Claude Mythos 5.1 has a different access boundary. Haiku 5.5 was announced as forthcoming; it is not a released model in this snapshot. Naming a product in a launch post does not establish open API access.
Specifications: context, output, modalities and reasoning
| Model / API ID | Context or input limit | Output limit | Native input → output | Reasoning / default |
|---|---|---|---|---|
| Gemini 4 Argon Public API ID not confirmed |
Input limit not specified in the launch announcement | 1M tokens announced | Multimodal launch evaluations; complete API contract pending | Highest thinking used in most launch evals; public API controls pending |
GPT-6 Astragpt-6-astra |
1,050,000 context; 922,000 max input | 128,000 | Text, image → text | Low, medium, high, xhigh, max |
GPT-6.1 Solgpt-6.1-sol |
1,050,000 context; 922,000 max input | 128,000 | Text, image → text | Low through max; medium default |
GPT-6 Lunagpt-6-luna |
1,050,000 context; 922,000 max input | 128,000 | Text, image → text | None, low through max; medium default |
Claude Opus 5.5claude-opus-5-5 |
1M context | 128,000 standard | Text, image → text | Adaptive, always on; medium API default |
Claude Sonnet 5.5claude-sonnet-5-5 |
1M context | 128,000 standard | Text, image → text | Adaptive; high API default |
Claude Fable 5.1claude-fable-5-1 |
1M context | 128,000 standard | Text, image → text | Adaptive, always on; high API default |
Gemini 3.8 Flashgemini-3.8-flash |
1,048,576 input | 65,536 | Text, image, video, audio, PDF → text | Low, medium, high |
Gemini 3.1 Pro Previewgemini-3.1-pro-preview |
1,048,576 input | 65,536 | Text, image, video, audio, PDF → text | Thinking supported |
Sources: the official model pages for Astra, Sol 6.1, Luna, Opus, Sonnet, Fable, Flash and Pro Preview, plus Argon’s announcement.
Published knowledge cutoffs are April 30, 2026 for Astra and Sol 6.1, May 18 for Luna, and June 2026 for these three Claude models, per their model pages. A cutoff describes training knowledge, not the freshness of an answer retrieved through search. We have not confirmed Argon’s cutoff in its launch announcement.
Context is the information a request can hold. Maximum input and maximum output are different constraints, and providers describe them differently. A million-token context window does not promise a million-token answer. Anthropic’s thinking and answer tokens share the output budget; its 300K-output beta is for Message Batches, not a blanket increase for ordinary interactive requests. Google’s Argon announcement describes a million-token output limit as room for long reasoning trajectories. We would wait for the API contract before designing a workflow around a million tokens of user-visible text.
The same caution applies to modalities. A text-output model can call an image-generation tool without natively generating images. An assistant can transcribe audio before passing text to its core model. Gemini 3.8 Flash’s native video/audio input makes it useful to evaluate those workflows directly; an Astra or Claude agent may use a different preprocessing pipeline. Keep the pipeline in the comparison.
The cited closed-model documentation does not disclose enough to compare parameter counts, training FLOPs, GPU inventories or exact architectural layouts across the leaders. We leave those entries unknown. A guessed trillion-parameter label would add apparent precision without helping anyone choose a model.
API token prices, caching and long-prompt rates
The following rates are per million tokens. “Fresh input” excludes cache reads. The output column can include billable reasoning tokens. These are standard direct-provider rates or announced rates, without tools, regional premiums, storage, tax or negotiated enterprise discounts.
| Model | Fresh input | Cache read | Output | Condition |
|---|---|---|---|---|
| Gemini 4 Argon, announced | $2 intro; $4 later | $0.10 intro | $10 intro; $20 later | Intro end date unspecified |
| GPT-6 Astra | $10 | $1 | $50 | ≤272K input |
| GPT-6.1 Sol | $2 | $0.10 | $10 | ≤272K input |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | ≤272K input |
| Claude Opus 5.5 | $4 | $0.20 | $20 | Full 1M context standard |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 | Full 1M context standard |
| Claude Fable 5.1 | $10 | $0.25 | $50 | Full 1M context standard |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | Through Dec 31, 2026 |
| Gemini 3.1 Pro Preview | $2 | $0.20 | $12 | ≤200K prompt |
| Grok 4.7 | $2 | $0.50 | $6 | Below 200K prompt |
| Qwen3.8-Max-0902 | $2 | $0.25 implicit | $6 | Explicit cache differs |
| DeepSeek V4.1-Flash | $0.30 peak | $0.006 peak | $1.20 peak | Offpeak rates are half |
| DeepSeek V4-Pro-0813 | $1.32 peak | $0.044 peak | $3.96 peak | Offpeak rates are half |
| Kimi K3 | $3 | $0.30 | $15 | Launch-listed direct rates |
| GLM-5.3, Z.ai direct | $1.40 | $0.26 | $4.40 | Hosted provider cache rates differ |
| Muse Spark 1.3 Standard | $1.25 | $0.15 | $4.25 | Contributor has different terms |
| Mistral Medium 3.5 | $1.50 | $0.15 | $7.50 | 256K context |
| Mistral Large 3 | $0.50 | $0.05 | $1.50 | 256K context |
| MiniMax M3 | $0.30 | $0.06 | $1.20 | ≤512K input |
| MiMo-V2.6-Pro | $0.435 | $0.0036 | $0.87 | Overseas rate |
| MiMo-V2.6-Flash | $0.14 | $0.0028 | $0.28 | Overseas rate |
| Step 5 Preview | $1 | $0.05 | $2.70 | Global direct API |
Additional price sources: Grok, Qwen, DeepSeek, Kimi, Z.ai, Meta, Mistral, MiniMax, MiMo Pro, MiMo Flash and StepFun. Exact billing conditions belong to the provider contract, not an aggregator’s summary.
OpenAI’s pricing schedule charges Astra and Sol 6.1 twice the short-context input/cache rates and 1.5 times the output rate when input exceeds 272K tokens, for the whole request. Batch and Flex are half Standard; Fast is twice Standard. Astra Ultrafast is six times Standard. Those choices change the price before a single retry.
Anthropic offers the full 1M context at standard rates on its current models. Batch input/output receives a 50% discount. Five-minute cache writes cost 1.25 times fresh input: $5 on Opus, $2.50 on Sonnet and $12.50 on Fable. One-hour writes cost $8, $4 and $20 respectively. A cheap read has value only when the request gets a cache hit; include the original write and new context in your bill.
Google’s current Gemini pricing has a 200K prompt threshold for Pro 3.1: $2/$12 becomes $4/$18 above it, with cache reads rising from $0.20 to $0.40. Flash 3.8’s $0.75/$3.75 and $0.075 cache rate lasts through December 31, 2026; January rates double. Cache storage also costs money. Google’s table does not establish Argon’s eventual batch, storage, tool or long-prompt billing rules.
Worked cost examples and the price of a successful result
Consider one hypothetical request with 100,000 fresh input tokens and 20,000 billable output tokens. The arithmetic is input tokens ÷ 1,000,000 × input rate, plus output tokens ÷ 1,000,000 × output rate. The figures below are calculations from the published rates, not measurements of how many tokens a model needs to finish a task.
| Model / condition | 100K fresh input + 20K output | 100K cache read + 20K output |
|---|---|---|
| Astra, short context | $2.00 | $1.10 |
| Sol 6.1, short context | $0.40 | $0.21 |
| Opus 5.5 | $0.80 | $0.42 |
| Sonnet 5.5 | $0.40 | $0.22 |
| Argon, introductory announced | $0.40 | $0.21 |
| Argon, later announced | $0.80 | Not calculated |
| Grok 4.7, short prompt | $0.32 | $0.17 |
| DeepSeek V4.1-Flash, peak | $0.054 | $0.0246 |
| MiMo-V2.6-Pro | $0.0609 | $0.01776 |
| Muse Spark 1.3 Standard | $0.21 | $0.10 |
Cache scenarios exclude the original write, any storage charges, tools, other billable input and retries. Argon’s later cache price is left unconfirmed here. The token budgets are illustrative; providers tokenize content differently, and hidden reasoning can add billed output beyond the visible answer.
At those fixed token counts, Astra costs five times Sol 6.1 and Opus costs twice Sonnet. But models choose different numbers of steps, tools and reasoning tokens. An inexpensive request that fails, retries and still needs human repair can become the expensive workflow. A premium model that finishes once can justify its higher rate.
A long-document example exposes another difference. At 500,000 fresh input tokens and 20,000 output tokens, Astra’s long rates yield $11.50, Sol 6.1’s $2.30, Opus’s standard rates $2.40 and Sonnet’s $1.20. Pro 3.1 yields $2.36. These are still equal-token scenarios, not equal-quality tasks. Retrieval can reduce the prompt, but a poor retrieval step can omit the evidence needed for the answer.
Track cost per attempted task and cost per accepted task separately. Suppose a model costs $0.40 per attempt and succeeds on 80% of attempts. Dividing $0.40 by 0.80 gives a simple $0.50 per success estimate. That assumes comparable attempts and excludes reviewer time. A more useful production calculation sums all model, tool and infrastructure costs plus human repair, then divides by the number of outputs that meet the acceptance criteria. Never count an unreviewed answer as a success merely because the API returned HTTP 200.
Subscriptions answer a different purchasing question. OpenAI lists Plus at $20/month and Pro tiers at $100, $200 and $500. Work and Codex share usage; tasks consume varying amounts depending on context, tools and model. Fast can use included allowance at 2.5 times Standard, while Astra Ultrafast uses it at eight times. Those are usage multipliers, not speed promises or fixed task counts. Compare an app subscription with the app experience you need, and an API with the metered workload you intend to run.
Gemini 4 Argon’s launch benchmarks against Astra, Opus and Fable
The table selects ten rows from Google’s published comparison. Entries are percentages. Google’s runs and other providers’ or public leaderboard results are mixed; Sol 6.1 and Sonnet 5.5 are absent.
| Evaluation / score | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|
| AutomationBench | 51.3 | 41.4 | 42.5 | 31.4 |
| DeepSWE v1.1 | 77.9 | 74.1 | 74.2 | 67.4 |
| FrontierSWE v2 | 55.0 | 65.5 | 62.3 | 56.3 |
| Terminal-Bench 4.0 | 57.4 | 58.2 | 66.4 | 57.9 |
| Terminal-Bench Science 0.1 | 57.6 | 68.1 | 63.3 | 52.6 |
| LABBench 2 | 88.8 | 85.4 | 73.1 | 68.6 |
| GraphWalks 256K–1M, BFS F1 | 84.2 | 71.8 | 66.8 | 65.0 |
| OSWorld 2.0 offline, partial score | 69.2 | 72.6 | Not reported | Not reported |
| LVBench video understanding | 91.7 | 87.5 | 83.7 | 79.7 |
| CWE-bench v1 | 68.0 | 68.0 | 67.0 | 58.0 |
Argon merits tests on multimodal evidence and long documents. Opus leads this Terminal-Bench comparison; Astra leads FrontierSWE and Science. Those differences argue for separate workload tests.
Google’s evaluation PDF specifies highest thinking and generally single attempts. Argon’s DeepSWE uses mini-swe; its Science run has six times the verifier timeout. OSWorld reports the maximum of three runs on the offline subset. LVBench uses unequal vision budgets: Gemini at one frame per second, Astra 800 frames, Opus 600 and Fable 300. These differences limit a causal claim that the underlying model alone explains every gap. The PDF labels its results “as of October, 2026”; we identify them by the September 30 launch rather than silently changing their date.
Independent evaluations: four different views of the frontier
Artificial Analysis measures a weighted task mix
Artificial Analysis currently uses Intelligence Index v4.3.2. It is mainly an English/text capability measure, with vision, speech and multilingual performance assessed separately. Its aggregate uncertainty estimate is below about one point, while individual evaluations can be less precise. A five-point composite lead and a one-point lead deserve different interpretations.
The v4.3 weighting allocates 30% to agents, 20% to coding, 30% to general capability and 20% to scientific reasoning. That is an explicit definition of the composite. A company whose work is 80% visual document review should not copy those weights into its purchasing decision.
| Evaluation | Opus 5.5 Max + fallback | Sonnet 5.5 Max + fallback | Astra Max | Sol 6.1 Max |
|---|---|---|---|---|
| Intelligence Index v4.3.2 | 58 | 56 | 53 | 52 |
| GDPval-AA v2.1, Elo | 1,846 | 1,844 | 1,542 | 1,575 |
| AA-Briefcase v1.1, Elo | 1,822 | 1,811 | 1,569 | 1,564 |
| AutomationBench-AA | 70% | 71% | 68% | 65% |
| Terminal-Bench 4.0, AA implementation | 60% | 64% | 59% | 56% |
| Humanity’s Last Exam | 61% | 55% | 55% | 53% |
| GDP.pdf, all-pass | 26% | 26% | 31% | 31% |
| AA-LCR v1.1 | 85% | 83% | 81% | 83% |
Sources for the component snapshot: AA’s Opus/Astra comparison, Sonnet comparison page and Sol Max comparison page. The values here consistently select Max configurations; the comparison-page URLs can name a different effort for the other model. SciCode and CritPt were marked under review and are omitted from this selected table.
Opus leads the composite among these four, yet Astra and Sol lead the GDP.pdf all-pass result. Sonnet leads AA’s Terminal-Bench implementation. This is precisely why a single index should start a shortlist rather than finish it. We did not find an Argon Intelligence Index result in the retrieved snapshot.
Historical AA numbers can use another index version. The v4.2 revision added Briefcase and GDP.pdf, removed saturated GPQA Diamond and changed long-context testing. Astra’s older launch-page score on v4.1.1 belongs to that older index. Placing it beside a v4.3.2 score would create a ranking that the evaluator never measured.
Arena measures human preference
In Arena’s September 30 Text snapshot, Gemini 4 Argon High ranks first at 1,525 ±9 with 4,942 votes. Opus 5.5 High has 1,504 ±10; Fable 5.1 Max 1,501 ±7; Muse Spark 1.3 Max 1,495 ±6; Kimi K3 Max 1,488 ±5; Qwen3.8 Max 1,481 ±5; Astra Max 1,476 ±7. Sol 6.1 and Sonnet 5.5 were absent from the retrieved table. These scores describe pairwise preference, not percent correct.
Arena reports uncertainty and rank spreads because neighbouring raw positions can overstate the evidence. Differences in style, helpfulness, verbosity and prompt mix affect preference. A model selected by users for a readable answer can still fail a tool workflow whose success is checked mechanically.
Agent Arena measures different signals from randomized agent configurations and interaction traces. Its net-improvement measure is relative to the average orchestrator, not a solved-task percentage. The September 29 Agent table lists Fable Max at 14.58% ±1.85 net improvement, Astra Max at 11.78% ±2.17 and Opus 5.5 High at 11.15% ±1.73. These intervals overlap for Astra and Opus. Sol in that table is the older GPT-6 Sol, not 6.1. We retain its separate date and metric.
Epoch combines capability evidence and audits benchmark quality
Epoch’s model page gives Opus 5.5 an ECI of 167 and first place; Astra also rounds to 167 and ranks second. Its administered tests show different strengths: Opus leads MirrorCode, while Astra leads the listed FrontierMath and Mystery Game Puzzles results. Rounded equality on a composite does not imply equal performance on each problem.
The Epoch Capabilities Index is a latent scale built from overlapping benchmarks. Its inputs can include Epoch-run, benchmark-author and model-developer results. It is not a percentage accuracy score, and refitting can revise historical numbers. That source mix matters when someone labels an entire ECI ranking “independent testing.”
Vals shows why token price does not settle task cost
| Model | Index score | Cost / test | Cost basis caveat |
|---|---|---|---|
| Gemini 4 Argon | 68.90% | $15.68 | $4/$20 rates, not introductory |
| Claude Sonnet 5.5 | 67.04% | $21.34 | Max; fallbacks included |
| Claude Opus 5.5 | 66.97% | $32.14 | Max; fallbacks included |
| Claude Fable 5.1 | 65.83% | $28.71 | Fallback policy applies |
| GPT-6 Astra | 63.13% | $18.46 | Vals suite |
| GPT-6.1 Sol | 61.15% | $3.24 | Vals suite |
Vals Index v2.1, updated September 29 and retrieved September 30, includes tax and Terminal-Bench 4.0 after its September 25 revision. Its task mix therefore differs from the dated launch prose still present on model pages. Argon’s test cost uses $4/$20 token rates; we preserve that basis rather than silently repricing it at Google’s introductory rate. The table supplies another useful view of capability and task cost, not a measurement of GDP impact.
On September 30, Vals’ Sonnet live card showed 67.04% ±0.92 and $21.34 per test. Its Opus live card showed 66.97% ±0.89 and $32.14. The score gap is far smaller than the stated uncertainty. Sonnet is cheaper on this suite. The pages’ dated launch prose reports different figures, so we use the live cards consistently.
Vals discloses Max effort and fallback models. A fallback-assisted system can complete a task that the named model would refuse. The reported success therefore needs the fallback policy alongside it. AA’s Max-effort task-cost snapshot gives the opposite Sonnet/Opus cost ordering, as the next section shows. Different workloads, reasoning lengths and scoring rules can produce both results without contradiction.
Speed, latency and reasoning effort
| AA configuration | Output tokens/second | Time to first token | Weighted cost / Index task |
|---|---|---|---|
| Opus 5.5 Max, default fallback | About 92 | About 703 seconds | $5.98 |
| Sonnet 5.5 Max, default fallback | About 139 | About 441 seconds | $7.62 |
| GPT-6 Astra Max | About 51 | About 320 seconds | $3.26 |
| GPT-6.1 Sol Max | About 66 | About 273 seconds | $0.72 |
Source snapshot: AA’s individual pages for Opus, Sonnet, Astra and Sol. The latency figures come from AA’s separate performance workload; they are not the duration of an Index task and are not an ordinary-app response-time promise. Measurements fluctuate, so we round them.
AA’s performance methodology uses a default 10K-input workload with at least 1,500 answer tokens, generally reported as 72-hour medians. Throughput excludes the initial wait; first-token latency can include reasoning. The observations originate from a US-central server, so network location also matters. Measure end-to-end task time separately.
Sonnet streams tokens faster here, but Opus costs less per Index task. AA reports about 410 million total output tokens in Sonnet’s Index run versus 260 million for Opus. Astra and Sol use far fewer, about 60 million and 67 million respectively. Those are totals across the evaluator’s run, not tokens per user prompt. They show how token consumption can overwhelm the base-price comparison.
AA’s Sol effort sweep reports Index scores of 42, 48, 50, 51 and 52 at Low, Medium, High, Xhigh and Max. The corresponding costs per task are $0.13, $0.21, $0.32, $0.39 and $0.72. Medium to Max buys four index points at more than three times the measured cost. Whether that buys value depends on the failures it prevents in your own workload.
Anthropic’s effort setting is a behavioural signal, not a strict compute budget. Provider labels do not calibrate equal compute across models. Max versus Max is useful documentation, but it does not equalize dollars, seconds or tokens. Run an effort sweep before assuming the largest setting is always appropriate.
OpenAI Fast mode can reduce latency for regular traffic, with ramp-rate limits and model/region constraints. Astra Fast has no latency SLA. Astra Ultrafast supports US/global processing, with persistent WebSockets recommended to reduce overhead in tool-heavy agents. Sol 6.1 Ultrafast was described as coming later; it is not a launch-day feature to price into a production promise.
Which models belong on your shortlist?
Gemini 4 Argon
If you obtain access, build an Argon pilot around concrete questions on your own videos and documents. Save the endpoint, thinking setting and token usage. Grade answers against the original evidence, and repeat difficult cases. These are proposed tests, not reported performance.
GPT-6 Astra and GPT-6.1 Sol
Astra belongs on difficult end-to-end tasks where the cost of a miss exceeds the token premium. OpenAI’s launch results include 97.6% on FrontierMath Tier 4 v2 and 99.9% on ARC-AGI-3, with maximum scores across effort settings and research/API conditions. Near-saturated scores leave little room to separate future models; they do not establish comparable mastery of every scientific problem.
Sol 6.1 is a sensible first candidate for repeated coding and professional work, with Astra kept as an escalation. In OpenAI’s Science 0.1 testing, Sol Max costs $5.47 per task versus $23.21 for Opus and $23.80 for Astra; Astra achieves the strongest tested score, 68.1%. These are vendor-reported benchmark costs. Their useful implication is to test whether Sol clears your quality floor before paying for Astra on every attempt.
The GPT-6 API supports asynchronous tool calls, mid-turn steering and reasoning changes that preserve the prompt prefix for caching. These can matter more to a production agent than a small benchmark gap. The application still runs tools and manages pending work. Astra and Sol 6.1 require Responses for tool calling; adopting a model ID without checking your API path can break an integration.
Claude Opus 5.5, Sonnet 5.5 and Fable 5.1
Opus is the strongest Claude default to test for sustained judgment across coding and knowledge work. Sonnet gives you half Opus’s fresh-input/output price and faster measured streaming. Fable remains a premium option when higher-effort Opus falls short, rather than an automatic replacement for a cheaper model that already succeeds.
Anthropic’s Sonnet launch illustrates the effort trap: Sonnet’s headline Terminal-Bench 4.0 score is 70.6% at Max, compared with Opus’s 66.4% at Xhigh. The Sonnet system card reports FrontierCode 1.1 Main at 46.2% at Max and 52.1% at Xhigh, with additional review sometimes timing out or exceeding scope. More reasoning can hurt a constrained task. The same launch flags a since-fixed structured-output bug in prerelease AA knowledge-work testing, expected to produce a small understatement.
Opus’s launch table reports 67.7% on Humanity’s Last Exam with tools and 81.8% on OSWorld 2.1 partial credit. Those are different protocols from Google’s OSWorld 2.0 offline comparison. We keep the numbers in their own source context instead of treating “OSWorld” as one interchangeable test.
Grok 4.7
Grok 4.7 is the current model to compare, rather than a cached Grok 4.6 listing. It takes text/images and produces text, with function calls and structured outputs. Its price doubles at 200K prompt tokens or more for input, cache reads and output. Web search and code execution add $5 per 1,000 calls; X Search has per-item billing. Tool-heavy research needs those costs in the calculation.
AA measures Grok Xhigh at an Index of 46 and $3.74 per task, while High also rounds to 46 at $2.73. That is a useful reason to test the default before increasing effort. Grok’s search integrations can be relevant to a live-information workflow, but a model score alone cannot evaluate the quality of retrieved evidence.
DeepSeek V4.1-Flash and V4-Pro-0813
The current DeepSeek contract lists separate Flash and Pro endpoints, despite an earlier announced plan to retire Pro. Flash has native vision; Pro is text-only. Peak pricing applies in two UTC windows on weekdays, excluding Chinese public holidays; other periods cost half. A price chart that ignores scheduling and caching can substantially misstate the bill.
AA’s Flash Max result is 39 at $0.27 per Index task, versus Pro Max at 36 and $0.67. The newer Flash is stronger in that composite despite its lower token price. The names “Flash” and “Pro” are product labels; measure the current versions.
DeepSeek’s model card also shows the effect of the agent: its DeepSWE result ranges from 65.5 with OpenCode to 74.2 with mini-SWE. Those are provider runs, and the difference comes with a harness change. Flash’s MIT weights are useful for deployment control, but its disclosed 552B backbone still has a substantial memory footprint.
Qwen3.8-Max-0902 and Kimi K3
Qwen’s hosted 0902 snapshot supports text, images and video, with a 1M context and separate output/reasoning limits. Implicit cache reads cost $0.25; explicit creation and reads cost $2.50 and $0.17. The floating Max alias moved to 0902, so use a pinned snapshot when comparing results across dates. The open 2.4T-A95B checkpoint is a separate target.
Kimi K3 is the broad current Kimi candidate. K2.7 Code remains a specialized 256K coding option rather than a synonym for K3. Current AA pages give Qwen Max 0902 an Index of 45 and K3 Max 44. Their measured costs, $5.41 and $2.00 per task, reinforce the need to inspect reasoning length rather than comparing the input rate alone.
GLM-5.3 and its cheaper variants
Z.ai lists GLM-5.3 at $1.40/$4.40, GLM-5.3-Flash at $0.15/$0.50 and FlashX at $0.37/$1.25. Context and output contracts should be checked for each variant rather than inherited from the flagship. Mistral’s hosted GLM card specifies 1M context and 128K output, with a different cache rate from Z.ai direct. AA’s GLM Max page reports Index 45 at $2.01 per task. It is a credible alternative to include in a coding/agent evaluation, not proof that every endpoint behaves identically.
Meta Muse Spark 1.3
Muse Spark 1.3 is Meta’s current hosted frontier model. Llama 4 remains an open deployment option, but “Llama 5” would be an invented replacement in this comparison. Spark has a broad input contract, although Meta warns that 1.3 audio understanding is not fully supported and quality may degrade. Its AA Max result is 48 at $1.60 per task, with about 184 output tokens per second in that snapshot.
Meta’s Contributor endpoint costs $0.10 input, $0.002 cached input and $0.20 output, much less than Standard. Prompts and completions may train future models, and Contributor lacks Max effort. That is a different data-use and capability bargain. A privacy-sensitive team should evaluate the appropriate contract, not rank the lowest price as interchangeable with Standard.
Mistral, MiniMax, Xiaomi and StepFun
Mistral Medium 3.5 combines instruct, reasoning and coding in a dense 128B checkpoint. Large 3 has 675B total and 41B active parameters, Apache 2.0 licensing and a lower hosted token rate. Total parameter count and release naming do not create a performance ranking; test both if their deployment terms fit your task.
MiniMax M3 is a low-cost multimodal candidate with a price doubling above 512K input. The model card describes sparse attention and roughly 428B total/23B active parameters. Its provider benchmark results span distinct tests such as SWE-bench Pro and Verified, which should retain their separate labels.
Xiaomi’s current family is MiMo V2.6 Pro and Flash, with text, image, video and audio input. AA’s Pro result is 46 at $0.13 per task. That combination warrants a place in a cost-sensitive evaluation. Its open MOPD checkpoint and hosted API are distinct configurations; preserve which one you tested.
Step 5 Preview supports tools, JSON Schema, multimodal input and a 64K maximum output. We use the provider’s limit because an intermediary listing can disagree. We have not confirmed downloadable Step 5 weights in the checked model contract. AA’s snapshot puts it at Index 44 and $0.72 per task.
The rest of the specifications
| Model / direct API ID | Context / output cap | Native input → output | Contract detail |
|---|---|---|---|
Grok 4.7 / grok-4.7 |
500K / not verified | Text, image → text | Low–xhigh; high default; no Batch |
Qwen3.8-Max-0902 / qwen3.8-max-0902 |
1M / 131K | Text, image, video → text | Separate 262K reasoning limit |
DeepSeek Flash / deepseek-flash |
1M / 384K | Text, image → text | V4.1-Flash current endpoint |
DeepSeek Pro / deepseek-v4-pro |
1M / 384K | Text → text | V4-Pro-0813; separate from Flash |
Kimi K3 / kimi-k3 |
1,048,576 / not verified | Text, image, video → text | Open checkpoint; provider-specific license |
GLM-5.3 / glm-5.3 |
1M / 128K on Mistral-hosted version | Text → text | Direct and hosted contracts differ |
Muse Spark 1.3 / muse-spark-1.3 |
1,048,576 / not verified | Text, image, video, PDF; limited audio → text | Hosted model; audio support warning |
Mistral Medium 3.5 / mistral-medium-3-5 |
256K / not verified | Text, image → text | Open weights; modified MIT terms |
Mistral Large 3 / mistral-large-2512 |
256K / not verified | Text, image → text | Open weights; Apache 2.0 |
MiniMax M3 / MiniMax-M3 |
1M / not verified | Text, image, video → text | Community license |
MiMo V2.6 Pro / mimo-v2.6-pro |
1M / 128K | Text, image, video, audio → text | Open checkpoints also released |
MiMo V2.6 Flash / mimo-v2.6-flash |
1M / 128K | Text, image, video, audio → text | Lower-cost member |
Step 5 Preview / step-5-preview |
1M / 64K | Text, image, video → text | Preview; weights not confirmed |
Specification sources for this table: Grok, Qwen, DeepSeek, Kimi, hosted GLM, Meta, Mistral Medium, Large, MiniMax, MiMo Pro, Flash and StepFun. “Not verified” avoids turning a missing output limit into an invented one.
Open weights, licenses and hardware cost
Open weights give you deployment control, but the licenses differ. Qwen3.8’s custom license and Kimi K3’s license contain separate-agreement and attribution conditions for specified business categories and scale. GLM-5.3 has a custom license too. Calling all three “MIT” would mislead a deployment decision.
Medium 3.5’s modified MIT license excludes companies or employers above $20M in consolidated monthly revenue under the stated test unless they obtain a separate license. MiniMax M3 requires commercial attribution and notice, with prior authorization for specified products/services above $20M annual revenue. MiMo Pro-MOPD is MIT, and Large 3 is Apache 2.0. Read the exact checkpoint license before deciding the deployment fits your organization.
Mixture-of-experts active parameters describe part of the per-token computation. Total weights still require storage and memory. Kimi K3’s disclosed 2.8T parameters would imply about 5.6TB for 16-bit weights or 1.4TB at four bits before overhead; that is our arithmetic, not a vendor minimum hardware requirement. KV cache, activations, networking and serving headroom add more. Self-hosting replaces the provider’s token meter with hardware, utilization and operations costs. Low traffic can make an API cheaper even when the weights are available.
What the benchmark labels can conceal
A terminal benchmark usually tests the complete agent: model, prompt, shell tools, environment, time limit and verifier. Terminal-Bench 4.0 changed resource calibration and task composition; major versions require rerunning. A 2.1 score is not an older measurement of the identical test. The benchmark’s run guide uses repeated trials, while other evaluators may choose different repetition counts and agents.
Repository tests can contain flawed acceptance checks. Epoch’s September DeepSWE v1.1 audit found confirmed false negatives in 23 of 113 tasks, including conflicts involving agent-created tests. That limits how much confidence to place in a small DeepSWE gap. Its SWE-bench Verified review also identifies contamination risk, verifier problems and harness dependence.
Epoch’s Terminal-Bench 4.0.0 audit identifies scoring defects affecting 30 of 66 tasks, including false positives, false negatives and leakage. The audit applies to that version, not every later revision. Its HLE review concerns original HLE, not HLE-Rolling or HLE-Verified. Benchmark criticism needs the same version precision as a benchmark score.
Even a familiar name can conceal a different subset. SWE-bench Verified has 500 instances, while Epoch’s evaluation uses 484 validated instances. “Same benchmark” needs the version and subset to mean the same thing. An Elo score, F1 score, all-pass percentage and partial-credit score are different measurements; averaging their raw values creates a meaningless result.
Our standard for reading a leaderboard is to keep the task, version, effort, harness, tools, timeout, fallback policy and evaluator beside the score. Look for uncertainty, repeated runs and error examples. A two-point lead can be weaker evidence than a large difference in completion cost or a failure that repeatedly affects your users.
Safety, reliability, privacy and deployment belong in the comparison
Safety evals and capability evals answer different questions. A cybersecurity score may involve restricted access or fewer production safeguards than the service you can buy. A production refusal can be appropriate policy enforcement, while an invented citation or unauthorized action is a different failure. Count those outcomes separately.
Google’s Argon access policy starts with defenders. Claude launch reports disclose fallbacks for selected sensitive tasks. OpenAI’s GPT-6 documentation describes monitoring that can pause or stop an agent. These controls affect task availability and latency, so evaluate the deployed configuration you will use. A model-card maximum is not a promise that an ordinary public endpoint will perform every action from the research test.
For reliability, test citation fidelity, abstention when evidence is missing, adherence to user constraints, recovery after tool failures and structured-output validity. A fluent answer can fail each one. Keep a set of adversarial documents and tool results that attempt to redirect an agent away from its user’s task. Give the agent only the permissions its workflow needs, and examine its behaviour when a tool reports a failure or denies an action.
For privacy, distinguish consumer apps, paid APIs, free tiers, contribution programs and negotiated enterprise terms. Google’s pricing documentation distinguishes free-tier data use from paid-tier treatment. Cloud marketplaces and regional endpoints can have different terms and prices. “OpenAI-compatible API” says something about the request format; it says nothing by itself about retention, training use or geographic processing.
For deployment, check quotas and rate limits with the account that will serve real traffic. A model can have a generous advertised context window yet insufficient tokens-per-minute capacity for your concurrent workload. Measure queueing, retries, cancellation, timeout recovery and median plus tail latency. Include tool and sandbox costs, and preserve exact model IDs where the provider supports them so an alias update does not silently change your evaluation target.
How to evaluate these models for your own work
A useful comparison begins with representative private tasks. Start with a manageable set of repository fixes, researched reports, spreadsheet reconciliations, long-document questions, visual interpretation and tool workflows drawn from actual work. Include easy cases, difficult cases and tasks where the correct result is to ask for missing evidence. Set success criteria before seeing the model outputs.
- Record the configuration. Save exact model IDs, date, effort, temperature where supported, output/turn/time budgets, tools, retrieval corpus, cache policy and fallback rules. Test a common agent harness first, then test complete products separately. A raw-model comparison and an app comparison answer different questions.
- Compare budgets as well as effort labels. Use low-cost, standard and high-compute bands. Record actual tokens, dollars and elapsed time. Equal labels across providers do not guarantee equal resources. Include cache misses and the original write in cost calculations.
- Grade the deliverable. Use executable tests, reconciled totals, verified citations and explicit rubrics. Blind human reviewers to model names and rotate output order. Use a second review for consequential mistakes and disagreement. Avoid a single model judge that rewards its own preferred style.
- Repeat uncertain tasks. Report sample size, pass rates and uncertainty. Save per-task outcomes and useful failure examples, including plausible answers with wrong details. A small sample should produce a provisional shortlist, not a universal winner.
- Separate failure types. Track task completion, factual errors, unsupported claims, valid formatting, appropriate refusals, unnecessary refusals, tool recovery and instruction violations. Mark whether a fallback handled the task. Keep the original model’s result distinct from the complete system’s result.
- Price accepted work. Report cost per attempt, cost per success, reviewer time, median latency and tail latency. Establish a quality floor before optimizing cost. Then test routing: a cheaper first attempt, a clear acceptance check and escalation for unresolved work.
- Keep the comparison current. Rerun affected cases after a model, prompt, retrieval or agent update. Keep anonymized traces and usage receipts so you can explain why a result changed. Avoid silently changing the model behind a fixed benchmark label.
This plan is our proposed evaluation method. It has not been run for this article. The tables above remain provider and independent-evaluator results under their respective configurations.
A practical buying shortlist
For frequent coding and document work, start with Sol 6.1 and Sonnet 5.5 at sensible effort, and test Opus or Astra on the failures. For difficult scientific tasks, compare Astra and Opus against the same verifiable deliverables. For visual and long-context work, add Argon when access permits, plus available Gemini alternatives. For high-volume extraction, classification and transformations, test the lower-cost options with the same quality floor.
For local or sovereign deployment, begin with license and hardware constraints, then evaluate suitable open weights. For an enterprise API, begin with data handling, regional availability, quotas and tool support. These constraints can eliminate a model before a benchmark score becomes relevant.
The next useful purchase decision is a small workload-specific evaluation with recorded costs and reviewed outputs. Choose the model that clears your acceptance criteria at the best total cost, and retain a stronger option for the tasks that still need it.
Editorial method. We checked official model documentation, launch materials, pricing pages and original evaluation sites on September 30, 2026. Sources are linked beside the relevant claims. Tables preserve versions, effort settings and missing results. Our cost examples are arithmetic scenarios; recommendations are our interpretation of the evidence. No paid model tests were commissioned and no hands-on performance is claimed.
