AI News

Mistral Large 4 (Le Chonk): Specs, Benchmarks, Pricing and the Open-Weight Promise

Mistral Large 4 is a serious new contender for legal agents, finance workflows and European deployment. Its first independent results do not establish it as the strongest general-purpose model. The useful story is more specific: a trillion-parameter mixture-of-experts model, unusually promising specialist results, aggressive preview pricing, and weights promised for later this month.

Mistral announced the model, nicknamed Le Chonk, on October 6, 2026. It is available through the API in public preview. The announced open-weight release is scheduled for the end of October; a downloadable release and final license were not available when we checked. That distinction matters for anyone budgeting a private deployment today.

Living article: Evidence updated October 6, 2026, at 09:08 PDT / 16:08 UTC. We are checking for material developments every two hours through October 9. New evaluations, pricing changes and release details appear in the dated update log below. Kingy local tests: 0. This is an analysis of published evidence, not a hands-on model review.

The headline numbers are 1.05 trillion total parameters, 49 billion active parameters, a 1.6 billion-parameter vision encoder and a manufacturer-listed one-million-token context window. The headline caveats are equally consequential: independent evaluators currently list roughly half that context, benchmark configurations differ, and the price shown in Mistral’s current documentation is half the rate used in the first independent evaluations.

The verdict at launch

Large 4 earns a place on an enterprise evaluation shortlist. It has a credible case for document-heavy legal work, financial research and organizations that value Mistral’s European infrastructure. The evidence is weaker for choosing it as the default model for every coding or knowledge task.

On Vals AI’s broad industry index, it scores 48.05%, behind GPT-6 Astra, GPT-6.1 Sol, Claude Sonnet 5.5, Claude Opus 5.5 and several other competitors. On the much narrower Harvey Legal Agent Benchmark, it scores 15.83%, ahead of those four models in that particular test. Finance Agent v2 also gives it a small advantage over Astra, although other models score higher. A specialist win and a broad-index loss can both be true. Vals model profile

The price makes these results more interesting. Mistral currently displays $0.68 per million input tokens, $0.07 per million cached-input tokens and $2.09 per million output tokens. Its card also displays crossed-out rates of $1.36, $0.14 and $4.18. Those higher input/output rates are what Vals and Artificial Analysis list for their initial results. We separate the current tariff from historical evaluator cost throughout this article. Mistral model card, Mistral pricing

The most defensible buying decision is therefore workload-specific: test the legal or finance workflow you actually operate, against the same tools and acceptance criteria, while keeping a stronger generalist available for tasks where Large 4 falls short. The broad benchmark evidence gives little reason to replace an entire model stack on announcement day.

Mistral Large 4 specifications

Item What is currently established Evidence and limits
Release October 6, 2026; public API preview Training and post-training are still developing
Architecture Granular mixture of experts Full architectural details have not yet been published
Total parameters 1.05 trillion Mistral’s launch post rounds this to 1T
Active parameters 49 billion Active compute is not the total memory needed to serve the model
Vision encoder 1.6 billion parameters Native multimodal model; supported inputs include text and images
Context window 1 million tokens on Mistral’s card Vals lists 512K; Artificial Analysis and Vercel’s live provider table list 524K; discrepancy remains unresolved
Maximum output Vals lists 256K; Vercel’s provider table lists 262K Evaluator/provider metadata, not a Kingy-verified API ceiling
API identifier mistral-large-4 Record the actual version used for reproducible comparisons
API features Structured outputs, function calling, document Q&A, prefix completion, batch processing, agents and conversations Feature availability does not establish benchmark quality
Weights Announced for the end of October Download and license sections currently say “Coming soon”
Training infrastructure 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centers No verified total training bill published
Language coverage Mistral says training data covers 160+ languages, including all official EU languages Coverage is not a per-language accuracy guarantee

Sources: official model card, launch announcement, Vals profile, Artificial Analysis profile, Vercel AI Gateway catalog.

Why 49B active parameters does not make this a 49B deployment

A mixture-of-experts model routes computation through selected experts rather than activating all its weights for every token. Dividing the advertised 49B active parameters by 1.05T total gives approximately 4.7%. That is a useful description of the architecture, not a prediction that serving it costs 4.7% as much as a dense model of the same total size.

The complete model still contains the other experts. A deployment must store them, distribute or load them, move activations and manage communication between devices. Context length, batch size, precision, routing implementation and parallelism all affect throughput and cost. Active parameter count alone cannot answer how many GPUs you need.

Simple weight arithmetic shows the scale. At two bytes per parameter, 1.05T parameters imply roughly 2.10 TB of raw weight data. At one byte, the figure is roughly 1.05 TB. Idealized four-bit packing gives roughly 525 GB. These are decimal payload estimates calculated from the advertised parameter count. They exclude runtime overhead, quantization metadata, KV cache, activations, redundancy and any uncertainty about which components the total includes.

Those estimates are not validated hardware requirements. A quantized checkpoint, a supported runtime and a working distributed configuration must exist before a credible deployment recommendation can be made. There is no basis here for promising a convenient consumer-GPU or laptop setup. Our local AI hardware guide explains the distinction between fitting weights and serving a useful workload.

The one-million-token claim needs an endpoint check

Mistral’s own model card is the primary source for the advertised 1M context. Vals currently lists 512K and Artificial Analysis 524K. We have not established whether those differences reflect an earlier preview configuration, evaluator settings, different endpoints or metadata that needs updating.

Provider update, 09:08 PDT: Vercel’s live AI Gateway provider table lists the Mistral route at 524K context and 262K maximum output, under gateway identifier mistral/mistral-large-4. It displays the same $0.68 input, $0.07 cache-read and $2.09 output rates per million tokens. A search-indexed version of the listing showed 1M context; we use the directly fetched provider table for this snapshot. These are published route specifications, not a Kingy API acceptance test. They strengthen the case for checking the chosen route rather than treating the model card’s 1M as a universal limit. Vercel’s Large 4 listing

For an implementation, verify the limit on the exact endpoint and model version you intend to use. Reserve room for the response and any tool messages. A context window is a shared budget; one million input tokens plus a substantial output may exceed a one-million-token total context.

Capacity also differs from useful recall. A model accepting a large document does not establish that it can reliably find every relevant clause, reconcile distant evidence or avoid inventing citations. Long-context retrieval tests and your own document acceptance tests remain necessary. We will update this section when the configuration discrepancy is resolved.

What the independent benchmarks actually show

Benchmarks measure tasks under a particular implementation. Agent scaffolding, tool permissions, reasoning effort, sampling, output limits and retry policies can change the result. We use the same evaluator’s tables for comparisons wherever possible and keep differently configured tests separate.

Broad capability: competitive, but behind the frontier leaders

Model Vals Index v2.1 Vals reported cost per test
Claude Opus 5.5 66.97% $32.14
Claude Sonnet 5.5 67.04% $21.34
Gemini 4 Argon 68.90% $15.68
GPT-6 Astra 63.13% $18.46
GPT-6.1 Sol 61.15% $3.24
GLM 5.3 53.51% $7.25
Kimi K3 50.30% $6.38
Qwen 3.8 Max 48.27% $4.96
DeepSeek V4.1 Flash 51.32% $0.33
Mistral Large 4 Preview 48.05% $13.78
DeepSeek V4 Pro 0813 47.63% $3.35

Snapshot: October 6. Vals uses its own evaluation configurations. The Large 4 profile specifies high reasoning effort, temperature 1 and top-p 0.95, while warning that individual benchmarks may use different settings. Vals Index

Separately, Artificial Analysis reports an Intelligence Index score of 38 for Large 4 Preview. Its v4.3.2 index incorporates ten evaluations, spanning agentic work, coding, knowledge, reasoning and long-context tasks. That score is on a different scale from Vals; the two must not be averaged or compared as percentages of the same test. Artificial Analysis profile

Vals reports Large 4 at 48.05% ±1.11, ranking 32nd of 44 models in the displayed index. The index combines industry-oriented evaluations across finance, coding, legal and tax work with economic weighting. It is not a universal probability that the model will complete your task.

The small numerical distance between Large 4 and Qwen 3.8 Max, or between Large 4 and DeepSeek V4 Pro, should not support a sweeping rank claim. The displayed uncertainty and different workload results matter more than a few tenths of a point. The larger gaps to the leading frontier models are harder to dismiss, although your specific application may favor Large 4.

Mistral’s announcement describes it as the best open-weight model from the US or Europe on aggregated benchmarks. That is a geographically bounded vendor claim, with an announced future weight release. It does not mean best among all models, best among all globally developed open models or downloadable today. The broader independent tables above are the more useful starting point for procurement.

Finance: a credible result with a qualified Astra comparison

Model Finance Agent v2 accuracy Reported cost per test Reported latency
Gemini 4 Argon 65.40% $4.38 16m 21s
Claude Opus 5.5 58.59% $9.22 33m 08s
Claude Sonnet 5.5 58.10% $6.55 31m 43s
GLM 5.3 55.84% $1.07 15m 50s
Mistral Large 4 54.68% $1.20 26m 02s
GPT-6 Astra 53.54% $6.82 11m 33s
Kimi K3 53.11% $1.91 4m 46s
GPT-6.1 Sol 52.03% $1.62 15m 25s
DeepSeek V4 Pro 0813 50.39% $0.88 16m 15s

Source: Vals Finance Agent v2. Latency describes this evaluator’s agent workloads, not chat response time or a service-level guarantee.

Large 4 exceeds Astra by 1.14 percentage points in this snapshot, with a substantially lower reported test cost. It also takes more than twice as long in the displayed latency measure. That tradeoff matters: a research job that runs overnight and an analyst waiting interactively have different requirements.

This result supports evaluating Large 4 for financial workflows; it does not establish a general finance crown. Gemini, both Claude models and GLM 5.3 score higher in the same table. Financial reasoning accuracy also says nothing by itself about permission controls, auditability, data provenance or the suitability of generated advice.

Model Harvey Legal Agent Benchmark task pass rate Reported cost per test
Gemini 4 Argon 19.58% $10.55
Mistral Large 4 15.83% $4.18
Kimi K3 12.92% $3.83
Qwen 3.8 Max 10.42% $2.37
GLM 5.3 8.33% $4.30
DeepSeek V4 Pro 0813 7.50% $0.17
GPT-6 Astra 5.42% $26.16
GPT-6.1 Sol 5.42% $3.93
Claude Opus 5.5 3.75% $21.38
Claude Sonnet 5.5 2.92% $16.61

Source: Vals Harvey Legal Agent Benchmark.

Large 4 ranks sixth of 75 models in the displayed results. That is a meaningful specialist signal and substantially better than its broad-index rank. The low absolute pass rates also show how demanding the task-level test is. A 15.83% result is not a claim that legal professionals can delegate the other 84.17% safely, nor a measurement of lawyer replacement.

Different legal scores can refer to different units: passing an entire task, satisfying an individual rubric criterion or answering a research question. Do not compare those percentages as though they measure the same outcome. Large 4’s separate Legal Research Bench result is 31.73%, ranking 38th of 74, which further cautions against converting one strong agent benchmark into a universal legal superiority claim. Vals model results

For buyers, the next step is a controlled evaluation of actual contracts, research questions and citation requirements, with qualified review. The benchmark makes that experiment worthwhile. It does not remove the review requirement.

Coding agents: the harness changes the conclusion

Mistral’s launch cites Artificial Analysis results of 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Their arithmetic mean is 49.8%, the cited Coding Agent Index. These are launch-reported AA results; we have not independently reproduced the runs. Mistral announcement

AA’s coding index methodology combines three benchmarks with equal benchmark weighting. Its current evaluation describes 113 DeepSWE tasks, 66 Terminal-Bench 4 tasks and 124 SWE-Atlas tasks, with results averaged across three attempts per task. A composite helps summarize that test set, but it can conceal a model’s much weaker terminal performance. Artificial Analysis coding agents

Vals’ separate Terminal-Bench 4 evaluation uses Mini-SWE-agent and reports a different Large 4 result:

Model Vals Terminal-Bench 4 accuracy Reported cost per test
Claude Opus 5.5 65.15% $13.20
Claude Sonnet 5.5 64.14% $16.51
GPT-6 Astra 59.60% $9.58
GPT-6.1 Sol 55.05% $1.72
GLM 5.3 38.89% $9.37
Qwen 3.8 Max 34.34% $10.63
Mistral Large 4 22.73% $7.90
Kimi K3 17.17% $12.02
DeepSeek V4 Pro 0813 14.14% $3.31

Source: Vals Terminal-Bench 4. This is a separate evaluation from the AA result cited by Mistral.

The 22.73% and 28.3% Large 4 numbers should remain separate. Averaging them would invent a new benchmark. Differences in task selection, scaffolding and execution budgets can explain differences between evaluators, but we have not isolated the cause here.

Vals also reports 78.40% on Vibe Code Bench v1.1, 30.56% on Code Migration and 14.26% on Vibe Code Bench 1–100. Those are distinct tasks and versions. The harder 1–100 score cannot be substituted for v1.1 or compared against a competitor’s v1.1 result without labeling the difference. Vals profile, Vibe Code benchmark

The practical conclusion is restrained. Large 4 can be a credible coding candidate, particularly where deployment constraints matter. The independent terminal table gives a strong reason to retain a frontier coding agent for difficult autonomous work and compare cost per accepted change, including repair attempts and reviewer time.

Cybersecurity, vision and safety: promising claims that need careful labels

Mistral reports 82% on CyberGym-E2E, 93% on Cybench and a top-five position in the Artificial Analysis Cyber Index. These are the company’s launch claims. They should not be presented as Kingy-run tests, and the 40-task Cybench result should not be treated as a comprehensive security audit. Mistral launch

AA’s Cyber Index focuses on defensive work using CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA: reproducing and patching real vulnerabilities. Its methodology also explains that safety refusals can receive zero scores. A lower score can therefore reflect a refusal policy as well as an inability to solve a problem. The benchmark does not collapse those two causes into an explanation of commercial suitability. AA Cyber Index, methodology

Mistral describes additional access for selected cybersecurity partners, with reduced moderation and expanded cyber capabilities. Results from such access must be distinguished from the public API preview. Model identity alone is insufficient if policy configuration or tool access differs. We are watching for independently published configurations and results that clarify this point.

For vision, Mistral reports 42% on Dense200, compared with 41% for GPT-6 Astra in its comparison. That is one percentage point on one test. Without uncertainty estimates and broader replicated evaluations, it supports a promising visual-grounding result, not a general claim that Large 4 is the best vision model.

The supported product story is text-and-image input with text output, including document Q&A and structured responses. Do not infer native audio, video generation or image generation from the word “multimodal.” A useful business evaluation would measure document extraction, bounding or grounding accuracy, small text, charts and citation consistency separately.

Mistral also reports 93.3% attack resistance on B3 and 1.691 out of 2 on KORABench. Those are named evaluations under reported conditions, not a promise that the model resists every prompt injection. Production agents still need scoped permissions, output validation and human approval where the action warrants it. A better model score does not make a tool connection safe by itself.

API pricing: the launch discount changes the comparison

All figures below are US dollars per million tokens, checked October 6. Prices exclude taxes and additional tools or infrastructure.

Mistral Large 4 token type Current displayed rate Crossed-out rate on the model card
Input $0.68 $1.36
Cached input $0.07 $0.14
Output $2.09 $4.18

The current values also appear in the Standard pricing table, so this is not simply a comparison between Standard and Batch modes. The displayed reduction is 50%. We have not verified an expiration date or a contractual guarantee that these preview rates will persist. Budget-sensitive deployments should retain both scenarios. Mistral pricing, model card

What a request costs

The uncached token formula is straightforward:

cost = input_tokens × input_rate / 1,000,000 + output_tokens × output_rate / 1,000,000

Illustrative workload Current displayed rates Crossed-out rates
10,000 input + 2,000 output tokens $0.01098 $0.02196
100,000 input + 10,000 output tokens $0.08890 $0.17780
900,000 input + 50,000 output tokens $0.71650 $1.43300
Monthly: 10M input + 2M output $10.98 $21.96
Monthly: 100M input + 20M output $109.80 $219.60
Monthly: 1B input + 200M output $1,098.00 $2,196.00

These are calculated token budgets, not measured invoices. The long request assumes an endpoint that actually supports the advertised 1M context and enough room for all messages and output. Agent loops can process the same growing conversation repeatedly, so a 100K document does not imply only 100K billed input tokens over an entire job.

Cached input can reduce a repeated-prefix workload substantially. Suppose a request has 100K input tokens, 90K qualify as cache hits and it produces 2K output tokens. At the displayed rates, 10K uncached input costs $0.00680, 90K cached input costs $0.00630 and output costs $0.00418: $0.01728 total, compared with $0.07218 if all the input were uncached.

That example assumes the cache hits actually qualify. It does not model cache writes, expiry, changed prefixes, tool charges or a guaranteed hit rate. Check the applicable caching rules before treating the saving as a forecast. Likewise, we do not stack a separate batch reduction on top of the preview discount without verified terms for the exact route.

Current token prices against frontier and lower-cost alternatives

Model/provider Uncached input /M Output /M Cost of 10K input + 2K output
Mistral Large 4, current displayed $0.68 $2.09 $0.01098
Mistral Large 4, crossed-out rates $1.36 $4.18 $0.02196
GPT-6 Astra, Standard short-input tier $10.00 $50.00 $0.20000
GPT-6.1 Sol, Standard short-input tier $2.00 $10.00 $0.04000
Claude Sonnet 5.5, Standard $2.00 $10.00 $0.04000
Claude Opus 5.5, Standard $4.00 $20.00 $0.08000
GLM 5.3, Z.ai $1.40 $4.40 $0.02280
DeepSeek V4.1 Flash, peak $0.30 $1.20 $0.00540
DeepSeek V4.1 Flash, off-peak $0.15 $0.60 $0.00270
DeepSeek V4 Pro 0813, peak $1.32 $3.96 $0.02112
DeepSeek V4 Pro 0813, off-peak $0.66 $1.98 $0.01056
Mistral Large 3 $0.50 $1.50 $0.00800

First-party sources: OpenAI Astra, OpenAI Sol, Anthropic pricing, Z.ai pricing, DeepSeek pricing, Mistral pricing.

The OpenAI examples use requests within the standard short-input tier; requests above 272K input tokens have different pricing multipliers. DeepSeek’s peak and off-peak schedules are separate tariffs, not prices you can select independently of when a request runs. Provider routes, caching, batch modes and priority service also matter. This table deliberately compares an uncached, short example rather than pretending every provider’s billing system is identical.

At current displayed rates, Large 4’s illustrative request costs about 3.64 times less than Sol or Sonnet and 18.21 times less than Astra. These ratios describe token prices for this input/output mix, not quality-adjusted savings. A harder request can change the ratio through reasoning output, retries, cache behavior and task completion.

For a deeper comparison of the expensive generalist alternative, see our GPT-6 Astra versus Claude Sonnet 5.5 analysis.

Why evaluator task cost is more revealing than the tariff

Vals reports $13.78 per Vals Index test for Large 4, compared with $3.24 for Sol, despite Sol’s higher current headline token rates. Large 4’s reported index latency is 105 minutes 50 seconds. The evaluator’s workload includes agent behavior and its own configuration; it is not a single small chat request.

Those results use the $1.36/$4.18 Large 4 rates listed in the evaluator’s snapshot. Applying a 50% tariff reduction to an identical token trace would mathematically halve its token charge. That would be a hypothetical recalculation, not a new observed Vals result. We therefore retain the published $13.78 and do not relabel it as a live invoice estimate.

For a business, the useful denominator is accepted work: total model, tool, infrastructure and review cost divided by successful, usable outcomes. Track failed attempts too. A lower-priced model can be expensive if it makes longer calls or requires more repair. A more expensive model can be economical when it finishes cleanly. The finance and legal tables give Large 4 stronger cost-quality arguments than the broad index does.

Artificial Analysis reports roughly 116.1 output tokens per second in its profile, alongside a reported $1.13 Intelligence Index task cost at the higher launch rates. Throughput is a useful serving signal but does not include all the time spent waiting, reasoning and calling tools in a business agent. It cannot be compared directly with Vals’ end-to-end job latency. AA profile

Open weights, licensing and the self-hosting cost question

The word “open” needs a date and an artifact. On October 6, the API is public, while the weights are promised for the end of the month. Mistral’s model documentation lists the download and license as coming soon. We cannot verify commercial redistribution rights, modification terms or supported deployment formats from an unreleased license.

Mistral Large 3 is an instructive predecessor: its documentation specifies 675B total parameters, 41B active parameters, a 256K context window and Apache 2.0 licensing. Large 4 increases the advertised total count by about 56%, the active count by about 20%, and the card’s context window by roughly four times. None of those changes guarantees a proportional quality increase, and Large 3’s license does not automatically become Large 4’s license. Large 3 model documentation

Comparisons with “open source” also need care. An API model associated with an open-model ecosystem is not necessarily a downloadable checkpoint. In particular, the Qwen 3.8 Max row above names an evaluated API model; it should not be silently replaced by a smaller open Qwen checkpoint. Model size, release version, license and inference provider belong in the comparison.

For a future self-hosted deployment, compare the entire operating bill: accelerator rental or depreciation, CPU and RAM, networking, storage, monitoring, redundancy, engineer time and utilization. Running a large reserved cluster for sporadic requests can be much more expensive than buying tokens. Sustained high utilization, private-data constraints or specialized customization can change that calculation.

We have no verified Large 4 self-hosting throughput, quantized quality measurements or production-serving recipe. Until the release arrives, there is no evidence-backed break-even price to publish. We also cannot estimate training cost from 3,800 GPUs alone: duration, utilization, ownership, power, networking and accounting treatment are missing.

European deployment and data residency

European training infrastructure is relevant to procurement, but it does not establish that every API request, log and account record remains in a chosen jurisdiction. Mistral offers regional inference routes for the EU and US and documents a 10% regional pricing uplift across applicable token charges. Its global endpoint is distinct from a region-specific route. Mistral regional inference documentation

The regional documentation specifies 1.1 times standard list pricing. Applied to the higher $0.02196 short-request example, that is $0.024156. We have not verified whether the preview discount also applies to the selected regional route. Confirm model availability and the applicable tariff before making a regional deployment budget.

The residency documentation also distinguishes inference processing from control-plane information such as accounts, billing and analytics. Buyers should check those boundaries against their actual requirements. European model development, regional inference and a fully controlled private deployment are related purchasing considerations with different operational consequences.

The potential attraction of Large 4 is having a European provider now and a possible private-deployment option later. Whether that option fits a particular organization depends on the released license, serving requirements and contractual controls.

How to evaluate it without wasting the preview window

Start with a narrow set of tasks drawn from your own workload. For legal work, require traceable citations and correct treatment of the relevant clauses. For finance, require sourced calculations and consistent units. For coding, require passing the project’s tests and a reviewable change. Define success before looking at the outputs.

Use the same evidence, tools, execution limits and acceptance criteria for each model. Record reasoning effort, model version, date, token usage, retries and wall-clock time. A powerful model with a weak agent wrapper and a weaker model with a well-tuned wrapper are not a clean model comparison.

The official usage example supports selecting high reasoning effort. A minimal illustrative API call is:

import os
from mistralai.client import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
    model="mistral-large-4",
    messages=[{
        "role": "user",
        "content": "Summarize the supplied evidence and cite each conclusion."
    }],
    reasoning_effort="high",
)
print(response.choices[0].message.content)

This follows the official model usage documentation. It is an illustrative starting point, not a tested integration or a completed benchmark. Set suitable budgets and output limits for your application, and avoid sending restricted information through an unapproved route.

For agents, inspect the failures rather than just the aggregate score. Was the answer wrong, a tool call malformed, a source unavailable, the budget exhausted or the action refused? Those failures require different remedies. Keep a frontier-model comparison and a cheaper-model comparison so the decision includes both quality and cost.

What we are watching over the next 72 hours

The initial publication leaves several questions open:

  1. Independent replication: additional evaluator results, revised tables and disclosed reasoning or agent settings, especially in legal, finance and cybersecurity.
  2. Context configuration: why Mistral advertises 1M while evaluators list roughly 512K, and what the public endpoint actually permits.
  3. Price durability: whether the current 50% reduction has an expiration date, how batch pricing applies and whether evaluator cost tables adopt the lower tariff.
  4. Weights and license: a concrete release date, downloadable artifacts, the final license and supported formats. The announced release is later than this initial 72-hour watch.
  5. Operational evidence: time to first token, throughput under realistic concurrency, tool-call reliability, long-document accuracy and total task costs.
  6. Preview changes: whether model updates materially change the scores or require identifying a new version rather than overwriting old results.

The article will preserve a dated record of material changes. A source check that finds nothing new will not create a cosmetic “updated” date. An unavailable source will not be treated as evidence that the underlying result is unchanged.

Dated update log

  • October 6, 2026, 09:08 PDT / 16:08 UTC — Added provider-specific limits and access details. Vercel’s directly fetched AI Gateway table lists 524K context, 262K maximum output and the discounted token rates for mistral/mistral-large-4. This differs from its search-indexed 1M label and does not resolve Mistral’s official 1M versus evaluator limits. Rechecked official prices and published evaluations; the cited figures remain unchanged. Mistral’s weights and license tabs still say “Coming soon.” Provider listing, official model card.
  • October 6, 2026, 06:50 PDT / 13:50 UTC — Initial evidence snapshot. Added the official architecture and preview status, current discounted rates and higher evaluator price basis, independent broad/finance/legal/coding comparisons, and the unresolved 1M versus roughly 512K context discrepancy. No Kingy model inference tests have been run.

Sources and reporting method

We inspected Mistral’s announcement, model documentation, pricing and regional inference terms; the independent Artificial Analysis profile and Vals results; and the first-party competitor pricing linked above. The original Mistral announcement on X is the starting point for this coverage.

Benchmark figures are published results attributed to their operator or, where explicitly labeled, Mistral’s launch claims. Dollar examples are our arithmetic using the cited rates. We have not run paid model calls, matched local benchmarks or self-hosting trials. Kingy local tests: 0. The launch assessment can change as the preview evolves and independent evidence expands.