Mistral Large 4 is a serious new contender for legal agents, finance workflows and European deployment. Its first independent results do not establish it as the strongest general-purpose model. The useful story is more specific: a trillion-parameter mixture-of-experts model, unusually promising specialist results, aggressive preview pricing, and weights promised for later this month.
Mistral announced the model, nicknamed Le Chonk, on October 6, 2026. It is available through the API in public preview. The announced open-weight release is scheduled for the end of October; a downloadable release and final license were not available when we checked. That distinction matters for anyone budgeting a private deployment today.
The headline numbers are 1.05 trillion total parameters, 49 billion active parameters, a 1.6 billion-parameter vision encoder and a manufacturer-listed one-million-token context window. The headline caveats are equally consequential: independent evaluators currently list roughly half that context, benchmark configurations differ, and the price shown in Mistral’s current documentation is half the rate used in the first independent evaluations.
The verdict at launch
Large 4 earns a place on an enterprise evaluation shortlist. It has a credible case for document-heavy legal work, financial research and organizations that value Mistral’s European infrastructure. The evidence is weaker for choosing it as the default model for every coding or knowledge task.
On Vals AI’s broad industry index, it scores 48.05%, behind GPT-6 Astra, GPT-6.1 Sol, Claude Sonnet 5.5, Claude Opus 5.5 and several other competitors. On the much narrower Harvey Legal Agent Benchmark, it scores 15.83%, ahead of those four models in that particular test. Finance Agent v2 also gives it a small advantage over Astra, although other models score higher. A specialist win and a broad-index loss can both be true. Vals model profile
The price makes these results more interesting. Mistral currently displays $0.68 per million input tokens, $0.07 per million cached-input tokens and $2.09 per million output tokens. Its card also displays crossed-out rates of $1.36, $0.14 and $4.18. Those higher input/output rates are what Vals and Artificial Analysis list for their initial results. We separate the current tariff from historical evaluator cost throughout this article. Mistral model card, Mistral pricing
The most defensible buying decision is therefore workload-specific: test the legal or finance workflow you actually operate, against the same tools and acceptance criteria, while keeping a stronger generalist available for tasks where Large 4 falls short. The broad benchmark evidence gives little reason to replace an entire model stack on announcement day.
Mistral Large 4 specifications
| Item | What is currently established | Evidence and limits |
|---|---|---|
| Release | October 6, 2026; public API preview | Training and post-training are still developing |
| Architecture | Granular mixture of experts | Full architectural details have not yet been published |
| Total parameters | 1.05 trillion | Mistral’s launch post rounds this to 1T |
| Active parameters | 49 billion | Active compute is not the total memory needed to serve the model |
| Vision encoder | 1.6 billion parameters | Native multimodal model; supported inputs include text and images |
| Context window | 1 million tokens on Mistral’s card | Vals lists 512K; Artificial Analysis and Vercel’s live provider table list 524K; discrepancy remains unresolved |
| Maximum output | Vals lists 256K; Vercel’s provider table lists 262K | Evaluator/provider metadata, not a Kingy-verified API ceiling |
| API identifier | mistral-large-4 |
Record the actual version used for reproducible comparisons |
| API features | Structured outputs, function calling, document Q&A, prefix completion, batch processing, agents and conversations | Feature availability does not establish benchmark quality |
| Weights | Announced for the end of October | Download and license sections currently say “Coming soon” |
| Training infrastructure | 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centers | No verified total training bill published |
| Language coverage | Mistral says training data covers 160+ languages, including all official EU languages | Coverage is not a per-language accuracy guarantee |
Sources: official model card, launch announcement, Vals profile, Artificial Analysis profile, Vercel AI Gateway catalog.
Why 49B active parameters does not make this a 49B deployment
A mixture-of-experts model routes computation through selected experts rather than activating all its weights for every token. Dividing the advertised 49B active parameters by 1.05T total gives approximately 4.7%. That is a useful description of the architecture, not a prediction that serving it costs 4.7% as much as a dense model of the same total size.
The complete model still contains the other experts. A deployment must store them, distribute or load them, move activations and manage communication between devices. Context length, batch size, precision, routing implementation and parallelism all affect throughput and cost. Active parameter count alone cannot answer how many GPUs you need.
Simple weight arithmetic shows the scale. At two bytes per parameter, 1.05T parameters imply roughly 2.10 TB of raw weight data. At one byte, the figure is roughly 1.05 TB. Idealized four-bit packing gives roughly 525 GB. These are decimal payload estimates calculated from the advertised parameter count. They exclude runtime overhead, quantization metadata, KV cache, activations, redundancy and any uncertainty about which components the total includes.
Those estimates are not validated hardware requirements. A quantized checkpoint, a supported runtime and a working distributed configuration must exist before a credible deployment recommendation can be made. There is no basis here for promising a convenient consumer-GPU or laptop setup. Our local AI hardware guide explains the distinction between fitting weights and serving a useful workload.
The one-million-token claim needs an endpoint check
Mistral’s own model card is the primary source for the advertised 1M context. Vals currently lists 512K and Artificial Analysis 524K. We have not established whether those differences reflect an earlier preview configuration, evaluator settings, different endpoints or metadata that needs updating.
Provider update, 09:08 PDT: Vercel’s live AI Gateway provider table lists the Mistral route at 524K context and 262K maximum output, under gateway identifier mistral/mistral-large-4. It displays the same $0.68 input, $0.07 cache-read and $2.09 output rates per million tokens. A search-indexed version of the listing showed 1M context; we use the directly fetched provider table for this snapshot. These are published route specifications, not a Kingy API acceptance test. They strengthen the case for checking the chosen route rather than treating the model card’s 1M as a universal limit. Vercel’s Large 4 listing
For an implementation, verify the limit on the exact endpoint and model version you intend to use. Reserve room for the response and any tool messages. A context window is a shared budget; one million input tokens plus a substantial output may exceed a one-million-token total context.
Capacity also differs from useful recall. A model accepting a large document does not establish that it can reliably find every relevant clause, reconcile distant evidence or avoid inventing citations. Long-context retrieval tests and your own document acceptance tests remain necessary. We will update this section when the configuration discrepancy is resolved.
What the independent benchmarks actually show
Benchmarks measure tasks under a particular implementation. Agent scaffolding, tool permissions, reasoning effort, sampling, output limits and retry policies can change the result. We use the same evaluator’s tables for comparisons wherever possible and keep differently configured tests separate.
Broad capability: competitive, but behind the frontier leaders
| Model | Vals Index v2.1 | Vals reported cost per test |
|---|---|---|
| Claude Opus 5.5 | 66.97% | $32.14 |
| Claude Sonnet 5.5 | 67.04% | $21.34 |
| Gemini 4 Argon | 68.90% | $15.68 |
| GPT-6 Astra | 63.13% | $18.46 |
| GPT-6.1 Sol | 61.15% | $3.24 |
| GLM 5.3 | 53.51% | $7.25 |
| Kimi K3 | 50.30% | $6.38 |
| Qwen 3.8 Max | 48.27% | $4.96 |
| DeepSeek V4.1 Flash | 51.32% | $0.33 |
| Mistral Large 4 Preview | 48.05% | $13.78 |
| DeepSeek V4 Pro 0813 | 47.63% | $3.35 |
Snapshot: October 6. Vals uses its own evaluation configurations. The Large 4 profile specifies high reasoning effort, temperature 1 and top-p 0.95, while warning that individual benchmarks may use different settings. Vals Index
Separately, Artificial Analysis reports an Intelligence Index score of 38 for Large 4 Preview. Its v4.3.2 index incorporates ten evaluations, spanning agentic work, coding, knowledge, reasoning and long-context tasks. That score is on a different scale from Vals; the two must not be averaged or compared as percentages of the same test. Artificial Analysis profile
Vals reports Large 4 at 48.05% ±1.11, ranking 32nd of 44 models in the displayed index. The index combines industry-oriented evaluations across finance, coding, legal and tax work with economic weighting. It is not a universal probability that the model will complete your task.
The small numerical distance between Large 4 and Qwen 3.8 Max, or between Large 4 and DeepSeek V4 Pro, should not support a sweeping rank claim. The displayed uncertainty and different workload results matter more than a few tenths of a point. The larger gaps to the leading frontier models are harder to dismiss, although your specific application may favor Large 4.
Mistral’s announcement describes it as the best open-weight model from the US or Europe on aggregated benchmarks. That is a geographically bounded vendor claim, with an announced future weight release. It does not mean best among all models, best among all globally developed open models or downloadable today. The broader independent tables above are the more useful starting point for procurement.
Finance: a credible result with a qualified Astra comparison
| Model | Finance Agent v2 accuracy | Reported cost per test | Reported latency |
|---|---|---|---|
| Gemini 4 Argon | 65.40% | $4.38 | 16m 21s |
| Claude Opus 5.5 | 58.59% | $9.22 | 33m 08s |
| Claude Sonnet 5.5 | 58.10% | $6.55 | 31m 43s |
| GLM 5.3 | 55.84% | $1.07 | 15m 50s |
| Mistral Large 4 | 54.68% | $1.20 | 26m 02s |
| GPT-6 Astra | 53.54% | $6.82 | 11m 33s |
| Kimi K3 | 53.11% | $1.91 | 4m 46s |
| GPT-6.1 Sol | 52.03% | $1.62 | 15m 25s |
| DeepSeek V4 Pro 0813 | 50.39% | $0.88 | 16m 15s |
Source: Vals Finance Agent v2. Latency describes this evaluator’s agent workloads, not chat response time or a service-level guarantee.
Large 4 exceeds Astra by 1.14 percentage points in this snapshot, with a substantially lower reported test cost. It also takes more than twice as long in the displayed latency measure. That tradeoff matters: a research job that runs overnight and an analyst waiting interactively have different requirements.
This result supports evaluating Large 4 for financial workflows; it does not establish a general finance crown. Gemini, both Claude models and GLM 5.3 score higher in the same table. Financial reasoning accuracy also says nothing by itself about permission controls, auditability, data provenance or the suitability of generated advice.
Legal agents: one of the most interesting launch results
| Model | Harvey Legal Agent Benchmark task pass rate | Reported cost per test |
|---|---|---|
| Gemini 4 Argon | 19.58% | $10.55 |
| Mistral Large 4 | 15.83% | $4.18 |
| Kimi K3 | 12.92% | $3.83 |
| Qwen 3.8 Max | 10.42% | $2.37 |
| GLM 5.3 | 8.33% | $4.30 |
| DeepSeek V4 Pro 0813 | 7.50% | $0.17 |
| GPT-6 Astra | 5.42% | $26.16 |
| GPT-6.1 Sol | 5.42% | $3.93 |
| Claude Opus 5.5 | 3.75% | $21.38 |
| Claude Sonnet 5.5 | 2.92% | $16.61 |
Source: Vals Harvey Legal Agent Benchmark.
Large 4 ranks sixth of 75 models in the displayed results. That is a meaningful specialist signal and substantially better than its broad-index rank. The low absolute pass rates also show how demanding the task-level test is. A 15.83% result is not a claim that legal professionals can delegate the other 84.17% safely, nor a measurement of lawyer replacement.
Different legal scores can refer to different units: passing an entire task, satisfying an individual rubric criterion or answering a research question. Do not compare those percentages as though they measure the same outcome. Large 4’s separate Legal Research Bench result is 31.73%, ranking 38th of 74, which further cautions against converting one strong agent benchmark into a universal legal superiority claim. Vals model results
For buyers, the next step is a controlled evaluation of actual contracts, research questions and citation requirements, with qualified review. The benchmark makes that experiment worthwhile. It does not remove the review requirement.
Coding agents: the harness changes the conclusion
Mistral’s launch cites Artificial Analysis results of 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Their arithmetic mean is 49.8%, the cited Coding Agent Index. These are launch-reported AA results; we have not independently reproduced the runs. Mistral announcement
AA’s coding index methodology combines three benchmarks with equal benchmark weighting. Its current evaluation describes 113 DeepSWE tasks, 66 Terminal-Bench 4 tasks and 124 SWE-Atlas tasks, with results averaged across three attempts per task. A composite helps summarize that test set, but it can conceal a model’s much weaker terminal performance. Artificial Analysis coding agents
Vals’ separate Terminal-Bench 4 evaluation uses Mini-SWE-agent and reports a different Large 4 result:
| Model | Vals Terminal-Bench 4 accuracy | Reported cost per test |
|---|---|---|
| Claude Opus 5.5 | 65.15% | $13.20 |
| Claude Sonnet 5.5 | 64.14% | $16.51 |
| GPT-6 Astra | 59.60% | $9.58 |
| GPT-6.1 Sol | 55.05% | $1.72 |
| GLM 5.3 | 38.89% | $9.37 |
| Qwen 3.8 Max | 34.34% | $10.63 |
| Mistral Large 4 | 22.73% | $7.90 |
| Kimi K3 | 17.17% | $12.02 |
| DeepSeek V4 Pro 0813 | 14.14% | $3.31 |
Source: Vals Terminal-Bench 4. This is a separate evaluation from the AA result cited by Mistral.
The 22.73% and 28.3% Large 4 numbers should remain separate. Averaging them would invent a new benchmark. Differences in task selection, scaffolding and execution budgets can explain differences between evaluators, but we have not isolated the cause here.
Vals also reports 78.40% on Vibe Code Bench v1.1, 30.56% on Code Migration and 14.26% on Vibe Code Bench 1–100. Those are distinct tasks and versions. The harder 1–100 score cannot be substituted for v1.1 or compared against a competitor’s v1.1 result without labeling the difference. Vals profile, Vibe Code benchmark
The practical conclusion is restrained. Large 4 can be a credible coding candidate, particularly where deployment constraints matter. The independent terminal table gives a strong reason to retain a frontier coding agent for difficult autonomous work and compare cost per accepted change, including repair attempts and reviewer time.
Cybersecurity, vision and safety: promising claims that need careful labels
Mistral reports 82% on CyberGym-E2E, 93% on Cybench and a top-five position in the Artificial Analysis Cyber Index. These are the company’s launch claims. They should not be presented as Kingy-run tests, and the 40-task Cybench result should not be treated as a comprehensive security audit. Mistral launch
AA’s Cyber Index focuses on defensive work using CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA: reproducing and patching real vulnerabilities. Its methodology also explains that safety refusals can receive zero scores. A lower score can therefore reflect a refusal policy as well as an inability to solve a problem. The benchmark does not collapse those two causes into an explanation of commercial suitability. AA Cyber Index, methodology
Mistral describes additional access for selected cybersecurity partners, with reduced moderation and expanded cyber capabilities. Results from such access must be distinguished from the public API preview. Model identity alone is insufficient if policy configuration or tool access differs. We are watching for independently published configurations and results that clarify this point.
For vision, Mistral reports 42% on Dense200, compared with 41% for GPT-6 Astra in its comparison. That is one percentage point on one test. Without uncertainty estimates and broader replicated evaluations, it supports a promising visual-grounding result, not a general claim that Large 4 is the best vision model.
The supported product story is text-and-image input with text output, including document Q&A and structured responses. Do not infer native audio, video generation or image generation from the word “multimodal.” A useful business evaluation would measure document extraction, bounding or grounding accuracy, small text, charts and citation consistency separately.
Mistral also reports 93.3% attack resistance on B3 and 1.691 out of 2 on KORABench. Those are named evaluations under reported conditions, not a promise that the model resists every prompt injection. Production agents still need scoped permissions, output validation and human approval where the action warrants it. A better model score does not make a tool connection safe by itself.
API pricing: the launch discount changes the comparison
All figures below are US dollars per million tokens, checked October 6. Prices exclude taxes and additional tools or infrastructure.
| Mistral Large 4 token type | Current displayed rate | Crossed-out rate on the model card |
|---|---|---|
| Input | $0.68 | $1.36 |
| Cached input | $0.07 | $0.14 |
| Output | $2.09 | $4.18 |
The current values also appear in the Standard pricing table, so this is not simply a comparison between Standard and Batch modes. The displayed reduction is 50%. We have not verified an expiration date or a contractual guarantee that these preview rates will persist. Budget-sensitive deployments should retain both scenarios. Mistral pricing, model card
What a request costs
The uncached token formula is straightforward:
cost = input_tokens × input_rate / 1,000,000 + output_tokens × output_rate / 1,000,000
| Illustrative workload | Current displayed rates | Crossed-out rates |
|---|---|---|
| 10,000 input + 2,000 output tokens | $0.01098 | $0.02196 |
| 100,000 input + 10,000 output tokens | $0.08890 | $0.17780 |
| 900,000 input + 50,000 output tokens | $0.71650 | $1.43300 |
| Monthly: 10M input + 2M output | $10.98 | $21.96 |
| Monthly: 100M input + 20M output | $109.80 | $219.60 |
| Monthly: 1B input + 200M output | $1,098.00 | $2,196.00 |
These are calculated token budgets, not measured invoices. The long request assumes an endpoint that actually supports the advertised 1M context and enough room for all messages and output. Agent loops can process the same growing conversation repeatedly, so a 100K document does not imply only 100K billed input tokens over an entire job.
Cached input can reduce a repeated-prefix workload substantially. Suppose a request has 100K input tokens, 90K qualify as cache hits and it produces 2K output tokens. At the displayed rates, 10K uncached input costs $0.00680, 90K cached input costs $0.00630 and output costs $0.00418: $0.01728 total, compared with $0.07218 if all the input were uncached.
That example assumes the cache hits actually qualify. It does not model cache writes, expiry, changed prefixes, tool charges or a guaranteed hit rate. Check the applicable caching rules before treating the saving as a forecast. Likewise, we do not stack a separate batch reduction on top of the preview discount without verified terms for the exact route.
Current token prices against frontier and lower-cost alternatives
| Model/provider | Uncached input /M | Output /M | Cost of 10K input + 2K output |
|---|---|---|---|
| Mistral Large 4, current displayed | $0.68 | $2.09 | $0.01098 |
| Mistral Large 4, crossed-out rates | $1.36 | $4.18 | $0.02196 |
| GPT-6 Astra, Standard short-input tier | $10.00 | $50.00 | $0.20000 |
| GPT-6.1 Sol, Standard short-input tier | $2.00 | $10.00 | $0.04000 |
| Claude Sonnet 5.5, Standard | $2.00 | $10.00 | $0.04000 |
| Claude Opus 5.5, Standard | $4.00 | $20.00 | $0.08000 |
| GLM 5.3, Z.ai | $1.40 | $4.40 | $0.02280 |
| DeepSeek V4.1 Flash, peak | $0.30 | $1.20 | $0.00540 |
| DeepSeek V4.1 Flash, off-peak | $0.15 | $0.60 | $0.00270 |
| DeepSeek V4 Pro 0813, peak | $1.32 | $3.96 | $0.02112 |
| DeepSeek V4 Pro 0813, off-peak | $0.66 | $1.98 | $0.01056 |
| Mistral Large 3 | $0.50 | $1.50 | $0.00800 |
First-party sources: OpenAI Astra, OpenAI Sol, Anthropic pricing, Z.ai pricing, DeepSeek pricing, Mistral pricing.
The OpenAI examples use requests within the standard short-input tier; requests above 272K input tokens have different pricing multipliers. DeepSeek’s peak and off-peak schedules are separate tariffs, not prices you can select independently of when a request runs. Provider routes, caching, batch modes and priority service also matter. This table deliberately compares an uncached, short example rather than pretending every provider’s billing system is identical.
At current displayed rates, Large 4’s illustrative request costs about 3.64 times less than Sol or Sonnet and 18.21 times less than Astra. These ratios describe token prices for this input/output mix, not quality-adjusted savings. A harder request can change the ratio through reasoning output, retries, cache behavior and task completion.
For a deeper comparison of the expensive generalist alternative, see our GPT-6 Astra versus Claude Sonnet 5.5 analysis.
Why evaluator task cost is more revealing than the tariff
Vals reports $13.78 per Vals Index test for Large 4, compared with $3.24 for Sol, despite Sol’s higher current headline token rates. Large 4’s reported index latency is 105 minutes 50 seconds. The evaluator’s workload includes agent behavior and its own configuration; it is not a single small chat request.
Those results use the $1.36/$4.18 Large 4 rates listed in the evaluator’s snapshot. Applying a 50% tariff reduction to an identical token trace would mathematically halve its token charge. That would be a hypothetical recalculation, not a new observed Vals result. We therefore retain the published $13.78 and do not relabel it as a live invoice estimate.
For a business, the useful denominator is accepted work: total model, tool, infrastructure and review cost divided by successful, usable outcomes. Track failed attempts too. A lower-priced model can be expensive if it makes longer calls or requires more repair. A more expensive model can be economical when it finishes cleanly. The finance and legal tables give Large 4 stronger cost-quality arguments than the broad index does.
Artificial Analysis reports roughly 116.1 output tokens per second in its profile, alongside a reported $1.13 Intelligence Index task cost at the higher launch rates. Throughput is a useful serving signal but does not include all the time spent waiting, reasoning and calling tools in a business agent. It cannot be compared directly with Vals’ end-to-end job latency. AA profile
Open weights, licensing and the self-hosting cost question
The word “open” needs a date and an artifact. On October 6, the API is public, while the weights are promised for the end of the month. Mistral’s model documentation lists the download and license as coming soon. We cannot verify commercial redistribution rights, modification terms or supported deployment formats from an unreleased license.
Mistral Large 3 is an instructive predecessor: its documentation specifies 675B total parameters, 41B active parameters, a 256K context window and Apache 2.0 licensing. Large 4 increases the advertised total count by about 56%, the active count by about 20%, and the card’s context window by roughly four times. None of those changes guarantees a proportional quality increase, and Large 3’s license does not automatically become Large 4’s license. Large 3 model documentation
Comparisons with “open source” also need care. An API model associated with an open-model ecosystem is not necessarily a downloadable checkpoint. In particular, the Qwen 3.8 Max row above names an evaluated API model; it should not be silently replaced by a smaller open Qwen checkpoint. Model size, release version, license and inference provider belong in the comparison.
For a future self-hosted deployment, compare the entire operating bill: accelerator rental or depreciation, CPU and RAM, networking, storage, monitoring, redundancy, engineer time and utilization. Running a large reserved cluster for sporadic requests can be much more expensive than buying tokens. Sustained high utilization, private-data constraints or specialized customization can change that calculation.
We have no verified Large 4 self-hosting throughput, quantized quality measurements or production-serving recipe. Until the release arrives, there is no evidence-backed break-even price to publish. We also cannot estimate training cost from 3,800 GPUs alone: duration, utilization, ownership, power, networking and accounting treatment are missing.
European deployment and data residency
European training infrastructure is relevant to procurement, but it does not establish that every API request, log and account record remains in a chosen jurisdiction. Mistral offers regional inference routes for the EU and US and documents a 10% regional pricing uplift across applicable token charges. Its global endpoint is distinct from a region-specific route. Mistral regional inference documentation
The regional documentation specifies 1.1 times standard list pricing. Applied to the higher $0.02196 short-request example, that is $0.024156. We have not verified whether the preview discount also applies to the selected regional route. Confirm model availability and the applicable tariff before making a regional deployment budget.
The residency documentation also distinguishes inference processing from control-plane information such as accounts, billing and analytics. Buyers should check those boundaries against their actual requirements. European model development, regional inference and a fully controlled private deployment are related purchasing considerations with different operational consequences.
The potential attraction of Large 4 is having a European provider now and a possible private-deployment option later. Whether that option fits a particular organization depends on the released license, serving requirements and contractual controls.
How to evaluate it without wasting the preview window
Start with a narrow set of tasks drawn from your own workload. For legal work, require traceable citations and correct treatment of the relevant clauses. For finance, require sourced calculations and consistent units. For coding, require passing the project’s tests and a reviewable change. Define success before looking at the outputs.
Use the same evidence, tools, execution limits and acceptance criteria for each model. Record reasoning effort, model version, date, token usage, retries and wall-clock time. A powerful model with a weak agent wrapper and a weaker model with a well-tuned wrapper are not a clean model comparison.
The official usage example supports selecting high reasoning effort. A minimal illustrative API call is:
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[{
"role": "user",
"content": "Summarize the supplied evidence and cite each conclusion."
}],
reasoning_effort="high",
)
print(response.choices[0].message.content)
This follows the official model usage documentation. It is an illustrative starting point, not a tested integration or a completed benchmark. Set suitable budgets and output limits for your application, and avoid sending restricted information through an unapproved route.
For agents, inspect the failures rather than just the aggregate score. Was the answer wrong, a tool call malformed, a source unavailable, the budget exhausted or the action refused? Those failures require different remedies. Keep a frontier-model comparison and a cheaper-model comparison so the decision includes both quality and cost.
What we are watching over the next 72 hours
The initial publication leaves several questions open:
- Independent replication: additional evaluator results, revised tables and disclosed reasoning or agent settings, especially in legal, finance and cybersecurity.
- Context configuration: why Mistral advertises 1M while evaluators list roughly 512K, and what the public endpoint actually permits.
- Price durability: whether the current 50% reduction has an expiration date, how batch pricing applies and whether evaluator cost tables adopt the lower tariff.
- Weights and license: a concrete release date, downloadable artifacts, the final license and supported formats. The announced release is later than this initial 72-hour watch.
- Operational evidence: time to first token, throughput under realistic concurrency, tool-call reliability, long-document accuracy and total task costs.
- Preview changes: whether model updates materially change the scores or require identifying a new version rather than overwriting old results.
The article will preserve a dated record of material changes. A source check that finds nothing new will not create a cosmetic “updated” date. An unavailable source will not be treated as evidence that the underlying result is unchanged.
Dated update log
- October 6, 2026, 09:08 PDT / 16:08 UTC — Added provider-specific limits and access details. Vercel’s directly fetched AI Gateway table lists 524K context, 262K maximum output and the discounted token rates for
mistral/mistral-large-4. This differs from its search-indexed 1M label and does not resolve Mistral’s official 1M versus evaluator limits. Rechecked official prices and published evaluations; the cited figures remain unchanged. Mistral’s weights and license tabs still say “Coming soon.” Provider listing, official model card. - October 6, 2026, 06:50 PDT / 13:50 UTC — Initial evidence snapshot. Added the official architecture and preview status, current discounted rates and higher evaluator price basis, independent broad/finance/legal/coding comparisons, and the unresolved 1M versus roughly 512K context discrepancy. No Kingy model inference tests have been run.
Sources and reporting method
We inspected Mistral’s announcement, model documentation, pricing and regional inference terms; the independent Artificial Analysis profile and Vals results; and the first-party competitor pricing linked above. The original Mistral announcement on X is the starting point for this coverage.
Benchmark figures are published results attributed to their operator or, where explicitly labeled, Mistral’s launch claims. Dollar examples are our arithmetic using the cited rates. We have not run paid model calls, matched local benchmarks or self-hosting trials. Kingy local tests: 0. The launch assessment can change as the preview evolves and independent evidence expands.
Trending on Kingy
Keep reading with the stories getting the most attention now.
The Kingy Brief
Get The Kingy Brief.
Every Friday: the AI launches worth your time, what we actually tested, and one thing to try. Free. Unsubscribe anytime.
Free · Double opt-in · Unsubscribe anytime
