AI News

Kimi K3 Benchmarks, Specs and Pricing: How It Ranks vs Frontier and Open Models

Evidence summary

  • What changed: Moonshot AI has now released Kimi K3 as a 2.8-trillion-parameter, native multimodal model with downloadable weights and 104 billion activated parameters.
  • Current access and cost: Kimi K3 is available through Kimi’s hosted products and API. The published API rates are $0.30 per million cached input tokens, $3 per million uncached input tokens and $15 per million output tokens.
  • Evidence status: Moonshot specifications and benchmark tables are vendor-reported. Artificial Analysis provides the named independent comparison snapshot. The official model repository now publishes the full weights, 104B activated-parameter count, deployment guidance and custom Kimi K3 License.
  • Kingy-tested: No hands-on K3 workflow test is claimed here. Kingy compared published results, pricing and methodology disclosures.
  • Last checked: July 27, 2026.
  • Related K3 guides: See the official-weight download guide, hardware and local-deployment guide, and Kimi K3 commercial-license analysis.

Published July 16 and updated July 27, 2026. This analysis separates independent evaluation data from Moonshot AI’s claims. The weights, activated-parameter count, licence and deployment guidance were rechecked against the official release.

Kimi K3 launches as the #4 tested configuration and effectively the #3 model family on Artificial Analysis Intelligence Index v4.1. Its 57.1 score trails Claude Fable 5 with Opus 4.8 fallback at 59.9 and GPT-5.6 Sol Max at 58.9, while sitting only 0.54 points behind Sol xhigh. Moonshot’s 2.8-trillion-parameter model also brings native image and video understanding, a 1,048,576-token context window and unusually competitive coding and agent scores. For the unresolved provenance question, read our separate audit of whether Kimi K3 distilled Claude Fable 5.

Moonshot’s detailed benchmark table mixes Kimi Code, Claude Code and Codex harnesses, and several K3 results were not visible on the referenced public leaderboards at launch. The weights are now available, but the broadest performance claims still need reproducible third-party testing.

What you need to know

  • Independent result: Artificial Analysis gives Kimi K3 Max a 57.1 score: #4 among tested configurations and effectively #3 when only the best configuration per model family is counted.
  • Scale: Moonshot reports 2.8 trillion total parameters and 104 billion activated parameters, with 16 of 896 routed experts selected per token.
  • Context and modality: K3 accepts text, images and video and offers a 1,048,576-token context window.
  • API price: $0.30 per million cached input tokens, $3 per million uncached input tokens and $15 per million output tokens.
  • Availability: K3 is live through Kimi products and the Kimi API. The current API supports low, high and max reasoning effort, with max as the default.
  • Open-weight status: Full weights are now downloadable from Moonshot’s official repository under the custom Kimi K3 License. The broad grant allows commercial use, modification and redistribution, but the licence adds conditions for certain large Model-as-a-Service businesses and very large products.

Evidence key: “Independent” means a result run by Artificial Analysis rather than the model maker. “Official” means a specification or score published by Moonshot, OpenAI or Anthropic and not necessarily independently reproduced. “Current release” refers to the July 27 weight and model-card update; benchmark results remain labelled by whether Moonshot or an independent evaluator produced them.

Kimi K3 specifications and pricing

Moonshot describes K3 as its most capable model and the first “open 3T-class” system. It combines Kimi Delta Attention (KDA), Attention Residuals (AttnRes) and Stable LatentMoE. KDA is intended to make attention scale more efficiently across long sequences, while AttnRes selectively retrieves representations from earlier layers instead of accumulating them uniformly through depth. Moonshot claims the overall design improves scaling efficiency by about 2.5 times over Kimi K2; that figure remains vendor-reported. The official model repository now includes the model card, full report and deployment details. The model-specific K3 API quickstart separately confirms image and video inputs, the one-million-token window and the launch-time request limits.

SpecificationKimi K3 launch detailEvidence/status
Total parameters2.8 trillionMoonshot launch disclosure
Activated parameters104 billion; 16 of 896 routed experts selected per tokenOfficial Kimi K3 model card
Core architectureKDA, AttnRes, Stable LatentMoE, Gated MLA; 93 layersOfficial model card and report
Context window1,048,576 tokensOfficial API and launch materials
Input modalitiesText, images and videoOfficial launch materials and API examples
Output modalityTextOfficial launch materials
Completion limit131,072 tokens by default; configurable up to the remaining 1,048,576-token contextOfficial API quickstart
Reasoning effortLow, high and max; max is the defaultOfficial model card and API guidance
Training/inference formatsQuantization-aware training from SFT; MXFP4 weights and MXFP8 activationsMoonshot architecture disclosure
API modelkimi-k3Available through Kimi API with tool calling and structured-output support
Weights and licenceFull weights available under the custom Kimi K3 LicenseReleased July 27, 2026
Recommended self-host deploymentvLLM, SGLang or TokenSpeed; Moonshot’s launch guidance recommends supernodes with 64 or more acceleratorsOfficial model card and launch guidance

Kimi K3 API pricing

Token typePrice per 1M tokensPractical implication
Cached input$0.30Favors stable prompts and reused repository or document prefixes
Uncached input$3.00Applies when the prompt prefix does not hit cache
Output$15.00Long reasoning traces can dominate total cost

Moonshot says its official API sees a cache-hit rate above 90% on coding workloads. Treat that as a service-side observation, not a guaranteed rate for every application. Your result will depend on prompt stability, agent design and how often the reusable prefix changes.

How the list price compares

ModelContext / max outputInput modalitiesUncached input / 1MOutput / 1MAvailability or pricing caveat
Kimi K31,048,576 / 131K default, configurable higherText, image, video$3$15Max reasoning only at launch; cached input is $0.30
GPT-5.6 Sol1.05M / 128KText, image$5$30Prompts above 272K receive a 2× input and 1.5× output multiplier for the full request
Claude Fable 51M / 128KText, image$10$50Safeguarded topics may fall back to Opus 4.8

At list price, K3 input is 40% cheaper than Sol and its output is 50% cheaper. Against Fable 5, both K3 rates are 70% lower. Those comparisons describe token prices, not completed-task cost: Artificial Analysis shows why K3’s substantially higher output-token use must be included in any budget model.

Independent benchmark: where Kimi K3 ranks

Artificial Analysis Intelligence Index v4.1 ranking of Kimi K3 against Claude Fable 5, GPT-5.6 Sol, Grok 4.5, Gemini 3.1 Pro and leading open-weight models
Artificial Analysis v4.1 places Kimi K3 Max fourth by configuration and effectively third by model family. Fable’s leading result allows an Opus 4.8 fallback, while two Sol effort settings occupy the next two configuration slots. Source: Artificial Analysis; chart by Kingy AI.

The strongest launch-day evidence is Artificial Analysis’s independent Kimi K3 evaluation. Its v4.1 index combines nine evaluations across agentic work, terminal use, scientific reasoning, knowledge and long-context reasoning rather than relying on one coding leaderboard. The published methodology weights agents at 34%, coding at 24%, scientific reasoning at 24% and general capability at 18%.

RankModel / effortAA v4.1Release statusHow to read it
1Claude Fable 5, adaptive Max with Opus 4.8 fallback59.9ProprietaryFallback-enabled production configuration
2GPT-5.6 Sol Max58.9ProprietaryBest single-model Sol setting
3GPT-5.6 Sol xhigh57.7ProprietarySame model family as rank 2
4Kimi K3 Max57.1API and open weights availableEffectively #3 model family
5Claude Opus 4.8, adaptive Max55.7ProprietaryPure Opus result without Fable fallback framing
6Grok 4.5 High53.8ProprietaryK3 leads by 3.3 points
7GLM-5.2 Max51.1Open weights · MITCurrent open-weight leader
8Muse Spark 1.1 xhigh50.6ProprietaryCheaper and faster, lower composite
9Gemini 3.1 Pro Preview46.5Proprietary previewStronger on several knowledge and vision measures
10MiniMax M344.4Open weights · restricted termsCommercial conditions apply
11DeepSeek V4 Pro Max44.3Open weights · MITLarge deployable open model
12Kimi K2.644.2Open weights · modified MITK3 improves by 12.9 points
13Inkling xhigh40.7Open weights · Apache 2.0One-million-token context
Why “#4 configuration” but “#3 model family”? The leaderboard counts Sol Max and Sol xhigh separately. If each model family contributes only its best setting, Fable is first, Sol second and K3 third. GPT-5.6 ultra is excluded because it is a multi-agent system rather than a like-for-like single-model run.

K3’s 0.54-point gap from Sol xhigh is too small to treat as a decisive capability difference. Artificial Analysis estimates the overall index’s 95% confidence interval at less than roughly one point. The 1.78-point gap from Sol Max and 2.75-point gap from the Fable configuration are clearer, although still not a substitute for workload-specific testing. The index is text-only and English-only; it cannot establish a multimodal or multilingual ranking by itself.

Cost and token use across the three leaders

ConfigurationAA v4.1Output tokens across the indexTotal evaluation cost
Claude Fable 5 Max, Opus 4.8 fallback59.987 million$5,630.52
GPT-5.6 Sol Max58.970 million$2,824.00
Kimi K3 Max57.1130 million$2,690.80

K3 used far more output tokens to complete the evaluation—about 1.9 times Sol’s total and 1.5 times Fable’s—yet its lower token prices kept the evaluation bill slightly below Sol and far below the Fable configuration. That is promising price-performance, not proof that K3 will be cheaper in every deployment: agent loops, cache reuse, retries and output length can move real costs sharply. For deeper context on the two proprietary leaders, see Kingy AI’s analyses of GPT-5.6 Sol’s benchmarks and specifications and Claude Fable 5’s benchmark and pricing picture.

Kimi K3 vs Grok 4.5, Muse Spark 1.1 and Gemini 3.1 Pro

The composite score hides important reversals. The table below uses Artificial Analysis’s standardized results for four frontier models that are easy to omit when focusing only on the top three. K3 leads all three alternatives on the overall, agentic and coding indices, but Gemini remains stronger on scientific knowledge, factual reliability and visual reasoning.

Independent metricKimi K3 MaxGrok 4.5 HighMuse Spark 1.1 xhighGemini 3.1 Pro Preview
AA Intelligence Index v4.157.153.850.646.5
Agentic Index50.145.737.521.4
Coding Index76.272.571.368.8
Terminal-Bench 2.185.0%81.7%77.9%73.8%
Humanity’s Last Exam, text only44.4%40.3%45.1%44.7%
GPQA Diamond93.5%93.1%89.8%94.1%
AA-Omniscience Index18.426.418.032.9
MMMU-Pro80.5%80.4%Not reported82.4%
Context window1.05M500K1.05M1M
Input / output price per 1M tokens$3 / $15$2 / $6$1.25 / $4.25$2 / $12

The practical read: K3 is the stronger candidate for coding agents and long-horizon tool work in this group. Gemini is the safer counterexample to any “K3 wins everything” claim; it leads on GPQA, AA-Omniscience and MMMU-Pro. Muse is materially cheaper and edges K3 on Humanity’s Last Exam, while Grok’s lower overall score still comes with better knowledge reliability. Model selection should follow the task profile, not the composite rank alone.

Moonshot’s coding results are impressive—and not apples to apples

Moonshot-reported coding benchmark comparison for Kimi K3, Claude Fable 5, GPT-5.6 Sol and GLM-5.2
Selected coding scores reported in Moonshot’s Kimi K3 launch post. Different models and tests use Kimi Code, Claude Code, Codex or other harnesses, so the bars should not be read as a controlled head-to-head. Chart by Kingy AI.
BenchmarkKimi K3 MaxFable 5 Max, fallback allowedGPT-5.6 Sol MaxGLM-5.2 Max
DeepSWE67.570.073.046.2
Program Bench77.876.877.663.7
Terminal-Bench 2.188.384.688.882.7
FrontierSWE81.286.671.367.3
SWE Marathon42.035.039.013.0

On Moonshot’s table, K3 leads Program Bench by 0.2 points over Sol and SWE Marathon by three points. It lands within 0.5 points of Sol on Terminal-Bench 2.1, while Fable leads FrontierSWE and Sol leads DeepSWE. This is not a story of K3 winning every coding benchmark; it is a story of the model staying competitive across several different kinds of long-horizon software work.

The testing conditions limit stronger conclusions. K3 uses the Kimi Code harness on DeepSWE, Program Bench, Terminal-Bench and FrontierSWE. Sol often uses Codex, while Fable and other Claude models use Claude Code or Terminus depending on the test. Fable’s safeguards can route some sessions to Opus 4.8, so its rows describe the production configuration rather than the underlying Fable model alone. On Terminal-Bench, Moonshot selects the best reported harness for several competitors. SWE Marathon mixes Claude Code for K3 and the Anthropic models with Codex for Sol. At publication time, the public DeepSWE and Program Bench pages also did not yet show a K3 entry that independently reproduced Moonshot’s number.

Moonshot says its K3 runs use Max reasoning, temperature 1.0 and top-p 1.0. The public API quickstart showed top-p 0.95 when checked on launch day. Even that small configuration mismatch is enough reason for evaluators to publish exact prompts, harness versions, tool permissions, token budgets and retry rules before treating a reproduction as definitive.

Agent and vision evaluations

Moonshot’s broader table suggests K3 is strongest when coding, browsing, tools and visual reasoning meet. On its reported agent evaluations, K3 scores 1,668 Elo on GDPval-AA v2, behind Fable at 1,760 and Sol at 1,748 but above Claude Opus 4.8 at 1,600. It reaches 1,548 on AA-Briefcase, second to Fable’s 1,583 and ahead of Sol’s 1,495. K3 also leads the compared group on Automation Bench at 30.8 and BrowseComp at 91.2, while Fable leads JobBench at 57.4.

Moonshot-reported evaluationKimi K3Fable 5, fallback allowedGPT-5.6 SolWhat it suggests
GDPval-AA v2 Elo1,6681,7601,748K3 is competitive but trails both frontier leaders
AA-Briefcase Elo1,5481,5831,495K3 sits between Fable and Sol on long-horizon knowledge work
Automation Bench30.829.129.7K3 holds a narrow reported lead
JobBench52.957.446.5Fable leads; K3 clears Sol
SpreadsheetBench 234.834.732.4Near tie with Fable under different harnesses
BrowseComp91.288.090.4K3 posts the best reported score

Do not combine Moonshot’s 30.8 Automation Bench figure with Artificial Analysis’s separate AutomationBench-AA result, where K3 scores 52.7. The names are similar, but the implementations and scales differ. The first is part of Moonshot’s launch table; the second is an independent AA workflow evaluation.

Native vision also appears useful rather than decorative. Moonshot reports K3 at 81.6 on MMMU-Pro versus 81.2 for Fable and 83.0 for Sol. With Python tools, K3 reaches 91.3 on CharXiv RQ, between Fable’s 93.5 and Sol’s 89.1; on MathVision with Python, K3 and Sol tie at 97.8 behind Fable at 98.6. K3’s 91.1 on OmniDocBench leads Fable at 89.8 and Sol at 85.8. These are still vendor-collected comparisons, and some visual tests are averaged over only three runs. They support a serious multimodal capability claim, not a universal vision crown.

Speed, latency, verbosity and real cost

K3’s API begins streaming quickly, but the final answer does not arrive in two seconds. Artificial Analysis measures a 1.99-second time to the first streamed chunk, followed by about 32.24 seconds of reasoning before the first answer token. On its standardized 500-token performance workload, total response time is about 42.30 seconds.

Independent API measureKimi K3Relevant comparisonInterpretation
Output speed62.0 tokens/s72.7 tokens/s peer medianBelow-median generation speed
Time to first streamed chunk1.99 seconds2.60-second peer medianFast initial stream, not a completed answer
Time to first answer token34.24 secondsIncludes 32.24 seconds of reasoningMore representative of perceived wait
Total standardized response time42.30 seconds500-token test responseReasoning dominates end-to-end latency
Index output-token use130 million63 million peer medianK3 is unusually verbose

What representative K3 requests cost

Illustrative workloadK3 with uncached inputK3 if all input hits cacheCalculation
100K input + 10K output$0.45$0.18Input plus output token charges
950K input + 50K output$3.60$1.04Fits inside the 1,048,576-token combined window

For the 100K-input, 10K-output example, GPT-5.6 Sol’s listed rates produce an $0.80 token bill and Fable 5’s produce $1.50, before tool charges or provider-specific caching. Near the context limit, Sol applies long-context multipliers above 272K input tokens, which makes a simple headline-rate multiplication misleading. These calculations are estimates, not task-cost guarantees: K3’s high reasoning-token use can erase part of its list-price advantage.

Is Kimi K3 really open source?

Kimi K3 is now available as open weights under a custom licence. Moonshot’s official model repository publishes the full 2.8-trillion-parameter checkpoint, a 104-billion activated-parameter count, the model card, configuration files and deployment guidance. That release is substantial, but “open weights” is more precise than unqualified “open source” because the custom licence includes commercial scale conditions. See our plain-English Kimi K3 licence analysis.

The released K3 checkpoint changes the open-weight comparison. Its 57.1 Artificial Analysis score is 6.0 points above GLM-5.2 Max. The table below distinguishes standard permissive licences from custom or restricted terms; a downloadable model is not automatically unconditionally open source.

ModelAA v4.1Total / activated parametersContextWeights and licence on July 27
Kimi K3 Max57.12.8T / 104B1MAvailable · custom Kimi K3 License
GLM-5.2 Max51.1753B / 40B1MAvailable · MIT
MiniMax M344.4428B / 23B1MAvailable · community license with commercial conditions
DeepSeek V4 Pro Max44.31.6T / 49B1MAvailable · MIT
Kimi K2.644.21T / 32B256KAvailable · modified MIT
Inkling xhigh40.7975B / 41B1MAvailable · Apache 2.0
Nemotron 3 Ultra37.8550B / 55B262KAvailable · OpenMDW
Qwen3.6 27B37.127.8B dense262KAvailable · Apache 2.0
Qwen3.5-397B-A17B33.7397B / 17B262KAvailable · Apache 2.0
Gemma 4 31B29.430.7B dense256KAvailable · Apache 2.0
gpt-oss-120b High23.8117B / 5.1B131KAvailable · Apache 2.0
Mistral Large 316.0675B / 41B256KAvailable · Apache 2.0

This table also explains why “best” is not synonymous with “largest.” Qwen3.6 27B is far easier to deploy than a multi-trillion-parameter MoE, even though its composite score is lower. MiniMax M3 and K3 are downloadable under custom terms rather than standard MIT or Apache 2.0 licences.

The remaining catch is infrastructure. The model card reports 104 billion activated parameters and native MXFP4 weights with MXFP8 activations, while Moonshot recommends vLLM, SGLang or TokenSpeed for serving. Its launch guidance points to supernodes with 64 or more accelerators. For most teams, open weights will mean specialist inference infrastructure rather than a workstation. See our download and official-weight guide, hardware and local-deployment guide, and open-weight model comparison.

Strengths, limitations and deployment verdict

Where Kimi K3 looks strongest

  • Near-frontier general capability: A 57.1 independent index score places K3 fourth by configuration and effectively third by model family.
  • Long-horizon coding: Its consistency across Program Bench, Terminal-Bench, FrontierSWE and SWE Marathon makes a credible case for repository-scale and tool-heavy work.
  • Vision inside the agent loop: Native visual input, a million-token window and strong chart, document and multimodal scores fit browser, frontend, research and document workflows.
  • API economics: K3 completed the Artificial Analysis suite at a slightly lower reported cost than Sol despite using substantially more output tokens.
  • Weight access: The released checkpoint gives researchers and infrastructure providers unusual control over a near-frontier model.

What should stop an immediate migration

  • Unusual licence terms: The custom Kimi K3 License allows broad commercial use but adds conditions for certain large Model-as-a-Service businesses and very large products.
  • Detailed comparisons are vendor-run: Mixed harnesses and unpublished K3 leaderboard entries make several coding wins provisional.
  • Benchmark-setting caveat: The headline evaluations use max reasoning; lower effort modes may change quality, latency and cost.
  • High token use: Artificial Analysis recorded 130 million output tokens for K3, making generation discipline and cache design important.
  • Knowledge reliability is not class-leading: K3’s 18.4 AA-Omniscience Index trails Grok 4.5 at 26.4 and Gemini 3.1 Pro Preview at 32.9.
  • The headline index is narrow by design: It is an English-only, text-only suite and cannot settle multimodal or multilingual performance.
  • Operational sensitivity: Moonshot warns that K3 expects its full thinking history to be preserved. Switching models mid-session or using an incompatible harness can destabilize quality.
  • Excessive proactivity: Moonshot says K3 may make unexpected decisions when instructions are ambiguous. High-impact tools need strict permissions, confirmations and explicit behavioral boundaries.
  • User experience gap: The company itself acknowledges that K3 still trails Fable 5 and GPT-5.6 Sol in overall user experience.

Verdict: Kimi K3 deserves an immediate controlled pilot for coding agents, deep research, visual document work and long-context automation. It does not yet justify a blanket replacement of Fable 5 or GPT-5.6 Sol. Run the same internal tasks through the same harness, cap tool permissions, record token and retry costs, and score completion quality—not just benchmark rank. For regulated or high-scale commercial deployment, inspect the custom licence, validate the released checkpoint under your own harness and budget for cluster-scale serving.

Frequently asked questions

What is Kimi K3?

Kimi K3 is Moonshot AI’s 2.8-trillion-parameter Mixture-of-Experts model for reasoning, coding, knowledge work and native visual understanding. It supports a 1,048,576-token context window and is available through Kimi products and the Kimi API.

How does Kimi K3 compare with Claude Fable 5 and GPT-5.6 Sol?

Artificial Analysis Intelligence Index v4.1 scores K3 Max at 57.1, GPT-5.6 Sol Max at 58.9 and Fable 5 Max with Opus 4.8 fallback at 59.9. K3 is fourth by configuration because Sol xhigh also ranks above it, but effectively third when only the best setting from each model family is counted.

How much does the Kimi K3 API cost?

The launch rate is $0.30 per million cached input tokens, $3 per million uncached input tokens and $15 per million output tokens. Actual workload cost will depend heavily on cache hits, output length and agent retries.

Is Kimi K3 open source or open weight?

Kimi K3 is now downloadable as open weights under the custom Kimi K3 License. The licence broadly permits commercial use, modification and redistribution, but it is not a standard MIT or Apache licence and includes conditions for certain high-scale commercial deployments.

Can Kimi K3 process images and video?

Yes. K3 accepts image and video input alongside text, and its launch table reports results on MMMU-Pro, CharXiv, MathVision, ZeroBench, OmniDocBench and other multimodal evaluations. Output is text.

Can most companies self-host Kimi K3?

Probably not economically on ordinary infrastructure. Sparse routing means only a subset of experts is selected per token, but the model still contains 2.8 trillion total parameters. Moonshot recommends supernodes with at least 64 accelerators, and it has not yet disclosed the active-parameter count or final distribution formats.

Should teams switch from GPT-5.6 Sol or Fable 5?

Not without an internal bake-off. K3 is close on the independent composite, cheaper than Sol across the Artificial Analysis evaluation run and strong on several agent tests. Sol and the Fable configuration still score higher overall, while K3 has launch-day reproducibility and UX caveats. Route representative production tasks through identical tools and score accuracy, intervention rate, latency and total cost.

Sources and evidence standard

Primary launch specifications, pricing, availability, benchmark settings and limitations come from Moonshot AI’s Kimi K3 launch post, the official K3 API quickstart and the company’s launch announcement on X. The independent composite score, sub-evaluations, token use, speed and cost come from Artificial Analysis’s K3 profile, its v4.1 leaderboard and its published methodology. Comparator values were checked against the AA profiles for Grok 4.5, Muse Spark 1.1 and Gemini 3.1 Pro Preview, plus official model cards and licenses for the open-weight field. Moonshot’s detailed benchmark scores remain vendor-reported unless explicitly identified as independent. Any claim dependent on the future weight release should be rechecked on or after July 27, 2026.