Trending on Kingy
Keep reading with the stories getting the most attention now.
There is no single open-weight coding crown today. There are three—and three different winners.
GLM-5.3 has the strongest vendor-reported launch case. DeepSeek V4 Pro 0813 is the best-documented production endpoint available right now. Kimi K3 is the strongest current checkpoint you can actually download, license and attempt to reproduce.
That split is not a rhetorical escape hatch. It is the only honest way to compare releases whose benchmarked service, public API and downloadable weights do not always refer to the same checkpoint.
Evidence cutoff: August 13, 2026, 11:30 p.m. PDT (August 14 in Z.ai’s publication timezone). This is a source-audited first edition, not a hands-on Kingy benchmark. Kingy did not run the three models through an identical repository-repair suite. Scores below are vendor-reported unless explicitly stated otherwise.
The three crowns
| Crown | Winner | Why | Confidence |
|---|---|---|---|
| Best vendor-reported performance | GLM-5.3 | Leads the three on Terminal-Bench 3.0, PostTrainBench, AutomationBench and the listed cyber evaluations; nearly ties Kimi on Terminal-Bench 2.1 and DeepSWE | Moderate: broad table, but Z.ai operated or assembled the comparison |
| Best model endpoint available today | DeepSeek V4 Pro 0813 | GA endpoint, pinned version disclosure, 1M context, 384K max output, native Responses API, Anthropic compatibility and transparent token pricing | Moderate-high for access and product facts; lower for real-world quality without a common Kingy run |
| Best weights you can download now | Kimi K3 | Exact K3 checkpoint is live, large artifacts are present, serving recipes are linked and the license covers use, modification and redistribution with conditions | High for availability; moderate for reproducibility at practical scale |
Who should choose what?
- Use GLM-5.3 if you want the most ambitious new coding/cyber model and can accept first-party benchmark evidence, a brand-new endpoint and weights that are promised rather than downloadable today.
- Use DeepSeek V4 Pro 0813 if you want a low-cost, production-labelled coding endpoint with strong agent results and broad API compatibility—and do not require the hosted GA model to match a downloadable checkpoint yet.
- Use Kimi K3 if local control, artifact inspection and checkpoint-level reproducibility matter more than having the cheapest API or winning every benchmark row.
What changed on launch day
Z.ai’s GLM-5.3 launch post says the model uses the same 743B base model as GLM-5.2 and attributes the gains to scaled post-training: more environments, more diverse long-horizon tasks and more compute. Z.ai says the exact weights will arrive two weeks after launch, following safety evaluation and hardening. The live release therefore qualifies as an open-weights commitment, not a downloadable open-weight artifact at this cutoff.
DeepSeek’s August 13 changelog records the GA rollout of DeepSeek V4 Pro across its app, web product and API. The unchanged alias deepseek-v4-pro now serves version DeepSeek-V4-Pro-0813. DeepSeek also added native Responses API support and low, high and max reasoning effort controls.
The crucial mismatch is that DeepSeek’s public V4 Pro repository still represents the earlier open checkpoint. Kingy found no official DeepSeek-V4-Pro-0813 weight repository at the cutoff. The GA service is real; exact GA weight parity is not yet established.
Kimi K3 is the mature control. Moonshot’s official Kimi K3 repository and Hugging Face checkpoint expose the model card, weight shards, deployment guidance and the custom Kimi K3 License. Its endpoint and released model share the K3 identity, although any hosted service can still contain operational changes that are not encoded in the checkpoint.
The benchmark table is nuanced—not a clean sweep
The most useful launch comparison is Z.ai’s own table because it puts GLM-5.3, GLM-5.2, Kimi K3 and DeepSeek V4 Pro 0813 in one place. It also undermines the lazy version of Z.ai’s marketing story: GLM does not win every row.
| Benchmark | GLM-5.3 | Kimi K3 | DeepSeek V4 Pro 0813 |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 88.3 | 87.9 |
| Terminal-Bench 3.0 | 28.3 | 17.4 | Not reported |
| DeepSWE v1.1 | 66.9 | 67.5 | 62.7 |
| NL2Repo | 58.0 | 58.0 | 61.1 |
| PostTrainBench | 39.8 | 32.0 | Not reported |
| AutomationBench v1.0.6 | 48.2 | 46.7 | 43.2 |
| CyberGym | 84.5 | 80.0 | 83.3 |
| ExploitGym, 2h / 6h | 105 / 130 | 36 / 70 | Not reported |
| ExploitBench | 54.4 | 32.2 | Not reported |
Source: Z.ai’s GLM-5.3 launch table, checked August 13, 2026. These are not Kingy-run results.
Kimi wins Terminal-Bench 2.1 by one tenth of a point and DeepSWE by six tenths. Those margins are too small to support sweeping claims about practical superiority. GLM’s more persuasive case is breadth: it leads on the newer Terminal-Bench 3.0, the post-training workload, automation and every cyber row where all three have a published result. DeepSeek takes NL2Repo, the repository-generation test.
Even this table is not a neutral tournament. Z.ai’s footnotes identify different harnesses, context limits, timeouts and scoring procedures by benchmark. Some comparisons use Claude Code; some use official evaluation services; some use a model-specific throughput normalization; some average three runs while others report a single run. Missing values mean “not reported in this table,” not zero.
There is another warning sign: DeepSeek’s own GA changelog reports 61.5 on NL2Repo and 31.8 on AutomationBench, while Z.ai’s comparison shows 61.1 and 43.2. The discrepancy likely reflects different evaluation revisions or harness settings, but neither vendor provides enough shared run-level evidence to reconcile it cleanly. Kingy therefore keeps each score tied to its source instead of averaging or selecting the friendliest number.
Why GLM-5.2 is the right baseline
GLM-5.2 is not a co-winner; it is the control that isolates what Z.ai says changed. The downloadable GLM-5.2 checkpoint is MIT-licensed, has a documented 1M-token context and uses the same model family and base architecture.
On Z.ai’s current table, GLM-5.3 moves from 81.0 to 88.2 on Terminal-Bench 2.1, 4.6 to 28.3 on Terminal-Bench 3.0, 46.2 to 66.9 on DeepSWE, 31.7 to 39.8 on PostTrainBench and 26.2 to 48.2 on AutomationBench. Cyber gains are also large: 24.4 to 54.4 on ExploitBench and 29/39 to 105/130 on the two ExploitGym budgets.
Those deltas are internally coherent with Z.ai’s claim of much heavier long-horizon post-training. They still need independent reproduction. A same-family before/after comparison is more informative than a victory lap across unrelated models, but it remains vendor evidence.
Checkpoint and license reality
| Model | Exact launch weights available? | License at cutoff | Reproduction status | Main catch |
|---|---|---|---|---|
| GLM-5.3 | No | Not yet published for 5.3 | Endpoint only; weights promised in two weeks | Open-weight crown cannot be awarded before files and terms exist |
| Kimi K3 | Yes | Kimi K3 License | Model card, shards and serving recipes available | 2.8T total parameters; custom commercial conditions |
| DeepSeek V4 Pro 0813 | No matching GA checkpoint found | GA weight license not applicable yet; prior V4 Pro weights are MIT | GA endpoint reproducible only as a hosted API call | Downloadable weights appear to be the earlier checkpoint, not 0813 |
| GLM-5.2 baseline | Yes | MIT | BF16 and FP8 artifacts plus serving guidance | It is the prior model, not a substitute for 5.3 results |
“Downloadable” is not the same as “easy to reproduce.” Kimi K3 has 2.8T total parameters and 104B active parameters, with MXFP4 weights and MXFP8 activations. Even with compressed weights, it is a data-centre-scale model. Reproduction requires capable inference software, substantial accelerator memory and bandwidth, the right quantization path and the benchmark’s exact harness and budget.
The Kimi license is broad—it permits use, modification, distribution, sublicensing and sale—but it adds conditions. A model-as-a-service business above $20 million in aggregate revenue over a consecutive 12-month period needs a separate Moonshot agreement for commercial use. Very large consumer products can also face attribution requirements. That makes Kimi genuinely downloadable and reusable, but not equivalent to an unconditional MIT release.
Endpoint crown: why DeepSeek wins today
This is the closest call and partly a buyer judgement. GLM-5.3 is callable through Z.ai’s API shape and has rolled out to GLM Coding Plan users, but its standalone pricing page had not yet added a GLM-5.3 row at the cutoff. Z.ai’s launch also forces thinking on and introduces a new migration requirement: clients using thinking.type: "disabled" must enable thinking and choose a reasoning effort before changing the model ID.
DeepSeek’s current model and pricing page is unusually explicit. It pins deepseek-v4-pro to DeepSeek-V4-Pro-0813, documents a 1M context window and 384K maximum output, and supports OpenAI Chat Completions, Anthropic-style messages and the Responses API. Until August 16 at 16:00 UTC, the listed price is $0.435 per million uncached input tokens and $0.87 per million output tokens. After that, off-peak Pro pricing becomes $0.66/$1.98 and peak pricing $1.32/$3.96.
That combination—pinned identity, GA status, multiple interfaces, clear limits and clear prices—earns DeepSeek the endpoint crown. It is not a claim that DeepSeek produces better code than GLM-5.3 or Kimi K3 in every environment. It is a claim that a team can understand what it is calling and budget the service today.
Secondary field: Nemotron 3 Ultra and MiniMax M3
Two models belong in the broader open-weight shortlist, but not in the main three-way launch fight.
| Model | Open artifact status | Scale / context | Why it matters here |
|---|---|---|---|
| NVIDIA Nemotron 3 Ultra | Downloadable BF16 and NVFP4 checkpoints; NVIDIA open-model terms | 550B total / 55B active; 1M context | Strong reproducibility story: NVIDIA publishes checkpoints, training data references, recipes and evaluation tooling |
| MiniMax M3 | Downloadable checkpoint under MiniMax Community License | About 428B total / 23B active; 1M context | Native multimodality plus coding and computer-use positioning in a materially smaller model |
NVIDIA’s Nemotron 3 Ultra model card is unusually detailed about architecture, training stages, datasets, hardware and evaluation tooling. MiniMax’s M3 release and checkpoint combine native image/video understanding, sparse million-token context and downloadable weights.
Neither replaces the main comparison. The launch-window question is specifically whether GLM-5.3 displaced Kimi K3 and DeepSeek V4 Pro in coding. Nemotron and MiniMax are useful alternatives for teams weighting reproducibility, efficiency or multimodality differently.
Efficiency control: Nemotron 3 Nano 30B
Nemotron 3 Nano 30B is a control, not a coequal frontier rival. NVIDIA lists it as a 30B-total, 3B-active hybrid MoE with a 1M context window and both downloadable weights and a free NVIDIA-hosted endpoint. It answers a different question: how much agentic coding utility can be delivered with dramatically less active compute?
If Nano completes a routine task reliably, running a 743B, 1.6T or 2.8T model is not a victory. It is overhead. Frontier comparisons should therefore report cost and accepted-task rate, not only peak capability.
Qwen3.8-Max belongs in an API-only sidebar
Qwen3.8-Max does not belong in the open-weight crown table until an exact checkpoint and governing license matching the benchmarked service are verifiable. Alibaba’s current catalog lists the production qwen3.8-max service and also labels a qwen3.8-2.4t-a95b entry as open source. Kingy’s earlier Qwen3.8-Max launch audit tracked the hosted release. At this cutoff, Kingy did not verify an official downloadable artifact demonstrating that the Max service and the open-source-labelled entry are the same checkpoint. That identity gap is another reason to keep endpoint results separate from weight claims.
Z.ai’s table gives the hosted Qwen line credible results—86.6 on Terminal-Bench 2.1 and 56.6 on DeepSWE—but a hosted benchmark row does not establish checkpoint parity. Until the exact files, license and serving recipe for the benchmarked Max model are verifiable, Qwen remains relevant to the endpoint market, not this article’s downloadable crown.
What happened to Meta’s open-model lead?
Meta is absent because Llama 4 is not a clean launch-window comparator for these repository-scale coding agents. It can support a separate analysis of how Meta lost mindshare in the open-model race, but inserting it here would turn a current product decision into a history lesson. The same rule applies to older open models with recognizable names but no current frontier coding case.
What this first edition does not prove
Kingy cross-audited official launch tables, endpoint documentation, model repositories, artifacts, license text and published harness notes. We did not:
- run an identical repository-repair suite across the three endpoints;
- verify benchmark trajectories or recompute vendor scores;
- deploy the full Kimi K3, DeepSeek V4 Pro or future GLM-5.3 weights;
- measure latency, throughput, tool-call reliability, context degradation or cost per accepted repair;
- test cybersecurity capability against live systems.
The first edition should not be called hands-on. A later update can add one identical, versioned repository-repair suite once GLM-5.3 has an exact accessible route suitable for a no-cost or sponsored-but-independent run. The harness, prompt fixture, tool permissions, reasoning effort, timeout, retry policy and acceptance tests must be held constant.
FAQ
Is GLM-5.3 open weight today?
Not yet. Z.ai says it will publish the weights two weeks after launch, following safety evaluation and hardening. Until the files and license are live, GLM-5.3 is an endpoint with an open-weight commitment.
Does DeepSeek V4 Pro 0813 have downloadable weights?
Kingy did not find an official checkpoint matching the 0813 GA endpoint at the evidence cutoff. The existing DeepSeek V4 Pro repository is downloadable and MIT-licensed, but it should not be assumed to contain the GA post-training update.
Why does Kimi K3 win the downloadable crown if GLM scores higher?
Because the crown is for the best current weights you can obtain, license and reproduce. GLM-5.3’s files do not exist publicly yet. Kimi K3’s do, and its vendor-reported coding results remain extremely close to GLM on Terminal-Bench 2.1 and DeepSWE.
Are these benchmark scores directly comparable?
Only directionally. Harnesses, context limits, timeouts, reasoning settings, revisions, tools and judges differ. Z.ai’s table is useful because it exposes those differences in footnotes; it is not a neutral Kingy rerun.
Which model should a small team try first?
DeepSeek V4 Pro 0813 is the easiest high-end endpoint to budget and integrate today. For local experiments, start with a smaller control such as Nemotron 3 Nano before committing to Kimi-scale infrastructure. Teams requiring exact downloadable frontier weights should evaluate Kimi K3 and MiniMax M3 against their own tasks.
Will the winner change when GLM-5.3 weights arrive?
Possibly. The downloadable crown should be reassessed only after the checkpoint, license, hashes, serving instructions and benchmark parity are public. A release date alone is not enough.
Final verdict
GLM-5.3 has the most impressive launch table, and the shape of its gains matters more than any single score: stronger results across newer terminal work, post-training, automation and cyber tasks suggest that Z.ai’s long-horizon training stack is paying off. But it cannot hold today’s open-weight crown without public weights.
DeepSeek V4 Pro 0813 is the endpoint winner. It is GA, inexpensive, version-labelled and compatible with the interfaces coding teams already use. Its weakness is checkpoint parity: the service advanced faster than the public artifact.
Kimi K3 holds the crown that “open weight” is supposed to mean. You can inspect the repository, obtain the massive checkpoint, read the license and attempt to reproduce the model’s behavior. Its scorecard is still frontier-grade, its artifacts are real and its limitations are visible.
So the answer is three names, not one: GLM-5.3 for the vendor-reported performance crown, DeepSeek V4 Pro 0813 for the endpoint crown, and Kimi K3 for the downloadable crown. Anyone collapsing those into a single leaderboard is comparing promises, services and files as though they were the same product. They are not.
Official sources
- Z.ai: GLM-5.3 launch post and benchmark methodology
- Z.ai: GLM-5.2 official checkpoint
- DeepSeek: V4 Pro 0813 GA changelog
- DeepSeek: current model versions, interfaces and pricing
- DeepSeek: V4 Pro open checkpoint repository
- Moonshot AI: Kimi K3 repository and benchmark notes
- Moonshot AI: Kimi K3 downloadable checkpoint
- Moonshot AI: Kimi K3 License
- NVIDIA: Nemotron 3 Ultra checkpoint
- NVIDIA: Nemotron 3 model endpoints
- MiniMax: M3 launch post
- MiniMax: M3 downloadable checkpoint
- Alibaba Cloud: Qwen3.8-Max model availability
Review by Curtis Pyke. See Kingy’s editorial and sponsorship standards. No vendor supplied hands-on access, benchmark credits or payment for this article.
