Z.ai has released the full GLM‑5.3 checkpoints. Here is what is actually downloadable, what the benchmark table does and does not prove, and why most readers should still use the API.
Originally published: August 13, 2026
Updated: August 28, 2026
Update note: Z.ai’s full GLM‑5.3 weights are live. This revision adds checkpoint sizes, shard counts, license terms, hardware guidance, and current download, vLLM, and API recipes. Ox Alpha became GLM‑5.3‑Flash, not full GLM‑5.3.
The verdict
Researchers and infrastructure teams can now inspect, adapt, and deploy Z.ai’s strongest coding model. It is not easy to run.
The official FP8 repository is a 755.7 GB transfer spread across 141 weight shards. The BF16 version is about 1.5 TB across 282 shards. Z.ai’s current vLLM deployment recipe starts with an eight-accelerator server; a normal gaming PC or Mac is not a realistic home for the official full checkpoint.
For most teams, Z.ai’s API is the rational first step. Self-host when control, data location, customization, or sustained utilization justifies datacenter infrastructure. Choose GLM‑5.3‑Flash when native vision, an MIT license, or a smaller checkpoint matters more than the full model’s coding profile.
One correction belongs above the fold: the full 744B-class GLM‑5.3 was never Ox Alpha. Z.ai says ox-alpha became the smaller, multimodal GLM‑5.3‑Flash. Kingy’s definitive GLM‑5.3‑Flash, formerly Ox Alpha record covers the identity, API IDs, pricing and weights.
What changed with the weight release
At GLM‑5.3’s launch, Z.ai made the model available through hosted interfaces but held back the full weights for roughly two weeks. The company said the interval would be used for safety evaluation and hardening after cyber capability improved faster than expected during post-training. By August 28, both official full-model repositories were publicly downloadable:
zai-org/GLM-5.3, the native FP8 release.zai-org/GLM-5.3-BF16, the much larger BF16 release.
Developers can now inspect the configuration, audit the license, measure the repositories, and use first-party serving recipes. Independent reproduction is possible, but the weights do not make Z.ai’s existing scores independent. That requires compatible hardware, matching harnesses and settings, and public run artifacts.
Z.ai had not published a detailed hardening report in the materials reviewed. The release followed the announced interval; outside validation of every safeguard did not.
GLM‑5.3 versus GLM‑5.3‑Flash—and where Ox Alpha fits
These related models have different architectures, modalities, licenses, and deployment profiles.
| Field | GLM‑5.3 | GLM‑5.3‑Flash |
|---|---|---|
| Official repository | zai-org/GLM-5.3 |
zai-org/GLM-5.3-Flash |
| Rounded parameters | 744B total / 40B active | 320B total / 18B active |
| Preview identity | None disclosed | Ox Alpha |
| Modalities | Text input and output | Native text, image, and video input |
| Architecture | GLM‑5.2-derived base; post-training upgrade | New base; hybrid linear and sparse attention; Manifold-Constrained Hyper-Connections |
| Context window | 1,048,576 tokens | 1,048,576 tokens |
| FP8 repository | 755.7 GB; 141 weight shards | 328.4 GB; 62 weight shards |
| BF16 repository | About 1.5 TB; 282 weight shards | 642.7 GB; 120 weight shards |
| Weight license | Custom GLM‑5.3 License | MIT |
| Best fit | Maximum coding, cyber, and long-horizon capability | Lower-cost multimodal agents and relatively more attainable self-hosting |
“More attainable” is relative: 328.4 GB remains server-scale. Our GLM‑5.3‑Flash review and pricing guide covers that model; its tests do not transfer to full GLM‑5.3. We also examined Flash inference on Chinese accelerators.
GLM‑5.3 specifications at a glance
| Specification | Released GLM‑5.3 |
|---|---|
| Developer | Z.ai |
| Exact repository | zai-org/GLM-5.3 |
| Model class | Mixture-of-experts causal language model |
| Modalities | Text input and text output |
| Family label | 744B total / 40B active |
| Repository parameter count | 753,329,940,480 parameters in Hugging Face safetensors metadata |
| Why the numbers differ | Z.ai and serving sources publish approximately 743B/39B or 744B/40B, while artifact metadata counts about 753B. These are rounded labels and/or different parameter-counting boundaries; Z.ai has not published a formal reconciliation. |
| Base model | Same base as GLM‑5.2 |
| Where improvements came from | Post-training, not a new pretraining run |
| Main transformer layers | 78, plus one next-token-prediction layer in the configuration |
| Hidden size | 6,144 |
| Attention | 64 attention heads and 64 key/value heads |
| Experts | 256 routed experts, one shared expert, eight routed experts selected per token |
| Native context | 1,048,576 tokens |
| Maximum documented hosted-API output | 128K tokens |
| Reasoning | Always enabled; low, high, or max; default max |
| Agent features | Streaming, function calling, structured output, context caching, and tool use |
| Default precision | Native FP8, E4M3 configuration |
| Default repository size | 755.7 GB in Hugging Face CLI dry-run output |
| BF16 repository | zai-org/GLM-5.3-BF16; about 1.5 TB |
| Supported local stacks | vLLM, SGLang, Transformers, KTransformers, Unsloth, TokenSpeed, plus documented Ascend stacks including vLLM-Ascend, xLLM, and SGLang |
| Hosted API price at cutoff | $1.40 input, $0.26 cached input, $4.40 output per million tokens |
| License | Custom GLM‑5.3 License—not MIT |
The 1M-token capability is not a promise that every local deployment serves it efficiently. KV-cache demand grows with length and concurrency; the full-context recipe uses eight 180 GB B200s and an FP8 KV cache.
What the weights contain—and why 40B active is not the download size
For each token, GLM‑5.3’s router activates a subset of its experts—roughly 40B parameters under Z.ai’s rounded label. That reduces computation, not storage: later tokens may select other experts, so the deployment still needs the entire checkpoint.
For this update, Kingy checked all four family repositories with Hugging Face’s hf download --dry-run; no weights were downloaded:
| Repository | Precision | Total files | Weight shards | Dry-run transfer size |
|---|---|---|---|---|
zai-org/GLM-5.3 |
Native FP8, with small non-FP8 tensors | 153 | 141 | 755.7 GB |
zai-org/GLM-5.3-BF16 |
BF16 | 291 | 282 | About 1.5 TB |
zai-org/GLM-5.3-Flash |
Native FP8 | 72 | 62 | 328.4 GB |
zai-org/GLM-5.3-Flash-BF16 |
BF16 | 130 | 120 | 642.7 GB |
These are decimal transfer sizes, not GiB memory figures. Leave room for download state, metadata, runtimes, container layers, logs, and caches.
GLM‑5.3 benchmarks: a large upgrade, not a clean sweep
Every same-family comparison in Z.ai’s table improves. Terminal-Bench 3.0 rises from 4.6 to 28.3; DeepSWE from 46.2 to 66.9; SWE-Marathon from 19.4 to 42.5; and AutomationBench from 26.2 to 48.2.
The table reproduces every category in the official model card. Figures are Z.ai-reported unless noted. The competitor column shows the highest non-GLM‑5.3 value in Z.ai’s table.
| Area | Benchmark | GLM‑5.3 | GLM‑5.2 | Best competing result shown |
|---|---|---|---|---|
| Terminal agents | Terminal-Bench 2.1 | 88.2 | 81.0 | GPT‑5.6 Sol: 88.8 |
| Terminal agents | Terminal-Bench 3.0 | 28.3 | 4.6 | GPT‑5.6 Sol: 34.6 |
| Software engineering | DeepSWE v1.1 | 66.9 | 46.2 | GPT‑5.6 Sol: 72.7 |
| Repository generation | NL2Repo | 58.0 | 48.9 | Opus 4.8: 69.7 |
| Program synthesis | ProgramBench Almost Solved | 19.0 | 9.5 | Fable 5 with fallback: 33.0 |
| Software engineering | FrontierSWE | 78.1 | 67.5 | Fable 5 with fallback: 88.2 |
| Long-horizon coding | SWE-Marathon v1.1 | 42.5 | 19.4 | Opus 4.8: 48.8 |
| Post-training research | PostTrainBench | 39.8 | 31.7 | Fable 5 with fallback: 41.8 |
| Vulnerability discovery | CyberGym | 84.5 | 77.2 | Fable 5 with fallback: 83.8 |
| Exploitation tasks | ExploitGym, 2h / 6h | 105 / 130 | 29 / 39 | GPT‑5.6 Sol: 216 / 293 |
| Exploitation coverage | ExploitBench | 54.4 | 24.4 | Fable 5 with fallback: 78.0 |
| Tool use | Toolathlon Verified | 73.0 | 59.9 | Kimi K3: 76.5 |
| Workflow automation | AutomationBench v1.0.6 | 48.2 | 26.2 | Kimi K3: 46.7 |
| Agent evaluation | Agents’ Last Exam CLI | 28.5 | 23.8 | GPT‑5.6 Sol: 28.6 |
| Knowledge and tools | HLE with Tools | 62.5 | 54.7 | GPT‑5.6 Sol: 64.5 |
| Professional tasks | GDPval-AA v2 Elo | 1769 | 1508 | Fable 5 with fallback: 1743 |
GLM‑5.3 does not win every row. GPT‑5.6 Sol leads five listed evaluations, Fable 5 leads four, and Kimi K3 leads Toolathlon. GLM‑5.3 leads CyberGym, AutomationBench, and GDPval-AA. A 0.1-point ALE gap is not decision-grade without uncertainty estimates.
Evidence quality and comparability
- Externally run: Z.ai attributes FrontierSWE to Proximal and GDPval-AA to Artificial Analysis. Full auditability still needs evaluator artifacts and versioning.
- Public suite, vendor-run: Most named-suite GLM‑5.3 scores were run or assembled by Z.ai. Public tasks do not make a vendor run independent.
- Private or incomplete evidence: Treat in-house tests and evaluations without public prompts, outputs, containers, judge traces, or logs as private evidence. We exclude Z.ai Code Bench’s in-house “50%” claim.
- Harnesses differ: Reported context spans 300K–1M, output caps 64K–163,840, and timeouts four hours to unlimited. Some scores average three runs; others are single-run Pass@1.
- Scores differ: ExploitGym counts tasks, GDPval-AA is Elo, and other rows use suite-specific measures. Never average them into one index.
- Missing means not reported: A dash in Z.ai’s full table is absent evidence, not a zero.
Our GLM‑5.3 versus Kimi K3 versus DeepSeek V4 Pro comparison adds competitive context. The official table remains mostly vendor-reported.
Cyber capability and the two-week delay
Z.ai tied the delay to unexpectedly fast cyber gains during post-training.
GLM‑5.3’s 84.5 is the highest CyberGym result in Z.ai’s displayed comparison, narrowly above Fable 5 at 83.8 and GPT‑5.6 Sol at 83.6. On deeper exploitation measurements, the order changes materially: GLM‑5.3 records 54.4 on ExploitBench versus 78.0 for Fable 5 and 76.5 for GPT‑5.6 Sol. On ExploitGym, GLM‑5.3 completes 105/130 tasks under the two- and six-hour budgets; GPT‑5.6 Sol records 216/293 and Fable 5 records 181/247.
These tests inform defensive research and release governance; they do not establish attack success on unknown targets. ExploitGym normalizes time using token-rate assumptions, while CyberGym is single-run Pass@1 over 1,507 tasks with unlimited per-task timeout. Such choices can move rankings.
Open weights aid audit and defensive work while lowering capability-access barriers. Security deployments need authorization boundaries, isolation, logging, and abuse monitoring. This article provides no exploitation instructions.
The GLM‑5.3 license: broad rights, one important threshold
Full GLM‑5.3 is open-weight under a custom GLM‑5.3 License. It is not MIT-licensed, and “open source” would conceal a material condition.
The license permits use, modification, distribution, sublicensing, sale, deployment, fine-tuning, and derivatives. Copies or substantial portions must retain the notice, and use must comply with law.
For Model-as-a-Service operators giving third parties meaningful control over inference or fine-tuning, aggregate licensee-and-affiliate revenue above $10 billion over any consecutive 12 months triggers Z.ai security review before commercial use. Some feature-bound end-user products and simple request relays are excluded from the MAAS definition.
GLM‑5.3‑Flash instead uses MIT. This is practical editorial analysis, not legal advice.
How to download GLM‑5.3 without an expensive mistake
Start with metadata and a dry run. Hugging Face’s CLI will show the files and transfer estimate without pulling three-quarters of a terabyte.
python -m pip install -U huggingface_hub
hf download zai-org/GLM-5.3 --dry-run
To inspect the architecture and tokenizer without downloading any weight shards:
hf download zai-org/GLM-5.3 \
config.json generation_config.json tokenizer.json tokenizer_config.json \
--local-dir ./GLM-5.3-metadata
When storage, bandwidth, and operational headroom are confirmed, download the native FP8 repository:
hf download zai-org/GLM-5.3 \
--local-dir ./GLM-5.3
The repository was public and ungated in Kingy’s check. Authentication can improve rate limits; pin a revision for production provenance. Follow Hugging Face’s download documentation.
Do not allocate exactly 756 GB. Budget for temporary state, caches, environments, containers, logs, and rollback space.
Hardware reality
The official checkpoint is datacenter-class even though only a fraction of its experts activate for each token.
| Deployment target | Practical starting point | What it means |
|---|---|---|
| Native FP8 | One node with 8×H200 or 8×H20, 141 GB each | The current vLLM recipe’s single-node starting point |
| Full 1M context | 8×B200, 180 GB each, with FP8 KV cache | More memory for long context; concurrency still affects capacity |
| AMD | 8×MI300X or MI355X-class accelerators with documented ROCm/AITER configuration | Official recipe is a starting point; validate kernels and limits on the exact platform |
| BF16 | Multi-node | The roughly 1.5 TB checkpoint does not fit the cited single-node configurations |
| Gaming PC or ordinary Mac | Not realistic for the official full checkpoint | Storage is only the first constraint; accelerator memory, bandwidth, and runtime support dominate |
| Community quantization | Hardware-specific, third-party | Audit the publisher, calibration, format, quality loss, and runtime compatibility separately |
One documented community option is Inferact/GLM-5.3-NVFP4. Kingy’s dry run showed about 464.9 GB. The vLLM recipe describes it as Blackwell-specific and quantizes the MoE expert linear layers while leaving other components at higher precision. It is an Inferact artifact, not an official Z.ai checkpoint.
Serving GLM‑5.3 with vLLM
The current official vLLM recipe requires vLLM 0.28.0 or newer and Transformers 5.15.0 or newer. On a supported eight-GPU FP8 node, the baseline recipe is:
uv venv
source .venv/bin/activate
uv pip install "vllm==0.28.0" --torch-backend=auto
uv pip install "transformers>=5.15.0"
vllm serve zai-org/GLM-5.3 \
--kv-cache-dtype fp8 \
--tensor-parallel-size 8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-auto-tool-choice \
--served-model-name glm-5.3
Validate drivers, CUDA or ROCm, networking, memory, revision, concurrency, and target context on the server. Full 1M-token operation needs the recipe’s B200/FP8-KV configuration.
SGLang’s cookbook documents NVIDIA and AMD alternatives. Transformers, KTransformers, Unsloth, TokenSpeed, and Ascend stacks appear in Z.ai’s family repository; support does not prove performance on every configuration.
Calling the local OpenAI-compatible endpoint
After vLLM starts, its default OpenAI-compatible endpoint is http://localhost:8000/v1. Install the OpenAI Python SDK and call the served model name:
from openai import OpenAI
client = OpenAI(
api_key="local-only",
base_url="http://localhost:8000/v1",
)
response = client.chat.completions.create(
model="glm-5.3",
messages=[
{"role": "user", "content": "Review this migration plan and identify the three highest-risk assumptions."}
],
max_tokens=1_024,
extra_body={
"chat_template_kwargs": {
"reasoning_effort": "high",
"clear_thinking": True,
}
},
)
print(response.choices[0].message.content)
reasoning_effort accepts low, high, or max; benchmarks and the default use max. For chat, test the documented clear_thinking=true behavior against the pinned tokenizer.
Calling Z.ai’s hosted API
Teams without an eight-GPU server can use Z.ai’s OpenAI-compatible API. Store the key in an environment variable; never paste a production key into source code.
export ZAI_API_KEY="replace-with-your-own-key"
python -m pip install -U openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ZAI_API_KEY"],
base_url="https://api.z.ai/api/paas/v4/",
)
response = client.chat.completions.create(
model="glm-5.3",
messages=[
{"role": "user", "content": "Design a rollback-safe PostgreSQL schema migration."}
],
max_tokens=4_096,
temperature=1.0,
extra_body={
"thinking": {"type": "enabled"},
"reasoning_effort": "high",
},
)
print(response.choices[0].message.content)
Z.ai’s developer guide documents reasoning, tools, structured output, streaming, and caching. The general base URL differs from the /api/coding/paas/v4/ Coding Plan endpoint.
At cutoff, Z.ai listed $1.40 input, $0.26 cached input, and $4.40 output per million tokens. Recheck before purchase. Retries, reasoning, tools, and correction determine cost per successful task.
API, self-host, or Flash?
| Choice | Choose it when | Avoid it when |
|---|---|---|
| Z.ai hosted GLM‑5.3 API | You are evaluating quality, have low-to-medium volume, need fast access, or do not want to operate distributed inference | Policy requires local weights or data residency; usage is sustained enough that owned infrastructure may be justified |
| Self-host full GLM‑5.3 | You need weight control, private adaptation, local data handling, reproducible versions, or consistently high utilization—and already operate suitable infrastructure | You lack an eight-accelerator node, distributed-inference expertise, or a clear utilization case |
| GLM‑5.3‑Flash | You need native image/video input, a smaller model, lower API pricing, or an MIT license | You specifically need the full model’s strongest coding, cyber, or long-horizon profile |
| Neither | Your tasks fit a smaller model, require consumer-local deployment, or do not benefit from 1M context and heavy reasoning | The workload has been measured and clearly needs this capability tier |
For a broader hardware and capability map, see Kingy’s guide to the best open-weight AI models, benchmarks, and hardware requirements.
Final verdict
GLM‑5.3 is now inspectable and deployable, but the operational answer remains: start with the API, measure task success, and self-host only when control or sustained economics earns the eight-GPU complexity. Choose Flash for multimodality, a smaller footprint, or MIT licensing.
Most importantly, keep the names straight: GLM‑5.3 is the full text model. GLM‑5.3‑Flash is the multimodal model previously tested as Ox Alpha.
Methodology and source cutoff
Source review ended August 28, 2026, Pacific time. Kingy inspected model cards, configs, licenses, and repository metadata, then ran unauthenticated dry runs against four official family repositories and Inferact NVFP4. No weights were downloaded.
We reviewed Z.ai’s docs, pricing, and current vLLM/SGLang recipes. Code was syntax-checked, not integration-tested.
Kingy verified the released artifacts, configuration, license, and deployment documentation but did not independently run the 744B checkpoint. We did not claim independent benchmark reproduction, throughput, latency, power use, or deployment cost. Repository sizes, software versions, pricing, and documentation can change after the cutoff.
FAQ
Are the GLM‑5.3 weights available to download?
Yes: FP8 is zai-org/GLM-5.3; BF16 is zai-org/GLM-5.3-BF16. Both were public in Kingy’s check.
Was GLM‑5.3 formerly called Ox Alpha?
No. Ox Alpha was the anonymous preview identity of GLM‑5.3‑Flash, the 320B/18B-active multimodal model. It was not the full 744B-class GLM‑5.3 text model.
How large is the GLM‑5.3 download?
The CLI reported 755.7 GB for FP8 and about 1.5 TB for BF16. Leave operational headroom.
Can GLM‑5.3 run on a Mac or gaming PC?
Not realistically for the official checkpoint. Z.ai’s vLLM path starts with eight high-memory datacenter accelerators; third-party quantizations carry separate risks.
Is GLM‑5.3 open source or MIT-licensed?
Full GLM‑5.3 is open-weight under a custom license; Flash uses MIT. Large MAAS operators should review the $10 billion security-review provision.
Does GLM‑5.3 support a one-million-token context window locally?
The config supports 1,048,576 positions, but serving it depends on KV cache, memory, concurrency, and settings. The full-context recipe uses eight B200s.
Does releasing the weights independently validate Z.ai’s benchmarks?
No. It lets others attempt reproduction. Most scores in the official table remain vendor-reported until independent evaluators publish compatible run artifacts and results.
Should I use the API or self-host GLM‑5.3?
Use the API for evaluation or low-to-medium volume. Self-host when control, customization, or sustained utilization justifies the infrastructure.
