AI News

GLM-5.3 Weights Are Out: Specs, Benchmarks, Download Guide and Hardware Requirements

Z.ai has released the full GLM‑5.3 checkpoints. Here is what is actually downloadable, what the benchmark table does and does not prove, and why most readers should still use the API.

Originally published: August 13, 2026
Updated: August 28, 2026

Update note: Z.ai’s full GLM‑5.3 weights are live. This revision adds checkpoint sizes, shard counts, license terms, hardware guidance, and current download, vLLM, and API recipes. Ox Alpha became GLM‑5.3‑Flash, not full GLM‑5.3.

The verdict

Researchers and infrastructure teams can now inspect, adapt, and deploy Z.ai’s strongest coding model. It is not easy to run.

The official FP8 repository is a 755.7 GB transfer spread across 141 weight shards. The BF16 version is about 1.5 TB across 282 shards. Z.ai’s current vLLM deployment recipe starts with an eight-accelerator server; a normal gaming PC or Mac is not a realistic home for the official full checkpoint.

For most teams, Z.ai’s API is the rational first step. Self-host when control, data location, customization, or sustained utilization justifies datacenter infrastructure. Choose GLM‑5.3‑Flash when native vision, an MIT license, or a smaller checkpoint matters more than the full model’s coding profile.

One correction belongs above the fold: the full 744B-class GLM‑5.3 was never Ox Alpha. Z.ai says ox-alpha became the smaller, multimodal GLM‑5.3‑Flash. Kingy’s definitive GLM‑5.3‑Flash, formerly Ox Alpha record covers the identity, API IDs, pricing and weights.

What changed with the weight release

At GLM‑5.3’s launch, Z.ai made the model available through hosted interfaces but held back the full weights for roughly two weeks. The company said the interval would be used for safety evaluation and hardening after cyber capability improved faster than expected during post-training. By August 28, both official full-model repositories were publicly downloadable:

Developers can now inspect the configuration, audit the license, measure the repositories, and use first-party serving recipes. Independent reproduction is possible, but the weights do not make Z.ai’s existing scores independent. That requires compatible hardware, matching harnesses and settings, and public run artifacts.

Z.ai had not published a detailed hardening report in the materials reviewed. The release followed the announced interval; outside validation of every safeguard did not.

GLM‑5.3 versus GLM‑5.3‑Flash—and where Ox Alpha fits

These related models have different architectures, modalities, licenses, and deployment profiles.

Field GLM‑5.3 GLM‑5.3‑Flash
Official repository zai-org/GLM-5.3 zai-org/GLM-5.3-Flash
Rounded parameters 744B total / 40B active 320B total / 18B active
Preview identity None disclosed Ox Alpha
Modalities Text input and output Native text, image, and video input
Architecture GLM‑5.2-derived base; post-training upgrade New base; hybrid linear and sparse attention; Manifold-Constrained Hyper-Connections
Context window 1,048,576 tokens 1,048,576 tokens
FP8 repository 755.7 GB; 141 weight shards 328.4 GB; 62 weight shards
BF16 repository About 1.5 TB; 282 weight shards 642.7 GB; 120 weight shards
Weight license Custom GLM‑5.3 License MIT
Best fit Maximum coding, cyber, and long-horizon capability Lower-cost multimodal agents and relatively more attainable self-hosting

“More attainable” is relative: 328.4 GB remains server-scale. Our GLM‑5.3‑Flash review and pricing guide covers that model; its tests do not transfer to full GLM‑5.3. We also examined Flash inference on Chinese accelerators.

GLM‑5.3 specifications at a glance

Specification Released GLM‑5.3
Developer Z.ai
Exact repository zai-org/GLM-5.3
Model class Mixture-of-experts causal language model
Modalities Text input and text output
Family label 744B total / 40B active
Repository parameter count 753,329,940,480 parameters in Hugging Face safetensors metadata
Why the numbers differ Z.ai and serving sources publish approximately 743B/39B or 744B/40B, while artifact metadata counts about 753B. These are rounded labels and/or different parameter-counting boundaries; Z.ai has not published a formal reconciliation.
Base model Same base as GLM‑5.2
Where improvements came from Post-training, not a new pretraining run
Main transformer layers 78, plus one next-token-prediction layer in the configuration
Hidden size 6,144
Attention 64 attention heads and 64 key/value heads
Experts 256 routed experts, one shared expert, eight routed experts selected per token
Native context 1,048,576 tokens
Maximum documented hosted-API output 128K tokens
Reasoning Always enabled; low, high, or max; default max
Agent features Streaming, function calling, structured output, context caching, and tool use
Default precision Native FP8, E4M3 configuration
Default repository size 755.7 GB in Hugging Face CLI dry-run output
BF16 repository zai-org/GLM-5.3-BF16; about 1.5 TB
Supported local stacks vLLM, SGLang, Transformers, KTransformers, Unsloth, TokenSpeed, plus documented Ascend stacks including vLLM-Ascend, xLLM, and SGLang
Hosted API price at cutoff $1.40 input, $0.26 cached input, $4.40 output per million tokens
License Custom GLM‑5.3 License—not MIT

The 1M-token capability is not a promise that every local deployment serves it efficiently. KV-cache demand grows with length and concurrency; the full-context recipe uses eight 180 GB B200s and an FP8 KV cache.

What the weights contain—and why 40B active is not the download size

For each token, GLM‑5.3’s router activates a subset of its experts—roughly 40B parameters under Z.ai’s rounded label. That reduces computation, not storage: later tokens may select other experts, so the deployment still needs the entire checkpoint.

For this update, Kingy checked all four family repositories with Hugging Face’s hf download --dry-run; no weights were downloaded:

Repository Precision Total files Weight shards Dry-run transfer size
zai-org/GLM-5.3 Native FP8, with small non-FP8 tensors 153 141 755.7 GB
zai-org/GLM-5.3-BF16 BF16 291 282 About 1.5 TB
zai-org/GLM-5.3-Flash Native FP8 72 62 328.4 GB
zai-org/GLM-5.3-Flash-BF16 BF16 130 120 642.7 GB

These are decimal transfer sizes, not GiB memory figures. Leave room for download state, metadata, runtimes, container layers, logs, and caches.

GLM‑5.3 benchmarks: a large upgrade, not a clean sweep

Every same-family comparison in Z.ai’s table improves. Terminal-Bench 3.0 rises from 4.6 to 28.3; DeepSWE from 46.2 to 66.9; SWE-Marathon from 19.4 to 42.5; and AutomationBench from 26.2 to 48.2.

The table reproduces every category in the official model card. Figures are Z.ai-reported unless noted. The competitor column shows the highest non-GLM‑5.3 value in Z.ai’s table.

Area Benchmark GLM‑5.3 GLM‑5.2 Best competing result shown
Terminal agents Terminal-Bench 2.1 88.2 81.0 GPT‑5.6 Sol: 88.8
Terminal agents Terminal-Bench 3.0 28.3 4.6 GPT‑5.6 Sol: 34.6
Software engineering DeepSWE v1.1 66.9 46.2 GPT‑5.6 Sol: 72.7
Repository generation NL2Repo 58.0 48.9 Opus 4.8: 69.7
Program synthesis ProgramBench Almost Solved 19.0 9.5 Fable 5 with fallback: 33.0
Software engineering FrontierSWE 78.1 67.5 Fable 5 with fallback: 88.2
Long-horizon coding SWE-Marathon v1.1 42.5 19.4 Opus 4.8: 48.8
Post-training research PostTrainBench 39.8 31.7 Fable 5 with fallback: 41.8
Vulnerability discovery CyberGym 84.5 77.2 Fable 5 with fallback: 83.8
Exploitation tasks ExploitGym, 2h / 6h 105 / 130 29 / 39 GPT‑5.6 Sol: 216 / 293
Exploitation coverage ExploitBench 54.4 24.4 Fable 5 with fallback: 78.0
Tool use Toolathlon Verified 73.0 59.9 Kimi K3: 76.5
Workflow automation AutomationBench v1.0.6 48.2 26.2 Kimi K3: 46.7
Agent evaluation Agents’ Last Exam CLI 28.5 23.8 GPT‑5.6 Sol: 28.6
Knowledge and tools HLE with Tools 62.5 54.7 GPT‑5.6 Sol: 64.5
Professional tasks GDPval-AA v2 Elo 1769 1508 Fable 5 with fallback: 1743

GLM‑5.3 does not win every row. GPT‑5.6 Sol leads five listed evaluations, Fable 5 leads four, and Kimi K3 leads Toolathlon. GLM‑5.3 leads CyberGym, AutomationBench, and GDPval-AA. A 0.1-point ALE gap is not decision-grade without uncertainty estimates.

Evidence quality and comparability

  • Externally run: Z.ai attributes FrontierSWE to Proximal and GDPval-AA to Artificial Analysis. Full auditability still needs evaluator artifacts and versioning.
  • Public suite, vendor-run: Most named-suite GLM‑5.3 scores were run or assembled by Z.ai. Public tasks do not make a vendor run independent.
  • Private or incomplete evidence: Treat in-house tests and evaluations without public prompts, outputs, containers, judge traces, or logs as private evidence. We exclude Z.ai Code Bench’s in-house “50%” claim.
  • Harnesses differ: Reported context spans 300K–1M, output caps 64K–163,840, and timeouts four hours to unlimited. Some scores average three runs; others are single-run Pass@1.
  • Scores differ: ExploitGym counts tasks, GDPval-AA is Elo, and other rows use suite-specific measures. Never average them into one index.
  • Missing means not reported: A dash in Z.ai’s full table is absent evidence, not a zero.

Our GLM‑5.3 versus Kimi K3 versus DeepSeek V4 Pro comparison adds competitive context. The official table remains mostly vendor-reported.

Cyber capability and the two-week delay

Z.ai tied the delay to unexpectedly fast cyber gains during post-training.

GLM‑5.3’s 84.5 is the highest CyberGym result in Z.ai’s displayed comparison, narrowly above Fable 5 at 83.8 and GPT‑5.6 Sol at 83.6. On deeper exploitation measurements, the order changes materially: GLM‑5.3 records 54.4 on ExploitBench versus 78.0 for Fable 5 and 76.5 for GPT‑5.6 Sol. On ExploitGym, GLM‑5.3 completes 105/130 tasks under the two- and six-hour budgets; GPT‑5.6 Sol records 216/293 and Fable 5 records 181/247.

These tests inform defensive research and release governance; they do not establish attack success on unknown targets. ExploitGym normalizes time using token-rate assumptions, while CyberGym is single-run Pass@1 over 1,507 tasks with unlimited per-task timeout. Such choices can move rankings.

Open weights aid audit and defensive work while lowering capability-access barriers. Security deployments need authorization boundaries, isolation, logging, and abuse monitoring. This article provides no exploitation instructions.

The GLM‑5.3 license: broad rights, one important threshold

Full GLM‑5.3 is open-weight under a custom GLM‑5.3 License. It is not MIT-licensed, and “open source” would conceal a material condition.

The license permits use, modification, distribution, sublicensing, sale, deployment, fine-tuning, and derivatives. Copies or substantial portions must retain the notice, and use must comply with law.

For Model-as-a-Service operators giving third parties meaningful control over inference or fine-tuning, aggregate licensee-and-affiliate revenue above $10 billion over any consecutive 12 months triggers Z.ai security review before commercial use. Some feature-bound end-user products and simple request relays are excluded from the MAAS definition.

GLM‑5.3‑Flash instead uses MIT. This is practical editorial analysis, not legal advice.

How to download GLM‑5.3 without an expensive mistake

Start with metadata and a dry run. Hugging Face’s CLI will show the files and transfer estimate without pulling three-quarters of a terabyte.

python -m pip install -U huggingface_hub

hf download zai-org/GLM-5.3 --dry-run

To inspect the architecture and tokenizer without downloading any weight shards:

hf download zai-org/GLM-5.3 \
  config.json generation_config.json tokenizer.json tokenizer_config.json \
  --local-dir ./GLM-5.3-metadata

When storage, bandwidth, and operational headroom are confirmed, download the native FP8 repository:

hf download zai-org/GLM-5.3 \
  --local-dir ./GLM-5.3

The repository was public and ungated in Kingy’s check. Authentication can improve rate limits; pin a revision for production provenance. Follow Hugging Face’s download documentation.

Do not allocate exactly 756 GB. Budget for temporary state, caches, environments, containers, logs, and rollback space.

Hardware reality

The official checkpoint is datacenter-class even though only a fraction of its experts activate for each token.

Deployment target Practical starting point What it means
Native FP8 One node with 8×H200 or 8×H20, 141 GB each The current vLLM recipe’s single-node starting point
Full 1M context 8×B200, 180 GB each, with FP8 KV cache More memory for long context; concurrency still affects capacity
AMD 8×MI300X or MI355X-class accelerators with documented ROCm/AITER configuration Official recipe is a starting point; validate kernels and limits on the exact platform
BF16 Multi-node The roughly 1.5 TB checkpoint does not fit the cited single-node configurations
Gaming PC or ordinary Mac Not realistic for the official full checkpoint Storage is only the first constraint; accelerator memory, bandwidth, and runtime support dominate
Community quantization Hardware-specific, third-party Audit the publisher, calibration, format, quality loss, and runtime compatibility separately

One documented community option is Inferact/GLM-5.3-NVFP4. Kingy’s dry run showed about 464.9 GB. The vLLM recipe describes it as Blackwell-specific and quantizes the MoE expert linear layers while leaving other components at higher precision. It is an Inferact artifact, not an official Z.ai checkpoint.

Serving GLM‑5.3 with vLLM

The current official vLLM recipe requires vLLM 0.28.0 or newer and Transformers 5.15.0 or newer. On a supported eight-GPU FP8 node, the baseline recipe is:

uv venv
source .venv/bin/activate
uv pip install "vllm==0.28.0" --torch-backend=auto
uv pip install "transformers>=5.15.0"

vllm serve zai-org/GLM-5.3 \
  --kv-cache-dtype fp8 \
  --tensor-parallel-size 8 \
  --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 5 \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --enable-auto-tool-choice \
  --served-model-name glm-5.3

Validate drivers, CUDA or ROCm, networking, memory, revision, concurrency, and target context on the server. Full 1M-token operation needs the recipe’s B200/FP8-KV configuration.

SGLang’s cookbook documents NVIDIA and AMD alternatives. Transformers, KTransformers, Unsloth, TokenSpeed, and Ascend stacks appear in Z.ai’s family repository; support does not prove performance on every configuration.

Calling the local OpenAI-compatible endpoint

After vLLM starts, its default OpenAI-compatible endpoint is http://localhost:8000/v1. Install the OpenAI Python SDK and call the served model name:

from openai import OpenAI

client = OpenAI(
    api_key="local-only",
    base_url="http://localhost:8000/v1",
)

response = client.chat.completions.create(
    model="glm-5.3",
    messages=[
        {"role": "user", "content": "Review this migration plan and identify the three highest-risk assumptions."}
    ],
    max_tokens=1_024,
    extra_body={
        "chat_template_kwargs": {
            "reasoning_effort": "high",
            "clear_thinking": True,
        }
    },
)

print(response.choices[0].message.content)

reasoning_effort accepts low, high, or max; benchmarks and the default use max. For chat, test the documented clear_thinking=true behavior against the pinned tokenizer.

Calling Z.ai’s hosted API

Teams without an eight-GPU server can use Z.ai’s OpenAI-compatible API. Store the key in an environment variable; never paste a production key into source code.

export ZAI_API_KEY="replace-with-your-own-key"
python -m pip install -U openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ZAI_API_KEY"],
    base_url="https://api.z.ai/api/paas/v4/",
)

response = client.chat.completions.create(
    model="glm-5.3",
    messages=[
        {"role": "user", "content": "Design a rollback-safe PostgreSQL schema migration."}
    ],
    max_tokens=4_096,
    temperature=1.0,
    extra_body={
        "thinking": {"type": "enabled"},
        "reasoning_effort": "high",
    },
)

print(response.choices[0].message.content)

Z.ai’s developer guide documents reasoning, tools, structured output, streaming, and caching. The general base URL differs from the /api/coding/paas/v4/ Coding Plan endpoint.

At cutoff, Z.ai listed $1.40 input, $0.26 cached input, and $4.40 output per million tokens. Recheck before purchase. Retries, reasoning, tools, and correction determine cost per successful task.

API, self-host, or Flash?

Choice Choose it when Avoid it when
Z.ai hosted GLM‑5.3 API You are evaluating quality, have low-to-medium volume, need fast access, or do not want to operate distributed inference Policy requires local weights or data residency; usage is sustained enough that owned infrastructure may be justified
Self-host full GLM‑5.3 You need weight control, private adaptation, local data handling, reproducible versions, or consistently high utilization—and already operate suitable infrastructure You lack an eight-accelerator node, distributed-inference expertise, or a clear utilization case
GLM‑5.3‑Flash You need native image/video input, a smaller model, lower API pricing, or an MIT license You specifically need the full model’s strongest coding, cyber, or long-horizon profile
Neither Your tasks fit a smaller model, require consumer-local deployment, or do not benefit from 1M context and heavy reasoning The workload has been measured and clearly needs this capability tier

For a broader hardware and capability map, see Kingy’s guide to the best open-weight AI models, benchmarks, and hardware requirements.

Final verdict

GLM‑5.3 is now inspectable and deployable, but the operational answer remains: start with the API, measure task success, and self-host only when control or sustained economics earns the eight-GPU complexity. Choose Flash for multimodality, a smaller footprint, or MIT licensing.

Most importantly, keep the names straight: GLM‑5.3 is the full text model. GLM‑5.3‑Flash is the multimodal model previously tested as Ox Alpha.

Methodology and source cutoff

Source review ended August 28, 2026, Pacific time. Kingy inspected model cards, configs, licenses, and repository metadata, then ran unauthenticated dry runs against four official family repositories and Inferact NVFP4. No weights were downloaded.

We reviewed Z.ai’s docs, pricing, and current vLLM/SGLang recipes. Code was syntax-checked, not integration-tested.

Kingy verified the released artifacts, configuration, license, and deployment documentation but did not independently run the 744B checkpoint. We did not claim independent benchmark reproduction, throughput, latency, power use, or deployment cost. Repository sizes, software versions, pricing, and documentation can change after the cutoff.

FAQ

Are the GLM‑5.3 weights available to download?

Yes: FP8 is zai-org/GLM-5.3; BF16 is zai-org/GLM-5.3-BF16. Both were public in Kingy’s check.

Was GLM‑5.3 formerly called Ox Alpha?

No. Ox Alpha was the anonymous preview identity of GLM‑5.3‑Flash, the 320B/18B-active multimodal model. It was not the full 744B-class GLM‑5.3 text model.

How large is the GLM‑5.3 download?

The CLI reported 755.7 GB for FP8 and about 1.5 TB for BF16. Leave operational headroom.

Can GLM‑5.3 run on a Mac or gaming PC?

Not realistically for the official checkpoint. Z.ai’s vLLM path starts with eight high-memory datacenter accelerators; third-party quantizations carry separate risks.

Is GLM‑5.3 open source or MIT-licensed?

Full GLM‑5.3 is open-weight under a custom license; Flash uses MIT. Large MAAS operators should review the $10 billion security-review provision.

Does GLM‑5.3 support a one-million-token context window locally?

The config supports 1,048,576 positions, but serving it depends on KV cache, memory, concurrency, and settings. The full-context recipe uses eight B200s.

Does releasing the weights independently validate Z.ai’s benchmarks?

No. It lets others attempt reproduction. Most scores in the official table remain vendor-reported until independent evaluators publish compatible run artifacts and results.

Should I use the API or self-host GLM‑5.3?

Use the API for evaluation or low-to-medium volume. Self-host when control, customization, or sustained utilization justifies the infrastructure.