AI News

Ox Alpha Was GLM-5.3-Flash: Tests, API, Pricing, Weights and Hardware

The answer is now definitive: Ox Alpha was the anonymous public preview of Z.ai’s GLM-5.3-Flash—not the full GLM-5.3 model. Z.ai says it tested Flash as ox-alpha on OpenCode and OpenRouter before the named launch. The official API, model documentation and MIT-licensed Hugging Face weights now replace the pre-reveal fingerprinting as the primary evidence.

That identity is only the beginning of the useful comparison. Full GLM-5.3 and GLM-5.3-Flash are different models with different bases, architectures, modalities, prices, licenses and deployment footprints. Full GLM-5.3 is the larger text-only coding model. Flash is the smaller, natively multimodal model built for much cheaper serving.

Kingy status — August 28, 2026
Identity: confirmed by Z.ai.
Named API: live as glm-5.3-flash.
Open weights: live as zai-org/GLM-5.3-Flash under MIT.
Kingy evidence: five narrow signed-out browser tests exist for Flash; one predeclared matched API packet has now been run once across full GLM-5.3 and Flash.
Remaining uncertainty: provider-specific serving limits, independent benchmark reproduction, exact preview-to-production checkpoint parity and undisclosed training hardware.

Commercial relationship for this article: None. This is based on Curtis Pyke’s direct confirmation on October 7, 2026.

The Ox Alpha reveal, in four dates

Date Event What it establishes
August 20, 2026 Ox Alpha appears as a free stealth preview through OpenCode and OpenRouter A real but unnamed multimodal reasoning endpoint existed
August 21–25 Community testers compare tokenization, error behavior and model capabilities Useful lineage clues, not identity proof
August 26 Z.ai launches GLM-5.3-Flash and says it tested the model anonymously as ox-alpha Direct first-party identity confirmation
August 26–28 Named API documentation, pricing and Hugging Face repositories go live Product specifications, access, weights and license become inspectable

Before the reveal, the strongest clues were GLM-family tokenization, the same always-on reasoning contract, a reported Z.ai-style error and an unusual mix of very long context plus visual input. Those observations were directionally useful. They could not prove which post-training checkpoint sat behind the endpoint.

They no longer have to. Z.ai’s own launch record names the preview. The responsible wording is therefore simple: Ox Alpha was GLM-5.3-Flash. It was not full GLM-5.3, and the retired stealth/ox-alpha route should not be used in new integrations.

Ox Alpha preview versus the named release

Field Ox Alpha preview Named GLM-5.3-Flash Current reading
Public identity Anonymous stealth model Z.ai GLM-5.3-Flash Confirmed by Z.ai
Access Free preview through OpenCode/OpenRouter Z.ai API, Coding Plan, OpenRouter and open weights Preview ended
Route stealth/ox-alpha Z.ai glm-5.3-flash; OpenRouter z-ai/glm-5.3-flash Use a named route
Context Listed near one million tokens Z.ai documents 1M; provider cards can expose different limits Check the chosen provider
Maximum output Preview/provider-specific Z.ai documents 128K Check the endpoint, not only the family name
Inputs Text, images and video were exposed Z.ai documents text, image, video and file input Native multimodality confirmed
Model scale Undisclosed 320B total / 18B active Vendor and repository record
Weights Unavailable FP8 and BF16 repositories on Hugging Face MIT-licensed
Price Zero-priced preview $0.15/M input and $0.50/M output list Launch promotion is temporary

Z.ai does not publicly state that every provider uses a numerically identical checkpoint, quantization and serving configuration. The identity is settled; exact endpoint equivalence is a narrower operational question.

GLM-5.3 versus GLM-5.3-Flash

The “Flash” suffix does not mean a cheaper serving tier for the same checkpoint. Z.ai says Flash starts from a newly trained base model. Full GLM-5.3 instead uses the same base as GLM-5.2 and attributes its improvements to scaled post-training.

Field GLM-5.3 GLM-5.3-Flash
Ox Alpha identity No Yes—the anonymous preview
Base/training relationship GLM-5.2 base with expanded post-training Newly trained base and 30T-token multimodal pre-training corpus
Rounded scale About 743B total / 39B active; often rounded 744B/40B 320B total / 18B active
Architecture GLM-5.2-derived MoE with dense/sparse attention design 45-layer multimodal MoE; hybrid linear and sparse attention, mHC and IndexPool
Modalities Text input and output Text, image, video and file input; text output
Z.ai context / maximum output 1M / 128K 1M / 128K
Reasoning Always on; low, high, max; default max Same
Z.ai model code glm-5.3 glm-5.3-flash
Z.ai list price per 1M tokens $1.40 input / $0.26 cached / $4.40 output $0.15 input / $0.03 cached / $0.50 output
Official FP8 repository zai-org/GLM-5.3 zai-org/GLM-5.3-Flash
Official BF16 repository zai-org/GLM-5.3-BF16 zai-org/GLM-5.3-Flash-BF16
Repository transfer size About 756 GB FP8 / 1.51 TB BF16 About 328 GB FP8 / 643 GB BF16
Weight license Custom GLM-5.3 License MIT
First-party/local stacks listed vLLM, SGLang, Transformers, KTransformers, Unsloth, TokenSpeed and documented Ascend options vLLM, SGLang, Transformers, KTransformers, Unsloth and TokenSpeed
Documented hardware starting point 8×H200/H20 for FP8; 8×B200 for full 1M context in the vLLM recipe FP8 TP4 example on a GB200 tray; AMD MI355X path; exact capacity depends on cache and concurrency
Best fit Maximum full-model coding, long-horizon and cyber profile Low-cost multimodal agents, document/UI work and an MIT-licensed deployment path

The repository size is not the production memory requirement. Operators also need headroom for the runtime, KV cache, multimodal encoder state, temporary buffers, parallelism, concurrency and failure recovery. “18B active” reduces per-token arithmetic; it does not let a machine discard the rest of the 320B checkpoint.

Benchmarks: shared names do not guarantee a matched experiment

Z.ai reports both models on several common suites. These figures are useful orientation, not Kingy-run head-to-head results. They come from vendor-published tables and inherit each benchmark’s harness, context, timeout, scoring and version choices.

Benchmark GLM-5.3 GLM-5.3-Flash Higher published result
Terminal-Bench 2.1 88.2 84.3 GLM-5.3
DeepSWE v1.1 66.9 63.4 GLM-5.3
NL2Repo 58.0 56.3 GLM-5.3
Toolathlon Verified 73.0 78.4 Flash
AutomationBench v1.0.6 48.2 48.8 Flash
Agents’ Last Exam CLI 28.5 26.3 GLM-5.3
HLE with Tools 62.5 55.3 GLM-5.3
GDPval-AA v2 1769 1773 Flash, by four reported Elo points

Do not average these rows. Terminal work, repository generation, tool use and a professional-task Elo rating do not share a denominator. Small gaps are not decision-grade without uncertainty and run-level artifacts.

The more actionable question is not “Which model has the larger table average?” It is: which model completes your task reliably at an acceptable latency and cost?

Kingy’s matched API result

Kingy previously ran five narrow tasks through the signed-out Z.ai chat interface with GLM-5.3-Flash. Flash passed structured output, JavaScript, simple prompt-injection resistance, exact cost arithmetic and retrieval in that small sample. Those runs were useful product observations, but they were not API benchmarks: there were no billed token counts, request IDs or server-side latency traces.

A matched API packet was predeclared for both named models and run once on August 29, 2026 UTC. It held the following constant:

  • identical text prompts and input records;
  • temperature: 1, top_p: 0.95, reasoning effort max and thinking enabled;
  • identical output caps, no tools and one fresh request per task;
  • deterministic graders declared before execution;
  • three runs per model;
  • latency, token use, cost, pass/fail and repair-loop recording.

The packet covers strict JSON, executable JavaScript, record retrieval, instruction-injection resistance, exact arithmetic and constrained evidence summarization. A seventh image test is Flash-only because Z.ai currently documents full GLM-5.3 as text-only. Reporting “unsupported” is more honest than converting the image into text and calling the result multimodal parity.

The run used Vercel AI Gateway model IDs zai/glm-5.3 and zai/glm-5.3-flash, with routing restricted on every request to provider zai. The runner issued 39 sequential requests with no retries. All 39 returned HTTP 200; the full model’s image case was recorded separately as unsupported.

Predeclared task GLM-5.3 GLM-5.3-Flash What the result supports
Strict JSON invoice 3/3 3/3 Both were exact in this task.
Executable JavaScript 0/3 0/3 All six calls exhausted the output cap in reasoning and returned no final code.
300-record retrieval 3/3 2/3 Full was exact; one Flash run ended with incomplete fenced JSON.
Prompt-injection triage 1/3 2/3 The three misses were empty or truncated at the cap.
Exact cost arithmetic 0/3 0/3 All six calls exhausted the 256-token cap in reasoning and returned no final JSON.
Evidence-bounded summary 0/3 1/3 Five calls exhausted the 384-token cap; Flash completed one exact answer.
Vision chart Unsupported 3/3 exact after audit Flash returned WEST, exactly as printed in the image; the predeclared expected value had the wrong capitalization.
Matched text total 7/18 (38.9%) 8/18 (44.4%) Too small and cap-sensitive for a broad quality ranking.

Median end-to-end latency across the 18 matched text calls was 4.296 seconds for full GLM-5.3 and 6.789 seconds for Flash. This includes Vercel Gateway overhead; it is not a direct-endpoint latency claim. Vercel reported 40,518 input tokens, 11,446 completion tokens, 22,528 cached prompt tokens and $0.03972980 total cost. A conservative reconstruction at uncached Z.ai list rates was $0.05638400, safely below the authorized $0.95 cap.

The most important limitation is the output contract. With reasoning set to max, 19 of the 39 calls consumed their full output allowance in reasoning and returned an empty final answer. The zeroes on JavaScript and arithmetic therefore show that these short caps are unsafe with max reasoning; they do not prove that either model lacks the underlying capability. No repair or higher-cap rerun was performed because the authorization allowed one execution without retries.

The immutable run file records the three Flash vision answers as failed because the predeclared fixture expected West. Final audit of the deterministic image showed that its label is WEST, and the prompt required preserving capitalization. Kingy corrected the published grade to 3/3 without changing or rerunning the raw evidence.

Kingy preserved the raw, run-level evidence in the publication record with SHA-256 fbdcf08b143bb1051e277d18be2812d0c472a2833040297ab7ff5d2601ee3e8f.

API model IDs and request examples

Use the exact model code for the endpoint you call:

Provider/path Full model Flash
Z.ai Model API glm-5.3 glm-5.3-flash
Vercel AI Gateway zai/glm-5.3 zai/glm-5.3-flash
OpenRouter z-ai/glm-5.3 z-ai/glm-5.3-flash
Retired stealth preview Not applicable stealth/ox-alpha — historical only
Hugging Face FP8 zai-org/GLM-5.3 zai-org/GLM-5.3-Flash

Z.ai exposes multiple compatible protocols and base URLs. The general model API and the Coding Plan endpoint are not interchangeable account products. Confirm which credential you have before copying a base URL.

An OpenAI-compatible text request through Z.ai can be structured as follows:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ZAI_API_KEY"],
    base_url="https://api.z.ai/api/paas/v4/",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",  # change to glm-5.3 for the full model
    messages=[
        {
            "role": "user",
            "content": "Return the three highest-risk assumptions in this rollout plan.",
        }
    ],
    thinking={"type": "enabled"},
    reasoning_effort="max",
    temperature=1.0,
    top_p=0.95,
    max_tokens=2048,
)

print(response.choices[0].message.content)

Never place a production API key in source code. Pin the provider, model ID and settings in evaluation logs; a family label alone is not enough for reproducibility.

Pricing: Flash is roughly one-tenth the list price

Z.ai’s pricing page listed these US-dollar rates on August 28, 2026:

Model Input / 1M Cached input / 1M Output / 1M
GLM-5.3 $1.40 $0.26 $4.40
GLM-5.3-Flash list $0.15 $0.03 $0.50
GLM-5.3-Flash launch promotion $0.075 $0.015 $0.25

Z.ai says the Flash promotion ends at 24:00 on September 9, 2026, UTC+8. Recheck the live pricing page before publication or purchase. OpenRouter prices and context limits are provider records, not automatic substitutes for Z.ai’s own API terms.

The token rate still does not determine workload cost. Reasoning tokens, long outputs, retries, cache hits, tool calls and repair loops can dominate. Compare cost per successful task, not the cheapest million-token headline.

Hugging Face weights and licenses

The official repositories are live:

Flash uses the MIT License. Full GLM-5.3 uses Z.ai’s custom GLM-5.3 License. “Both are MIT” is false. “Weights are downloadable” is accurate; “fully open source” can overstate what is available because the complete training corpus and a reproducible end-to-end training recipe are not published.

Inspect metadata before downloading hundreds of gigabytes:

hf download zai-org/GLM-5.3-Flash --dry-run
hf download zai-org/GLM-5.3 --dry-run

For production, pin a repository revision, verify the license and hashes, and budget storage above the displayed transfer size.

Hardware reality

Neither official checkpoint is a normal laptop model.

The current vLLM recipe says full GLM-5.3 FP8 fits on one eight-accelerator H200 or H20 node, with eight B200s and FP8 KV cache used for the full one-million-token context example. The BF16 checkpoint is a multi-node deployment at roughly 1.51 TB before runtime overhead. Kingy’s full GLM-5.3 weights and deployment guide covers the larger model in more detail.

The Flash recipe is smaller but still server-scale. It lists roughly 306 GiB of default FP8 weight memory before runtime and KV-cache overhead, demonstrates tensor parallelism across four accelerators on a GB200 tray and documents an AMD MI355X path. A four-device launch example is a validated recipe, not a universal minimum for every context length, concurrency target or precision.

Deployment question Practical answer
Can the official Flash FP8 checkpoint fit on a gaming GPU? No; the weights alone are about 328 GB as a transfer
Can a large-memory Mac run the official model comfortably? Not as a normal supported production configuration
Does 18B active mean 18B of storage? No; all routed experts must remain available
Does 1M context always fit? No; KV cache, precision and concurrent sequences determine the serving envelope
Are community quantizations equivalent? No; validate publisher, calibration, runtime support and quality loss separately

Self-host when weight control, data location, customization or sustained utilization justifies distributed-inference operations. For evaluation and intermittent workloads, the hosted API is the rational first step.

What the Chinese-chip run proves—and what it does not

Z.ai says the Ox Alpha preview traffic was served on a large cluster of Chinese AI chips. It describes a custom SGLang-based inference stack, W8A8 quantization, mixed cache formats, Layer Split and disaggregated encode, prefill and decode pools. The company reports a threefold end-to-end serving improvement over its starting point on the same hardware; Kingy’s Chinese-chip inference analysis separates that serving evidence from unsupported training-hardware claims.

That is meaningful evidence of large-scale non-Nvidia inference. It does not establish:

  • the exact accelerator vendor or bill of materials;
  • absolute throughput, latency, utilization or power;
  • a controlled comparison with a named Nvidia system;
  • that GLM-5.3-Flash was trained entirely on Chinese chips.

The precise claim is “Ox Alpha inference was served on Chinese AI accelerators.” The broader training claim is unsupported by the disclosed evidence.

Which model should you choose?

Choose full GLM-5.3 when text-only coding, long-horizon engineering or the full model’s published capability profile matters more than price, multimodality or an MIT license.

Choose GLM-5.3-Flash when you need image, video or file input; low API prices; smaller—but still very large—weights; or MIT licensing.

Choose neither on reputation alone. Build 20–50 tasks from your own failures, declare the graders, run multiple attempts and count every repair loop. The winning endpoint is the one with the best reliable outcome at an acceptable latency and total cost.

Frequently asked questions

Was Ox Alpha GLM-5.3?

It was GLM-5.3-Flash, not the full GLM-5.3 model. Z.ai says it tested Flash anonymously as ox-alpha before launch.

Which API model ID should I use?

Use glm-5.3-flash on Z.ai or z-ai/glm-5.3-flash on OpenRouter. The stealth/ox-alpha preview route is historical.

Is GLM-5.3-Flash cheaper than GLM-5.3?

Yes. At the August 28 list rates, Flash was $0.15/M input and $0.50/M output versus $1.40/M and $4.40/M for full GLM-5.3. A temporary Flash promotion halved the Flash rates through September 9, 2026, UTC+8.

Are the GLM-5.3-Flash weights on Hugging Face?

Yes. Z.ai publishes FP8 and BF16 repositories under the MIT License. The displayed transfers are about 328 GB and 643 GB.

Can GLM-5.3-Flash run on a normal PC or Mac?

Not as an ordinary official deployment. The FP8 weights alone are about 328 GB, and a serving system also needs accelerator memory for runtime state and cache. Treat community quantizations as separate artifacts requiring their own validation.

Does GLM-5.3-Flash support images and video?

Z.ai documents text, image, video and file input with text output. Full GLM-5.3 is currently documented as text-only input.

Was GLM-5.3-Flash trained on Chinese chips?

Z.ai’s disclosed evidence is about inference: it says Ox Alpha traffic was served on Chinese AI chips. The launch materials reviewed here do not establish the hardware used for the full pre-training run.

Do the open weights independently validate Z.ai’s benchmark scores?

No. Open weights make independent reproduction possible. Validation still requires matching harnesses, versions, settings, hardware and public run artifacts.

Methodology and limitations

Source review ended August 28, 2026, America/Vancouver. Kingy prioritized Z.ai’s launch announcement and developer documentation, official Hugging Face repositories and licenses, current vLLM recipes, then provider records such as OpenRouter for provider-specific routes and limits.

Identity and repository facts are confirmed from first-party or official platform records. Architecture and performance claims remain attributed to Z.ai. The benchmark table is not a Kingy reproduction. Repository sizes are displayed transfer sizes, not guaranteed runtime memory. Pricing and provider limits can change after the cutoff.

Kingy’s earlier Flash review used five signed-out browser runs. It did not expose API token counts, request IDs, provider routing or server-side latency. The newer matched packet used Vercel AI Gateway with a hard zai provider allowlist, so its token accounting and request IDs are available but its end-to-end latency includes the Gateway. The packet ran once with three samples per supported task, no retries and no repair loop. Its exact-grade results are configuration-specific, especially because max reasoning consumed the small output caps on 19 calls. Do not generalize this packet into a broad model-quality ranking.

Primary sources