Start Building →
Guide

GLM-5.3-Flash Explained: The Stealth-Launched Multimodal Model, Pricing, and Self-Hosting

It ran anonymously as Ox Alpha, reportedly drew heavy usage on OpenRouter, then turned out to be a different model from the flagship. Here is what it actually is.

Author photo
packet.ai Team
October 5, 2026

GLM-5.3-Flash spent its first week anonymous. Under the name Ox Alpha it ran on OpenRouter and OpenCode, reportedly became one of the most-used models of that week, and was only revealed as a Z.ai release on August 26, 2026. It is not a trimmed copy of the GLM-5.3 flagship. It is a separate 320B-A18B mixture-of-experts model with a newly trained base, native image and video input, a 1M-token context window, and MIT-licensed weights. This guide covers what GLM 5.3 Flash actually is, how it differs from the flagship, what it costs, where to get it, and what self-hosting it takes.

Key takeaways

  • GLM-5.3-Flash launched August 26, 2026 as roughly a 320B-total, 18B-active mixture-of-experts model. It is the first natively multimodal model in the GLM-5 series, accepts text, image, and video input, and was trained from a new base rather than post-trained from GLM-5.2
  • Do not confuse it with the older "Flash" in Z.ai's lineup: GLM-4.7-Flash, released in January 2026, is a 30B-total, 3B-active model with a 200K context window. GLM-5.3-Flash is a much larger, newer model
  • Z.ai's official list price is $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens. A 50% launch promotion ended September 9, 2026, and third-party gateways set their own rates
  • Compared with the GLM-5.3 flagship, Flash costs roughly one-ninth as much through the API and ships under a plain MIT license, while the flagship uses a custom license. Z.ai's published results are strong for the price, but the same tables show the flagship still leading on several of the hardest evaluations, and one hands-on test found Flash spending more time and tokens on a hard task
  • Despite 18B active parameters, self-hosting needs the whole checkpoint in memory: about 306 GiB of weights for the native FP8 version, roughly double that for BF16, and about 203 GB for one published NVFP4 quantization that runs in a four-GPU tensor-parallel setup

What GLM-5.3-Flash Actually Is

GLM-5.3-Flash is the first natively multimodal model in Z.ai's GLM-5 series. Z.ai describes it as 320B total parameters with 18B active per token; the Hugging Face checkpoint lists roughly 321B. According to vLLM's serving recipe, the 45-layer language model combines KDA linear-attention layers with NoPE sparse MLA layers, routes each token through 8 of 288 experts, and includes one multi-token-prediction draft layer. Z.ai reports about 3x less attention compute and a 4.4x smaller KV cache than the GLM-5.3 flagship, which is the design choice that makes a 1M-token context affordable to serve.

It was trained on a new base model over a 30-trillion-token multimodal corpus, rather than being post-trained from GLM-5.2's base the way the GLM-5.3 flagship was. That matters for how to read the name: "Flash" here signals a cost tier, not a distilled copy of a bigger sibling.

The name is also a source of confusion. Anyone searching for a glm flash model may land on GLM-4.7-Flash, which Z.ai released in January 2026 as a 30B-total, 3B-active model with a 200K context window and a free API tier. That is a different, much smaller model. Z.ai's cadence has also been fast enough that "the glm new model" can mean different things a month apart: GLM-5.2 in June, GLM-5.3 in mid-August, and GLM-5.3-Flash on August 26. As of this writing, GLM-5.3-Flash is the glm latest model in the Flash line, but check Z.ai's own model list before assuming that still holds.

The Ox Alpha Stealth Launch

The model appeared on OpenRouter and OpenCode on August 20, 2026 with no disclosed owner. OpenRouter's listing described it as developed and operated by a third-party provider that had chosen to remain anonymous during the preview, and OpenCode offered it free for the following week. Multiple reports say it became one of the most-used models on those platforms during that time. On August 26, Z.ai revealed that the ox alpha model was GLM-5.3-Flash, and reports describe the free access during that week as a deliberate load test.

For anyone who searched for the Ox Alpha name: it is the same model, not a separate release. The codename was only a testing label, and the model's behavior during the stealth week is the same one now available under the glm-5.3-flash name. One practical lesson from the reveal: the anonymous stealth/ox-alpha ID was retired afterward, so anything pinned to it had to move to the named model ID.

GLM 5.3 Flash vs GLM 5.3: What Actually Differs

GLM-5.3-FlashGLM-5.3 (flagship)
LaunchAugust 26, 2026Mid-August 2026 via API; weights followed August 28
Size~320B total, 18B active~753B total, ~40B active (per the Hugging Face checkpoint)
Base modelNewly trainedReused the GLM-5.2 base, with scaled post-training
LicenseMITGLM-5.3 License (Z.ai's own terms)
API list price (input / output per 1M)$0.15 / $0.50$1.40 / $4.40

The license difference is easy to miss and worth knowing. Flash's weights are plain MIT. The flagship's weights, released two days after Flash, use a custom GLM-5.3 License, which includes a clause requiring companies above $10 billion in revenue that host the model to pass a Z.ai security review. For most teams that clause changes nothing, but it is a real difference from GLM-5.2 and from Flash.

On capability, the honest picture is "strong for the price, not a flagship replacement." Z.ai's own launch numbers have Flash well ahead of GLM-5.2 on agentic coding benchmarks, and Z.ai says it approaches Claude Opus 4.8 on coding and agentic tests, which is a comparison against an earlier Opus generation rather than Anthropic's current one. Any glm 5.3 flash benchmark figure from Z.ai is a vendor claim until third parties reproduce it. The one independent number most sources cite is Artificial Analysis's Intelligence Index, reported at 57 for Flash against 60 for the flagship. Vision is the model's weaker side: published launch results show it trailing Google's Gemini 3.7 Flash on some vision benchmarks.

⚡ Cheap per token is not always cheap per solved task

In a hands-on comparison by The New Stack, Flash matched the flagship's score on the tasks tested, but on the hardest task it took 7.6 minutes and 38,677 tokens to reach an answer the flagship produced in under three minutes. One test is not a benchmark, but it illustrates the general point: a model that needs more tokens or more retries to finish a hard job can erase part of a per-token price advantage. Measure cost on your own task, not just the rate card.

Launch-week coverage also compared Flash with Alibaba's Qwen3.8-Flash-Next preview, so glm 5.3 flash vs qwen write-ups exist, but treat any head-to-head from that week as provisional.

GLM 5.3 Flash Pricing

Z.ai's current list pricing, set in its August 26, 2026 announcement, is $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens. This is the glm 5.3 flash pricing most relevant going forward. For launch, Z.ai ran a 50% promotion at $0.075 input, $0.25 output, and $0.015 cached, which ended at the close of September 9, 2026 (UTC+8). Any article, calculator, or reseller listing still showing the promo rates is out of date.

Third-party gateways set their own prices. A glm 5.3 flash openrouter listing, for example, shows per-provider rates that can sit above or below Z.ai's list price and shift over time, and gateway listings can also show different context limits or provider-specific configurations. Some listings included promo-era rates after Z.ai ended its own. For exact cost, check the live listing for the specific provider you will route through rather than relying on a figure from an article, including this one.

On whether the model is free: it was free during the Ox Alpha stealth week, but Z.ai's API is paid now. If you are asking whether glm 5.3 flash free access still exists, the practical answer is no for the first-party API. The weights are free to download under the MIT license, but running them takes hardware that is not free.

Where to Get GLM-5.3-Flash

The glm 5.3 flash api is available directly from Z.ai under the model name glm-5.3-flash, and through gateways and hosts including OpenRouter, Cloudflare Workers AI, Vercel AI Gateway, Baseten, and Deep Infra, according to several launch-period reports. It is also included in Z.ai's GLM Coding Plan tiers. Availability and pricing differ by host, so confirm both before choosing one.

For self-hosting, the weights are public on Hugging Face as zai-org/GLM-5.3-Flash under the MIT license, with BF16, native FP8, and community-quantized variants available. Strictly speaking, this is open-weight rather than fully open-source: the weights and license are open, but that says nothing about the training data or training code. Searches for glm 5.3 flash huggingface or glm 5.3 flash open source are generally looking for the zai-org repository.

One practical consideration for teams with data-residency requirements: Z.ai is a Beijing-based company, and its first-party API is operated by Z.ai. If where requests are processed matters for your compliance posture, check that with Z.ai directly, or consider a third-party host in a region you control, or running the open weights yourself.

Context Window and Multimodal Inputs

The model configuration specifies a glm 5.3 flash context window of 1,048,576 tokens, matching the flagship's 1M window, with no context-length price tiering on Z.ai's list rates. Gateway listings can show different limits, so check the specific provider. The model is glm 5.3 flash multimodal in the strict sense: image and video are part of its training rather than added through a separate vision adapter, and Z.ai's documentation lists text, image, video, and file input with text output. Native vision is the headline change from GLM-5.2, which was text-only.

The hybrid attention design is what makes the long context practical: with a KV cache Z.ai reports as about 4.4x smaller than the flagship's, a million-token window costs less memory per request. That cost still grows with how much context you actually use, so a 1M-token window is a ceiling, not a default setting to leave on.

GLM 5.3 Flash GPU and Hardware Requirements

The 18B active-parameter figure describes compute per token, not memory. As with every large mixture-of-experts model, self-hosting means loading all the experts, so glm 5.3 flash gpu requirements are driven by the full checkpoint, not the active parameters. This is the same distinction covered in the packet.ai DeepSeek V4.1 Flash guide, and the sizing logic carries over.

CheckpointApproximate checkpoint (weight) sizeNotes
Native FP8 (default zai-org release)~306 GiBPer vLLM's serving recipe, before runtime and KV-cache overhead
BF16Roughly twice FP8Per the same recipe
NVFP4 (community and partner quantizations)~203 GB (one published quantization)Published recipes use four-GPU tensor parallelism; quantization can change outputs

These are checkpoint (weight) sizes from public model cards and serving recipes, not recommended total VRAM or guaranteed minimums. Serving also needs memory for the KV cache and runtime state, which depends on context length, batch size, KV-cache precision, and serving engine.

What has actually been validated is more useful than a checkpoint size. vLLM's recipe for the model (vLLM 0.29.0 or newer) lists H100, B200, GB200 NVL4, MI355X, and Ascend 950PR configurations. Its headline NVIDIA example serves the FP8 checkpoint with tensor parallelism of 4 on a single GB200 tray, with multi-token-prediction speculative decoding and an FP8 KV cache enabled. The other documented setups are below.

HardwareDocumented configurationNotes
GB200 NVL4 (one tray)FP8 checkpoint, TP4, MTP, FP8 KV cacheThe headline example in vLLM's recipe
4x AMD MI355XTP4 with MTPValidated per the recipe; MTP on MI355X needs a vLLM build containing a specific fix
4x H200TP4Used in the recipe's KV-cache offloading validation, a text-only test at a 16,384-token maximum length, not a full-context benchmark
NVFP4 variant (RedHatAI)4-bit expertsRequires NVIDIA Blackwell GPUs, per the recipe

The recipe also notes that Hopper GPUs do not support an FP8 KV cache for this model and must run a BF16 KV cache, which uses more memory at long context than the FP8 KV cache available on Blackwell.

The practical takeaway: native FP8 and BF16 deployments are firmly multi-GPU, and even a roughly 200 GB quantization is more weight than most single GPUs can hold, which is why published recipes use four-GPU setups. Those NVFP4 versions are third-party quantizations of Z.ai's weights, so validate output quality on your own workload before relying on one. For sizing the previous generation, the packet.ai GLM-5.2 self-hosting guide covers the same questions, and the self-hosting vs API break-even guide covers when running the weights yourself beats paying per token. At Flash's list price, the API is usually the cheaper route until volume is very high or data cannot leave your infrastructure.

Running GLM-5.3-Flash Through a Managed Endpoint

If you want the open weights without provisioning the cluster, a managed endpoint is the middle path between paying a first-party API and operating the GPUs yourself. packet.ai's Token Factory (see what a token factory is) is being built as an OpenAI-compatible endpoint for open-weight models. It is currently in private preview, with the model catalog and pricing still being finalized, so GLM-5.3-Flash's inclusion is not confirmed until that catalog is published. For the GLM-5.3 flagship, see the existing packet.ai GLM-5.3 API guide.

Join the waitlist for early access once it opens.

Sources and Further Reading

Frequently asked questions

GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model, released August 26, 2026. It is a 320B-total, 18B-active mixture-of-experts model with a 1M-token context window, text, image, and video input, and MIT-licensed open weights. It first ran anonymously under the codename Ox Alpha.
The weights are free to download under the MIT license, but hosted API access is paid. It was free during the Ox Alpha stealth week, and a 50% launch promotion ended September 9, 2026. Z.ai's current list pricing is $0.15 per million input tokens and $0.50 per million output tokens, and running the weights yourself requires multi-GPU hardware.
Flash is a smaller, separately trained model (about 320B total, 18B active) that is natively multimodal, MIT-licensed, and priced at roughly one-ninth of the flagship through the API. GLM-5.3 is much larger, reused the GLM-5.2 base, and uses a custom GLM-5.3 License. Z.ai's published tables show the flagship still leading on several of the hardest evaluations.
It is open-weight under a plain MIT license, with weights on Hugging Face as zai-org/GLM-5.3-Flash. That covers the weights and license, not necessarily the training data or training code. The GLM-5.3 flagship is not MIT; it uses a custom GLM-5.3 License.
There is no single universal number. The native FP8 checkpoint is about 306 GiB of weights, BF16 is roughly double, and one published NVFP4 quantization is about 203 GB. Those are weight sizes, not total VRAM: serving also needs KV-cache and runtime memory, and the published recipes use four-GPU setups. Do not size a deployment from the 18B active-parameter figure.

Last reviewed: October 5, 2026. Z.ai pricing, model availability, and third-party gateway rates have changed repeatedly since the August 26, 2026 launch, including the end of the launch promotion; verify current figures against Z.ai's own documentation and the specific host you plan to use before building on this model, and verify hardware support before provisioning.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog