AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Z.ai released GLM-5.3-Flash, a 320B-parameter mixture-of-experts multimodal model with a one-million-token context window, under an MIT license with open weights. API pricing is roughly $0.15 per million input tokens and $0.50 output, but all 320 billion weights must be stored to self-host, making it cheap to serve yet expensive to run on your own hardware.

Z.ai has released GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with open weights under an MIT license, at API prices the company positions at roughly one-tenth the serving cost of GLM-5.2. But the pricing story splits sharply depending on how you use it: via the API, the model is genuinely inexpensive at approximately $0.15 per million input tokens and $0.50 per million output tokens, while self-hosting requires loading all 320 billion weights — hardware territory that rules out most individual workstations.

According to Z.ai’s release materials and reporting by ThorstenMeyerAI.com, GLM-5.3-Flash is a 320B-A18B mixture-of-experts model: 320 billion total parameters with only 18 billion active per token, down from 32 billion active in GLM-4.5. The weights were published on HuggingFace at launch, in contrast to Z.ai’s flagship GLM-5.3 text model two weeks earlier, whose weights were staged pending a cyber-safety review. The model carries a one-million-token context window and is the first natively multimodal model in the GLM-5 series, accepting text, images, and video.

The Flash tier’s listed API pricing runs around $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 for cached input, though users are advised to verify against Z.ai’s live pricing page since Flash tiers have historically been repriced. On the Artificial Analysis Intelligence Index, Z.ai reports a score of 57 at roughly $0.045 per task. Z.ai also says the model beats GLM-5.2 across its own benchmark suite while costing about a tenth as much to serve.

Z.ai claims the model was trained on a 30-trillion-token multimodal corpus and runs entirely on Chinese AI chips — a hardware-sovereignty claim that remains the company’s own assertion. The company also confirmed that the model briefly available free on OpenRouter as “Ox Alpha” was an early version, and says the official release is stronger and more stable.

At a glance
analysisWhen: released at launch; pricing figures cur…
The developmentZ.ai launched GLM-5.3-Flash with open weights and aggressive API pricing, prompting scrutiny of what ‘cheap’ actually means for the model’s different classes of users.

Why Agentic Workloads Drive the Price Pitch

The pricing matters because agent workflows have a specific cost shape: a single agent run can involve dozens of sequential steps — tool calls, repository reads, browser control, screenshot inspection, self-correction — each consuming tokens against a large working context. That workload does not reward paying frontier-model prices per step; it rewards a model that is strong enough, stable enough, and cheap enough that the agent can afford to take all of those steps.

Native multimodality compounds the value. As ThorstenMeyerAI.com notes, an agent that can read a screenshot and detect a broken layout closes a loop that previously required a human in the middle — a capability aimed directly at browser agents and coding agents that verify their own UI output.

The distinction between cheap-to-serve and cheap-to-self-host, however, determines who actually benefits. The MoE efficiency shows up on datacenter GPUs and flows to API customers as a low price. Self-hosters face the full memory footprint of a 320B model.

Amazon

high performance AI hardware for self-hosting

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From GLM-4.5 MoE to the Flash Variant

: “

GLM-5.3-Flash is built on a newly trained base redesigned for efficiency rather than a post-training pass on an older model, according to Z.ai. Its architecture pairs linear attention for local dependencies with sparse attention for global context, plus long-context optimizations to manage latency and memory at the million-token end.

The model continues Z.ai’s MoE lineage from GLM-4.5, but cuts active parameters per token from 32 billion to 18 billion while growing the total to 320 billion. The earlier Ox Alpha preview on OpenRouter gave outside analysts early access; independent reads from that preview placed it at roughly GLM-5.3’s level, vision aside — strong for the price, but not a leap past frontier models.

“GLM-5.3-Flash delivers roughly one-tenth the serving cost of GLM-5.2 while beating it across our benchmark suite.”

— Z.ai (release materials)

Amazon

large scale AI model server hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Vendor Benchmarks and Unverified Claims

Every headline comparison — including figures in the low-to-mid 80s on a terminal-coding benchmark said to approach Claude Opus 4.8, and a leading score on a knowledge-work benchmark — comes from Z.ai’s own testing, on company-chosen harnesses, context limits, and generation settings, against a company-selected comparison set. Different harnesses produce different numbers, so these should be treated as claims to verify rather than settled results.

Z.ai’s assertion that training and inference run entirely on Chinese AI chips is also unverified independently. The exact live pricing may shift, as Flash tiers have been repriced before. Independent evaluations of the official release, as opposed to the Ox Alpha preview, are still limited.

Amazon

multimodal AI model GPU setup

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Testing and Pricing Stability

Third-party benchmarking efforts should follow now that weights are publicly available on HuggingFace, giving the first independent read on whether the official release matches the Ox Alpha preview and Z.ai’s reported numbers. Z.ai’s live pricing page should be watched for adjustments to the Flash tier. Separately, the fate of the flagship GLM-5.3 text model’s weights — currently held for a cyber-safety review — remains an open question that will shape how Z.ai’s open-release strategy is perceived.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How much does GLM-5.3-Flash cost via the API?

Reported pricing is roughly $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 for cached input. Verify against Z.ai’s live pricing page, as Flash tiers have been repriced before.

Can I run GLM-5.3-Flash on my own hardware cheaply because only 18 billion parameters are active?

No. The 18 billion active parameters reduce per-token compute, but all 320 billion weights must be stored and loaded, requiring datacenter-grade VRAM. The efficiency advantage appears as a low API price, not as a home-workstation model.

Is GLM-5.3-Flash the same model as ‘Ox Alpha’ on OpenRouter?

Z.ai confirmed Ox Alpha was an early version of the model and says the official release is stronger and more stable.

Are the benchmark claims independently verified?

No. The headline results are Z.ai’s own measurements on its chosen harnesses and comparison sets. Independent analysts who tested the Ox Alpha preview rated it very good for the price but not a frontier-level leap.

What makes this model suited for agents specifically?

It combines native multimodality (text, image, video), a one-million-token context window, and low per-token pricing — allowing multi-step agent loops that call tools, inspect screenshots, and self-correct without incurring frontier-model costs.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Private AI Prompt Workspace For Sensitive Teams

IdeaNavigator AI launches a private, local-first prompt workspace designed for small regulated teams handling sensitive AI workflows, with pilot testing underway.

Show HN: Cactus Hybrid: We Taught Gemma 4 To Know When It’s Wrong

Cactus developers reveal Gemma 4, a small on-device AI model trained to identify when it makes mistakes, balancing privacy and accuracy.

The conversion. What turning the largest nonprofit into a company did to charity law.

OpenAI transformed from a nonprofit into a company retaining control, challenging traditional charity laws and raising questions about future conversions.

The Delegation Ladder: The Four Agentic Loops, and What Each One Lets You Stop Doing

An analysis of the four agentic loops in AI design, explaining what each allows you to stop doing and how they shape AI processes.