GLM-5.3-Flash opens with MIT weights, native multimodality, and Flash-tier pricing

GLM-5.3-Flash opens with MIT weights, native multimodality, and Flash-tier pricing

Z.ai released GLM-5.3-Flash with MIT open weights, native image input, 320B/18B MoE architecture, and API pricing far below the GLM-5.3 flagship tier.

Z.ai released GLM-5.3-Flash on August 26, 2026, calling it the first natively multimodal model in the GLM-5 series. The company published MIT-licensed weights on Hugging Face and opened hosted access under the API model ID glm-5.3-flash. Z.ai describes a newly trained 320B-parameter Mixture-of-Experts model with 18B active parameters, a 1M-token context window, and company-reported coding and agent scores that approach Claude Opus 4.8 while pricing well below GLM-5.3. 123

What launched

SignalConfirmed detailWhy it matters
ReleaseOfficial research blog dated 2026-08-26; Hugging Face repo zai-org/GLM-5.3-Flash last modified 2026-08-26T13:50:10Z with MIT license and safetensors weights. 12Hosted and self-host paths open on the same day.
Architecture320B total / 18B active parameters; hybrid sparse and linear attention plus Manifold-Constrained Hyper-Connections; max position embeddings of 1,048,576 tokens in the published config. 12A new base, separate from the earlier text-only GLM-5.3 post-training upgrade.
Multimodal APIImage inputs use type: image_url content blocks; recommended settings include temperature: 1, top_p: 0.95, and reasoning_effort: max, with thinking required. 3Screenshot and UI feedback can enter the same coding/agent loop as text.
PriceList API price is $0.15 input / $0.50 output per 1M tokens; a 50% promotion runs through 24:00 on September 9, 2026 (UTC+8) at $0.075 / $0.25. GLM-5.3 remains $1.4 / $4.4. Coding Plan users get 3× the usable quota of GLM-5.3. 14Cost is the clearest product differentiator versus the flagship tier.
Local servingOfficial cookbooks list SGLang, vLLM, TokenSpeed, and KTransformers. 2Self-hosting is documented, though a 320B MoE still needs substantial hardware.

What the company-reported scores show

Z.ai bar chart comparing GLM-5.3-Flash with GLM-5.2, DeepSeek-V4-Vision-Exp, Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash across coding and agent benchmarks
Z.ai's launch chart reports Terminal Bench 2.1 at 84.3, DeepSWE v1.1 at 63.4, AutomationBench v1.0.6 at 48.8, and Agents' Last Exam at 26.3 for GLM-5.3-Flash, with footnotes for harness and sampling settings. 1
On Z.ai's table, GLM-5.3-Flash beats GLM-5.2 by wide margins on DeepSWE (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2) and sits near Claude Opus 4.8 on several coding and agent rows. Vision rows such as OfficeQA Pro (62.4) and Chartography with tools (78.0) are part of the same company table. These figures are Z.ai-reported unless a third party reruns them under the same harness. 1
Before the named release, Z.ai says it ran the model anonymously as ox-alpha on OpenCode and OpenRouter and served that traffic on Chinese AI chips. The blog claims a 3× end-to-end serving gain on that domestic-chip stack versus its own baseline. 1

Why this matters

GLM-5.3-Flash is a different product from the August 14 GLM-5.3 text flagship: new base, native vision, open MIT weights on day one, and API list pricing about one-tenth of GLM-5.3's. For teams already on the GLM Coding Plan or evaluating open multimodal coding agents, the immediate path is a small hosted trial with screenshots and tool loops, plus a hardware check before downloading the full MoE checkpoint. Independent benchmark reruns and real workload latency will decide whether the company table holds outside Z.ai's harness.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content