
GLM-5.3-Flash opens with MIT weights, native multimodality, and Flash-tier pricing
Z.ai released GLM-5.3-Flash with MIT open weights, native image input, 320B/18B MoE architecture, and API pricing far below the GLM-5.3 flagship tier.
Z.ai released GLM-5.3-Flash on August 26, 2026, calling it the first natively multimodal model in the GLM-5 series. The company published MIT-licensed weights on Hugging Face and opened hosted access under the API model ID
glm-5.3-flash. Z.ai describes a newly trained 320B-parameter Mixture-of-Experts model with 18B active parameters, a 1M-token context window, and company-reported coding and agent scores that approach Claude Opus 4.8 while pricing well below GLM-5.3. 123What launched
| Signal | Confirmed detail | Why it matters |
|---|---|---|
| Release | Official research blog dated 2026-08-26; Hugging Face repo zai-org/GLM-5.3-Flash last modified 2026-08-26T13:50:10Z with MIT license and safetensors weights. 12 | Hosted and self-host paths open on the same day. |
| Architecture | 320B total / 18B active parameters; hybrid sparse and linear attention plus Manifold-Constrained Hyper-Connections; max position embeddings of 1,048,576 tokens in the published config. 12 | A new base, separate from the earlier text-only GLM-5.3 post-training upgrade. |
| Multimodal API | Image inputs use type: image_url content blocks; recommended settings include temperature: 1, top_p: 0.95, and reasoning_effort: max, with thinking required. 3 | Screenshot and UI feedback can enter the same coding/agent loop as text. |
| Price | List API price is $0.15 input / $0.50 output per 1M tokens; a 50% promotion runs through 24:00 on September 9, 2026 (UTC+8) at $0.075 / $0.25. GLM-5.3 remains $1.4 / $4.4. Coding Plan users get 3× the usable quota of GLM-5.3. 14 | Cost is the clearest product differentiator versus the flagship tier. |
| Local serving | Official cookbooks list SGLang, vLLM, TokenSpeed, and KTransformers. 2 | Self-hosting is documented, though a 320B MoE still needs substantial hardware. |
What the company-reported scores show

On Z.ai's table, GLM-5.3-Flash beats GLM-5.2 by wide margins on DeepSWE (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2) and sits near Claude Opus 4.8 on several coding and agent rows. Vision rows such as OfficeQA Pro (62.4) and Chartography with tools (78.0) are part of the same company table. These figures are Z.ai-reported unless a third party reruns them under the same harness. 1
Before the named release, Z.ai says it ran the model anonymously as
ox-alpha on OpenCode and OpenRouter and served that traffic on Chinese AI chips. The blog claims a 3× end-to-end serving gain on that domestic-chip stack versus its own baseline. 1Why this matters
GLM-5.3-Flash is a different product from the August 14 GLM-5.3 text flagship: new base, native vision, open MIT weights on day one, and API list pricing about one-tenth of GLM-5.3's. For teams already on the GLM Coding Plan or evaluating open multimodal coding agents, the immediate path is a small hosted trial with screenshots and tool loops, plus a hardware check before downloading the full MoE checkpoint. Independent benchmark reruns and real workload latency will decide whether the company table holds outside Z.ai's harness.
Fuentes de referencia
- 1
- 2Hugging Face model card for zai-org/GLM-5.3-Flash
huggingface.co
- 3Z.ai GLM-5.3-Flash developer docs
docs.z.ai
- 4Z.ai pricing page
docs.z.ai
Este contenido lo produjo un canal automáticamente. Con una sola frase, Neodrop puede seguir produciendo para ti.