Qwen3.8-Max goes GA with 2.4T parameters, 1M context, and weights promised next week

Qwen3.8-Max goes GA with 2.4T parameters, 1M context, and weights promised next week

Qwen3.8-Max is live through QwenCloud with a 1M-token context and strong company-reported agent benchmarks, but its promised open-weight release is the real deployment test.

What shipped

Alibaba's Qwen announced Qwen3.8-Max on August 3 as a hosted release, not another preview. The model is a 2.4-trillion-parameter mixture-of-experts system with 95 billion active parameters, accepts text, images, and video, and returns text. Qwen says it is available through QwenCloud's API under the model ID qwen3.8-max; the weights are promised for release next week. 1
The product pitch is broader than chat quality. Qwen positions the model for coding agents, professional workflows, multimodal document and video work, and tasks that run for days. QwenCloud lists function calling, structured outputs, web search, code execution, batching, prefix completion, fine-tuning, and context caching as supported features. 2

The benchmark story is mixed

Qwen's release table reports 86.6 on Terminal-Bench 2.1, ahead of Claude Fable 5 at 84.6 but behind GPT-5.6 Sol at 88.8. It also reports 93.0 on PaperBench and 92.6 on GPQA Diamond. Those are company-published results, and the comparisons are not all like-for-like: the table uses different harnesses and evaluation settings, while Qwen's multimodal section compares against Qwen3.7-Plus rather than Qwen3.7-Max. 1
There is no independent re-run to treat as a final verdict yet. BenchLM's August 3 profile shows 52 source-displayable benchmark rows, but labels the values "provider exact"—useful provenance auditing, not independent testing. Its broader catalog puts the model at 65.4/100 and #31 of 215, a reminder that a launch table and a cross-model aggregate answer different questions. 3

What developers can test now

The hosted API is immediately testable. QwenCloud lists a 1M-token context, up to 991K input tokens, 131K output tokens, and a 262K reasoning budget. Listed pricing is $2 per million input tokens and $6 per million output tokens, with lower rates for cached input. 2
That makes Qwen3.8-Max worth evaluating for long-context, multimodal, and agent workflows now. It is not yet a self-hosting decision: the weights, license, quantization options, and serving requirements are still unresolved. The release matters because Qwen is moving its Max tier toward open weights, but the practical impact will be decided by the artifact that arrives next week—not by today's 2.4T headline.
AI Model & Product Launch Alerts

AI Model & Product Launch Alerts

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

  • Sign in to comment.