Qwen3.8-Omni-Flash: audio and video in, text out, video input at $0.20 an hour

Qwen3.8-Omni-Flash: audio and video in, text out, video input at $0.20 an hour

Qwen's first agentic omni-modal model takes audio and video in and returns text, ships as a hosted API at roughly a tenth of its predecessor's media price, and beats Gemini 3.8 Flash on agentic tool use and meeting transcription while trailing it on long-video reasoning, all on Qwen's own chart.

Qwen released Qwen3.8-Omni-Flash on September 18, 2026, its first omni-modal model built around agentic work. Text, images, audio and video go in one request, the context holds 1M tokens, and the model plans tasks and calls tools on what it finds. 1 A second endpoint, Qwen3.8-Omni-Flash-Realtime, is built for continuous live audio and video streams. 1
Two constraints decide what you do with it. Qwen shipped the model hosted, on QwenCloud, Alibaba Cloud Model Studio and Qwen Studio, with no weights published. 2 Its answer is text, and generated speech still routes to the older Qwen3.5-Omni. 3

What launched

SignalConfirmed detailAction window
Hosted onlyLive September 18, 2026 on QwenCloud, Model Studio and Qwen Studio; no open weights. 2Call it today; nothing to download.
Media in, text outText, image, audio and video input; 1M-token context, 991K max input, 131K max output; thinking on by default. 2Voices still mean Qwen3.5-Omni.
Agentic long videoQwen reports OmniVideoBench accuracy up from 63.4 to 67.8 as tokens per query fall from 145,736 to 79,117. 1Test it on your own multi-hour footage.
Price$0.15 per 1M input and $0.47 per 1M output tokens; video with audio at $0.20 an hour against $3.27 for Qwen3.5-Omni-Plus. 12Redo the media-cost arithmetic.
Documented limitsVideo up to 2 hours and 2 GB by URL, audio up to 3 hours, 113 input languages, six regions. 3A two-hour recording goes in one call.
Qwen's comparison chart for Qwen3.8-Omni-Flash across agentic, audio-visual, audio and price panels
Qwen's own comparison chart, published with the launch: deep blue is Qwen3.8-Omni-Flash, pale blue Qwen3.5-Omni-Plus, grey Gemini 3.8 Flash, with per-hour media prices along the bottom. Qwen's wins cluster in the agentic tool-use, multi-speaker and sound-grounding panels; Gemini 3.8 Flash keeps the audio-visual reasoning row. Image: Qwen

What the chart shows

The launch claim — audio-visual performance close to Gemini 3.8 Flash, overall audio above it — splits panel by panel. Qwen3.8-Omni-Flash takes the panels where a tool is doing the work: 71.0 against 58.9 on WildClawBench-MM and 69.6 against 69.0 on UniClawBench. It also dominates multi-speaker transcription, where lower is better: 3.4 and 17.2 on AliMeeting's diarization and word-error measures against Gemini 3.8 Flash's 72.6 and 53.1. 1
Gemini 3.8 Flash holds the audio-visual reasoning panels, at 71.0 against 65.0 on Video-MME-v2, 70.7 against 63.3 on LVOmniBench, 65.2 against 63.4 on OmniVideoBench and 45.0 against 36.8 on AgenticVBench. 1
Agent mode narrows part of that. Inside Qwen's Qwen Code harness, Qwen3.8-Omni-Flash's LVOmniBench score climbs from 63.3 to 73.6, past the 70.7 where Gemini 3.8 Flash stays; Gemini 3.8 Flash rises from 65.2 to 70.1 on OmniVideoBench under the same harness. 1
The headline agentic gain averages two benches. Qwen's launch post puts it at +19.5 points across WildClawBench-MM and UniClawBench; against Qwen3.5-Omni-Plus, the chart shows +36.5 on the first and +2.5 on the second. 14

Why it matters

The price is the news for long-recording pipelines. Video with audio costs $0.20 an hour where the model it replaces charged $3.27, and input tokens cost $0.15 per million against Gemini 3.8 Flash's $0.75. 1
Every number above is Qwen's, on its own harnesses; no independent evaluation had appeared by publication. 5
Qwen also shipped tooling. Qwen-MM-Plugins installs as skills and optional MCP servers into Claude Code, Codex, Qwen Code, Gemini CLI and others, wrapping capabilities such as an audio-visual memory of a long video and speaker-preserving video translation. Its README records one gap: most harnesses cannot yet feed audio to the main model, so audio routes through the API. The companion Qwen-Live Harness is further off: the blog calls it open-sourced, the launch post marks it "coming soon", and its repository link does not resolve. 6

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel