
Qwen-Audio-3.0-TTS splits the launch between speed and voice quality
Alibaba's new hosted TTS line pairs a real-time Flash tier with a higher-quality Plus tier, adding multilingual control while leaving throughput and local deployment as practical tradeoffs.
Qwen now has a hosted TTS line
Alibaba's Qwen announced Qwen-Audio-3.0-TTS at 12:33 UTC on July 23, with two versions: Flash for real-time interaction and Plus for higher-quality generation. The launch post says the system supports 16 languages, one-pass long-form synthesis up to three minutes, and 48 kHz output. 1 2
This is a hosted product release rather than a new downloadable Qwen checkpoint. A detailed report identifies the Model Studio model IDs as
qwen-audio-3.0-tts-flash and qwen-audio-3.0-tts-plus, with bidirectional WebSocket streaming and PCM, WAV, MP3, and Opus output. 3The control layer is the main change
Qwen says the model accepts free-form instructions such as asking for a slower bedtime-story delivery, alongside 86 newly added inline tags. Those tags can steer a phrase or word toward effects such as whispering, anger, breathing, coughing, or laughter. The system also targets zero-shot voice cloning, cross-lingual generation, emotional delivery, Chinese dialect synthesis, and difficult text normalization. 2
The architecture is tuned for production constraints as well. Qwen describes a 12.5 Hz low-frame-rate speech tokenizer intended to reduce inference cost, plus generation from noisy, reverberant, or unclear reference audio without an explicit denoising mode. Vocoder super-resolution is used for 48 kHz output. These are useful details for applications that need controllable voices, but they do not remove the need to test latency and audio quality on real traffic.
The ranking needs context
Qwen says Qwen-Audio-3.0-TTS-Plus is first on the independent Artificial Analysis Text-to-Speech Leaderboard, after evaluations covering SEED-TTS-Eval, CV3-Eval, instruction following, long-form generation, and acoustic robustness. 2 A July 20 report puts Plus at roughly 1,236 Elo, just above Simba 3.2 at 1,234, while noting that their confidence intervals overlap. 3 The result is a strong early signal, not a decisive claim that it is better for every voice or workload.
The tradeoffs are practical. The same report describes Plus throughput at about 16 characters per second, below several competing systems, and says emotion and rich-language tags are limited to unidirectional streaming. Both versions are API-only, so teams that need local deployment, predictable high concurrency, or full control over weights will still need another option. For developers, the important split is therefore simple: Flash is the latency choice, Plus is the quality choice, and neither should be adopted on a leaderboard position alone.
Cargando tarjeta de contenido…
Fuentes de referencia
Contenido relacionado
- Inicia sesión para comentar.
More from this channel›
- Claude Opus 5 brings near-Fable performance to the Opus tier
- Kimi K3's second week turns a model launch into a compute and policy test
- OpenAI Presence brings managed voice and chat agents to enterprise workflows
- Google launches Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber
- Qwen3.8 Max preview lands with 2.4T parameters, but proof is still to come
- Thinking Machines launches Inkling, a 975B open-weight model built for customization