Summary
Fahd Mirza puts Qwen Audio 3.0 TTS — released by Alibaba’s Tongyi team in Flash and Plus tiers, built on a 12.5 Hz low-frame-rate speech tokenizer with a five-stage progressive training pipeline — through a systematic evaluation after the model claimed the top position on an independent TTS leaderboard. Mirza notes upfront that Alibaba has moved to API-key-only access rather than releasing open weights, a shift he criticizes as disappointing for the open-source community.
The evaluation runs three distinct test suites across the model’s claimed 16 supported languages (Spanish, German, Arabic, Chinese, Vietnamese, Japanese, Italian, Indonesian, French, Korean, and others). The first suite tests basic cross-lingual synthesis quality benchmarked against ElevenLabs, MiniMax, and Vox CPM 2 on the CommonVoice CV3 evaluation set. The second suite tests instruction-controlled delivery, where natural-language prompts like “speak slowly and sadly” or “speak like a news anchor” are passed alongside the text — a feature Alibaba has only officially documented in Chinese and English. The third suite tests fine-grained inline emotion tags (whisper, angry, happy) injected mid-sentence to shift delivery at a specific word rather than setting a single tone for the whole clip.
Results show strong instruction following in supported languages, with Japanese, Italian, and Indonesian performing particularly well. Arabic synthesis failed on the first attempt. Mirza flags clearly where tested behavior goes beyond Alibaba’s documented guarantees, making this a candid and practically useful reference for developers evaluating multilingual TTS pipelines.
📺 Source: Fahd Mirza · Published July 24, 2026
🏷️ Format: Benchmark Test







