Qwen's production speech-recognition system, documented in a 45-author technical report on 7 September 2026: a Mixture-of-Experts, LLM-based ASR model built on the Qwen backbone and trained on tens of millions of hours of speech, framed around the gap between benchmark accuracy and production utility. It transcribes 30 languages and 16 Chinese dialect varieties across eight dialect regions, and adds the production features the report singles out: industry-domain entity recognition, hierarchical hotword customisation, single-pass transcription polishing, and long-audio contextual modelling. A dedicated Qwen-Audio-3.0-ASR-Streaming variant targets latency-sensitive use. The report claims state-of-the-art or highly competitive results on Chinese, English, multilingual, and industrial test sets against leading commercial systems. It is the served successor to the small open Qwen3-ASR models; no weights are released with the report.

Model Details

Architecture MOE

Paper

speechaudiomultilingualmoeresearch

Related