Alibaba's Qwen team has released a new lineup of voice AI models called Qwen-Audio 3.1, covering speech recognition, speech synthesis, and real-time voice interaction. The five-model refresh brings performance improvements alongside steep price cuts, with API costs down by as much as 95 percent depending on the use case.
Five Models Covering Recognition, Synthesis, and Conversation
The new lineup consists of five models: an automatic speech recognition model (ASR), a higher-tier ASR-Next, a text-to-speech model (TTS), a higher-tier TTS-Next, and a bidirectional real-time interaction model. The release replaces the previous Qwen-Audio generation and is offered exclusively as a hosted cloud API through Alibaba Cloud rather than as downloadable model weights.
New Features: Emotion Detection and Ambient Sound Awareness
The ASR model improves recognition accuracy across multiple languages and dialects, and automatically cleans up filler words and repeated phrases when producing a transcript. ASR-Next builds on this with speaker diarization that labels who said what with timestamps, emotion estimation, and recognition of background and ambient sounds.
TTS supports natural cross-language voice transfer, letting developers control tone and delivery through text prompts alone. TTS-Next combines a language model with a diffusion-based approach, generating voice, sound effects, and ambient audio together in a single pass.
The real-time model is designed to listen and speak simultaneously, much like a human conversation, and can respond instantly when interrupted mid-sentence. Alibaba also says it can gauge a speaker's mood from their voice and adjust its response speed and empathy accordingly, though no benchmark figures for this capability have been published.
API Pricing Cut by as Much as 95%
Alongside the new models, Alibaba slashed API pricing: speech recognition costs are down by up to 95 percent, text-to-speech by about 70 percent, and real-time interaction by roughly 85 percent. The cuts appear aimed at lowering the barrier for companies looking to build voice AI features without high development and operating costs.
Voice AI has become a competitive battleground, with OpenAI, Google, and other major Western AI companies rapidly expanding their own conversational voice features. Chinese tech companies are ramping up investment in the space as well, and Alibaba's aggressive price cuts appear to reflect that broader competition on both performance and cost.
Summary
On September 23, 2026, Alibaba released Qwen-Audio 3.1, a new lineup covering speech recognition, speech synthesis, and real-time voice interaction. With new capabilities like speaker diarization and emotion estimation, combined with API price cuts of up to 95 percent, the barrier to building voice AI into products just got lower. How much traction Chinese voice AI can gain in practical deployments remains one to watch.
