Back
Alibaba's Qwen launches Audio 3.1 with five new models and cuts audio API prices by up to 95%
SiTech AI Team2 წთ. საკითხავი

Alibaba's Qwen launches Audio 3.1 with five new models and cuts audio API prices by up to 95%

Alibaba's Qwen team released Qwen Audio 3.1, a five-model family for speech recognition, text-to-speech and real-time voice interaction, while cutting prices by about 70 percent for TTS, 85 percent for the realtime model and up to 95 percent for ASR.

Alibaba's Qwen team has released Qwen Audio 3.1, a family of five models covering speech recognition, text-to-speech and real-time voice interaction. The release came with steep price cuts on the company's audio APIs: text-to-speech drops by about 70 percent, the realtime model by roughly 85 percent and speech recognition by up to 95 percent.

What the five models do

The core speech recognition model (ASR) improves multilingual and dialect recognition and automatically cleans up filler words and repetitions, which matters for transcribing interviews, calls and meeting recordings without heavy post-processing. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds and machine noise in the audio.

On the synthesis side, the TTS model produces multilingual speech with what Qwen calls natural cross-language voice transfer. Emotion, speed and style are set through plain text prompts, for example "Read this with a sharp, commanding tone, demanding respect." TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects and background audio in a single pass.

The fifth model focuses on real-time interaction: it supports speaking and listening at the same time and allows instant interruption. According to Qwen, when it detects a low mood it answers more slowly and with more empathy.

Price cuts of up to 95 percent

The reductions are the sharpest part of the announcement. Speech recognition, usually the highest-volume workload in voice products, becomes up to 95 percent cheaper; the realtime model falls about 85 percent and text-to-speech about 70 percent. For teams running large-scale transcription or voice agents, that changes the economics of keeping audio pipelines running around the clock.

Qwen published technical details on its blog and offers the models on Qwen Cloud under the Qwen Audio 3.1 name.

Why it matters

Voice interfaces have become one of the most competitive layers of the AI stack, and providers now compete on price as much as on quality. Cuts of this size in recognition and realtime APIs raise the bar for other vendors and make always-on voice features — call centres, assistants, accessibility tools — cheaper to operate.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.