TREE NEWS update: Alibaba’s Qwen released the Qwen-Audio-3.1 series of speech models, upgrading its ASR, TTS and Realtime voice interaction models and adding the new Qwen-Audio-3.1-TTS-Next audio creation model and Qwen-Audio-3.1-ASR-Next audio understanding model. Five new speech models were launched together, covering understanding, generation, interaction and creation. Prices across the Qwen-Audio line were cut, with TTS down about 70%, Realtime about 85% and ASR 95%.
Alibaba’s Qwen Releases Qwen-Audio-3.1, Cuts Speech Model Prices by Up to 95%
The headline number is the price cut, but the more consequential move is the breadth of the release: understanding, generation, interaction and creation models shipped together, which turns speech from a single API call into a stack. That matters most for developers building voice agents, where ASR and TTS costs have been a real constraint on always-on interaction. Whether rivals match these prices, and whether cheap speech plus capable models pulls more builders into voice-first products, is the open question.
Generated by AI for reference only.
Share on WeChat
Open WeChat → Scan → then tap "…" to send to a chat or Moments.
Tap "…" in the top-right corner to send to a chat or share to Moments.