The Maivia Gazette

Verified AI news, every morning

Models

Alibaba's Qwen Audio 3.1 adds five speech models and cuts audio prices by up to 95 percent

The lineup covers transcription, speech synthesis and real-time conversation, with speech-recognition prices dropping the most.

Five tape reels on a descending staircase of yellow blocks, with tape ribbons flowing down the steps.
AI-generated illustration, not event photography. The motion is AI-generated from the still.

Alibaba's Qwen team released Qwen-Audio-3.1, a set of five models for speech recognition, text-to-speech and real-time interaction. The base speech-recognition model improves recognition of multiple languages and dialects and automatically removes filler words and repetitions. ASR-Next adds identification of multiple speakers with timestamps, and detects emotions, ambient sounds and machine noise. The TTS model handles multilingual synthesis with voice transfer across languages. Users steer emotion, speed and style through text prompts. TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects and background audio in a single pass. The real-time model can speak and listen at the same time and can be interrupted instantly. According to Qwen, it responds more slowly and with more empathy when it detects a low mood. Alibaba is also cutting prices: text-to-speech by about 70 percent, the real-time model by about 85 percent and speech recognition by up to 95 percent. The cuts follow Qwen's earlier aggressive pricing on multimodal models and add pressure on rival speech API providers. The Decoder's report does not include independent benchmark results.

Sources

  1. The DecoderAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percentPublished · fetched

Also in this edition