Qwen-Audio-3.1 Speech Family + Qwen3.8-LiveTranslate
Alibaba's next-gen speech stack announced at Yunqi 2026: ASR-Next transcribes with emotion, music, and ambience understanding; TTS-Next generates dialogue fused with environmental audio; Realtime does listen-think-speak; LiveTranslate cuts simultaneous-interpretation latency below 2.5s, with ASR prices cut up to 95%.
Why it matters
One family covers the full audio pipeline (understanding to cinematic generation to real-time interpretation), and the 95% ASR price cut aggressively undercuts dedicated STT APIs; dialogue-plus-ambience TTS is a new capability for audiobook, podcast, and video workflows.
What you could build with it
Multilingual podcast localization service: ASR-Next transcribes source episodes with emotion and diarization metadata, then TTS-Next regenerates each episode in 10+ languages preserving dialogue and background ambience; the ASR price cut makes per-episode cost negligible for indie podcasters.
Does it hold up?
Too early to judge: announcement-stage with no independent hands-on usage yet. The aggressive ASR price cut is the most concrete, immediately usable part.
Built with Qwen-Audio-3.1 Speech Family +…
- AIBase daily: Qwen-Audio-3.1 and LiveTranslate breakdown (Sept 23)news · AIBase's roundup of the Yunqi speech announcements and price cuts.
- Alibaba unveils roadmap: full-stack AI strategy (CNA, Sept 23)news · Channel NewsAsia coverage of Alibaba's full-stack AI roadmap including the audio family.
- THE DECODER: five-model breakdown with sample audionews · Hands-on style breakdown of each model (ASR, ASR-Next, TTS, TTS-Next, Realtime) with embedded sample audio and price-cut analysis.
- ArtRealmAI: TTS-Next script-to-soundscape for audio creatorsblog · Creator-focused writeup on using TTS-Next to turn a text script into a full mix of dialogue plus ambient sound in one pass.
- AlphaSignal: Alibaba slashes Qwen-Audio API prices up to 95%news · Signal-style analysis of the price cuts (TTS ~70%, Realtime ~85%, ASR up to 95%) and the one-stack voice-loop integration argument.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.