ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variant
ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline natural-language direction tags like [laughs] and [whispers], 90+ languages, 10-second voice cloning, scene-level context for dialogue, and a Turbo variant at ~100ms median inference latency for voice agents; ranked #1 by Artificial Analysis.
Why it matters
Splits TTS into two lanes in one launch: a quality flagship with scene-level context and stackable direction tags, plus a ~100ms streaming Turbo aimed squarely at live voice agents; Artificial Analysis #1 ranking is third-party validation.
What you could build with it
A voice-agent startup could ship a multilingual support agent on v4 Turbo's ~100ms streaming for turn-taking that feels live, while a creator-tools team could offer director-style audiobook narration where authors embed [laughs] and [whispers] inline instead of tuning SSML.
Does it hold up?
Hours old: third-party validation exists (Artificial Analysis #1, 1,674 samples), but no independent production deployments yet; the 10-second cloning and tag-following claims are vendor-reported and unverified outside their tests.
Built with ElevenLabs Eleven v4 + v4 Turbo
- What the Eleven v4 latency numbers actually measurearticle · Independent read separating the ~100ms median inference figure (network excluded) from the 150ms time-to-first-speech claim, and noting the 75% blind-preference figure is ElevenLabs' own test.
- ElevenLabs official v4 launch postx · Official September 28 launch post on X with demo clips of the new architecture.
- Hands-on v4 test: AI dragon audition, emotion tags, 10s voice cloningyoutube · Independent creator puts Eleven v4 through real-world tests: character acting with audio tags, laughing/whispering, multi-speaker conversation, multilingual cloning from a 10-second sample, and v4 Turbo speed.
Learn more
First spotted on article: source.
More AI models
VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66Claude Sonnet 5.5: 30% faster, up to 30% cheaper per taskAnthropic's new mid-tier: 70.6% on Terminal-Bench 4.0 (vs 10.3% for Sonnet 5), 46.2% FrontierCode 1.1 max-effort…models · JEV 0.65
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.