Radar / AI media tools / Kandinsky 6.0 Video

Kandinsky 6.0 Video: open MIT-licensed audio-video generation models

Sber's team released Kandinsky 6.0 Video: Lite (3B) and Pro (29B) diffusion models generating 5-second clips with synchronized 44kHz audio and lip-sync, dual-stream CrossDiT architecture, code and checkpoints under MIT.

new launchmedia · imageJEV traction 0.59added 2026-10-07
Open Kandinsky 6.0 Video →View on the radar

Why it matters

A genuinely open (MIT) synchronized text-to-audio-video model family with diffusers integration, claiming Pro beats Kandinsky 5.0 and competes with leading closed audio-video models on speech quality.

What you could build with it

An indie developer or small studio could build a multilingual lip-synced dubbing service for short-form video creators on top of the MIT-licensed Lite model, producing localized ad variants with native audio. A marketing agency could generate product videos with synchronized voiceovers in multiple languages without reshooting.

Does it hold up?

Too early to judge. The paper appeared October 4 and code/checkpoints were only just released; no independent community builds or benchmark reproductions are visible yet.

Built with Kandinsky 6.0 Video

Learn more

paper summary →

First spotted on arxiv: source.

More AI media tools

OpenMontage: open-source agentic video production systemTurns an AI coding assistant into a full video studio: 12 pipelines and 100+ tools for scripting, footage, narration…media · JEV 0.73HyperFramesHeyGen's open-source framework that turns HTML/CSS/GSAP into deterministic MP4 video via 21 agent skills and a router…media · JEV 0.69Meshy crosses $100M ARR and launches official iOS/Android appsMeshy announced Sep 30 that ARR surpassed $100M (up 100x in under two years, first AI-3D company at that milestone)…media · JEV 0.61FLUX 3 ImageBlack Forest Labs' new multimodal image model with bounding-box layout control, up to 10 reference images, and native…media · JEV 0.57UniMate: one model to animate any skeleton from textSIGGRAPH Asia 2026 paper and open code: a topology-aware diffusion transformer that generates motion for any rigged 3D…media · JEV 0.43Tencent Hy Image 3.5 PreviewTencent's latest image model with multi-turn conversational editing, direct 4K output, and up to 20 reference images.media · JEV 0.43

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.