Radar / AI media tools / Kandinsky 6.0 Video
Kandinsky 6.0 Video: open MIT-licensed audio-video generation models
Sber's team released Kandinsky 6.0 Video: Lite (3B) and Pro (29B) diffusion models generating 5-second clips with synchronized 44kHz audio and lip-sync, dual-stream CrossDiT architecture, code and checkpoints under MIT.
Why it matters
A genuinely open (MIT) synchronized text-to-audio-video model family with diffusers integration, claiming Pro beats Kandinsky 5.0 and competes with leading closed audio-video models on speech quality.
What you could build with it
An indie developer or small studio could build a multilingual lip-synced dubbing service for short-form video creators on top of the MIT-licensed Lite model, producing localized ad variants with native audio. A marketing agency could generate product videos with synchronized voiceovers in multiple languages without reshooting.
Does it hold up?
Too early to judge. The paper appeared October 4 and code/checkpoints were only just released; no independent community builds or benchmark reproductions are visible yet.
Built with Kandinsky 6.0 Video
- Kandinsky 6.0 Video paper page with abstract and methodsPapers with Code · Canonical paper page summarizing the dual-stream CrossDiT architecture, continuous pretraining strategy, and human-evaluation results versus Kandinsky 5.0.
- Sberbank Kandinsky Video product overviewTAdviser · Product overview of the Kandinsky line, noting open developer access to the earlier Kandinsky Video Lite and the open-license positioning of the family.
Learn more
First spotted on arxiv: source.
More AI media tools
OpenMontage: open-source agentic video production systemTurns an AI coding assistant into a full video studio: 12 pipelines and 100+ tools for scripting, footage, narration…media · JEV 0.73HyperFramesHeyGen's open-source framework that turns HTML/CSS/GSAP into deterministic MP4 video via 21 agent skills and a router…media · JEV 0.69Meshy crosses $100M ARR and launches official iOS/Android appsMeshy announced Sep 30 that ARR surpassed $100M (up 100x in under two years, first AI-3D company at that milestone)…media · JEV 0.61FLUX 3 ImageBlack Forest Labs' new multimodal image model with bounding-box layout control, up to 10 reference images, and native…media · JEV 0.57UniMate: one model to animate any skeleton from textSIGGRAPH Asia 2026 paper and open code: a topology-aware diffusion transformer that generates motion for any rigged 3D…media · JEV 0.43Tencent Hy Image 3.5 PreviewTencent's latest image model with multi-turn conversational editing, direct 4K output, and up to 20 reference images.media · JEV 0.43
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.