Kandinsky 6.0 brings video, audio, and lip-sync together in an open-source generation model

Kandinsky 6.0 moves into video with synchronized audio, lip-sync and Full HD output across two open-source Lite and Pro models.

Kandinsky Lab is moving into a new phase with Kandinsky 6.0 Video, a family of models designed to generate video and its soundtrack simultaneously. Introduced in a research paper published on arXiv and now available as open source on GitHub, the family comes in two versions: Lite, with 3 billion parameters, and Pro, with 29 billion.

Both models generate five-second clips from either a text prompt or an input image, with synchronized 44 kHz audio. Kandinsky Lab also supports lip-sync, while a dedicated super-resolution component can bring the final output up to Full HD at 1920 × 1080. Motion, dialogue, ambient sound and audio effects can therefore be generated as part of the same process rather than adding a soundtrack after the video has been created.

Kandinsky 6.0 uses a dual-stream CrossDiT architecture to connect the two modalities. A pretrained video stream is paired with an audio stream through bidirectional cross-attention intended to maintain temporal and semantic consistency between sound and visuals. Training combines joint pretraining, supervised fine-tuning, reinforcement learning-based post-training and distillation. According to the team's evaluations, the Pro version clearly improves on Kandinsky 5.0 Pro and remains competitive with other audio-video generation systems, particularly for speech quality.

Kandinsky Lab has released the code, checkpoints and Diffusers integration under the MIT license. Pro, Pro Distill, Lite and Lite Distill are among the variants already available through the Kandinsky 6.0 collection on Hugging Face. The distilled Pro checkpoint, for example, runs generation in ten inference steps. ComfyUI integration and vLLM-Omni support are also part of the release.

The release is already moving beyond local deployment. fal has made both Kandinsky 6.0 Lite for text-to-video and Kandinsky 6.0 Pro available through its platform, alongside image-to-video endpoints and separate VSR models for video super-resolution. fal currently lists a five-second 480p Lite generation at $0.16, compared with $1.35 for Pro, reflecting the split between faster iteration and the heavier flagship model.

With Kandinsky 6.0, Kandinsky Lab is joining the broader shift toward native video systems that handle visuals, motion and sound within a single generation process. The simultaneous release of the code and model checkpoints is arguably the more distinctive part of the launch, making the technology available for direct testing, integration and adaptation rather than restricting it to a proprietary interface.