Sber has announced the release of the Kandinsky WM 1.0 family of generative models, designed to produce video materials with a high degree of physical plausibility. These videos are necessary for training Physical AI systems, including autonomous vehicles and various robots.
According to a press release obtained by N + 1, the code and weights of these models were published under the MIT license, making them free for developers to use.
Physical AI systems are defined as intelligent systems capable of interacting with the real physical environment, including pedestrians, cars, industrial equipment, and other robots. The foundation of training such systems is extensive sets of video data. However, collecting this data is often costly and labor-intensive, and recording certain critical scenarios, such as accidents, equipment failures, or extreme weather events, in real conditions is impossible.
Furthermore, fine-tuning robots, such as loaders or self-driving taxis, directly in the operational process presents not only high costs but also potential danger.
Previously, the use of videos generated by neural networks for training Physical AI faced difficulties due to numerous artifacts: geometric distortions, perspective issues, and the appearance of objects with unnatural, 'soft' deformation. Sber developers, led by Denis Dimitrov, solved this problem by creating a specialized family of models for generating physically accurate video scenes.
The new models are based on Kandinsky 5.0 Video Lite but have undergone additional training on millions of video recordings collected from robot cameras, autonomous vehicles, and during industrial processes. According to the developers, this has significantly reduced common errors in video generation.
Thanks to these Kandinsky WM models, it has learned to create physically correct video clips lasting about five seconds. These clips can demonstrate various actions and situations, such as object manipulation, road scenes, or stages of production cycles. The resulting materials are useful for training computer vision systems and simulators for autonomous transport.
To improve the quality of automotive scenes, an additional reinforcement learning stage was implemented: Kandinsky WM generated videos, and other neural networks acted as evaluators, checking the preservation of object shape, the geometric accuracy of the scene, and the naturalness of movements, providing feedback.
The Kandinsky WM 1.0 models are already available in the open repository of the Kandinsky project. The press release notes that the AIRI institute uses these models in its road simulator, and Sber's Robotics Center applies them to train its own robots.
Another part of the material mentions that the Airbus consortium presented the U145 mockup—an unmanned version of the H145 helicopter. In this modification, the pilot cabin was replaced with a set of sensors with an AI control system, and instead of the cockpit, there are doors with a folding table for loading. The first flight is planned for the end of 2026, and full commissioning is expected at the beginning of the next decade.



