Embodied AI has become the most prominent race in robotics in 2026. At the WAIC exhibition last month, demonstrations of humanoid robots dancing and fighting gave way to displays of robots operating on real production lines, performing assembly, sorting, and parts picking tasks on full-scale factory replicas.
Investors closely monitored this area: in the first half of 2026, over 300 deals in embodied AI and related supply chains in China exceeded 90 billion yuan. Furthermore, the number of employees involved in embodied AI roles increased fourfold in just the first four months.
However, beneath all this activity lies a common problem—data. Models with limited computational power and algorithms with limited capabilities are now facing the fact that data is what restricts their development potential. Robots cannot automatically absorb accumulated internet content, such as text, images, and videos. They require a different array of information: an understanding of friction between materials, forces generated when handling different weights, and the causal relationship between grasping and an object—all the subtle physics that websites do not provide.
According to industry experts, the global supply of high-quality manipulation data amounts to only a few hundred thousand hours, whereas supporting a general-purpose robot likely requires tens of millions of hours.
Teleoperation allows for the collection of high-precision data, but it is a costly process. First-person videography scales cheaply, but it lacks detail regarding joints, forces, and trajectories. Pure simulation, on the other hand, fails when transferred from the virtual environment to the real world. Liu Shenxiang, founder and CEO of WUWENAI, who previously led autonomous driving data and testing projects at Baidu, views these methods not as alternatives, but as a unified pipeline.
The system includes multimodal real-world data collection complexes for continuous acquisition of physical interaction data. Concurrently, a generative world model expands scenarios and edge cases in simulation, and both elements are continuously used for training and further data generation. WUWENAI calls this a 'closed loop for merging virtual and real data,' built upon three components: the Data Factory, the World Model, and the World Simulator, which enables large-scale training, evaluation, and feedback.
Liu believes the company goes beyond simply being a data provisioning platform. Once a high-precision world simulator can reliably reproduce real physics, a robot will be able to perform an AlphaGo-level maneuver, conduct millions of inexpensive trial-and-error cycles in simulation, and transition from imitation learning to reinforcement learning. WUWENAI began supplying this Real2Sim2Real stack several months before SceniX was acquired by World Labs.
Liu emphasizes the parallel as a signal to the industry: world models only become infrastructure when integrated with real-world data, robotic simulation, scaled training, and real-world feedback. Video and 3D generation alone are insufficient.
The company's bet is structural, not product-based. Embodied AI models are limited not so much by the number of parameters as by the volume of data they consume, and the data industry is transforming from simple raw data collection to the creation of orchestrated infrastructure. The next test for the Chinese startup WUWENAI is whether it can become a world-class company in physical AI infrastructure amidst giants like World Labs, Tesla with Optimus, and Google DeepMind, which are developing their own stacks in parallel. The road is being paved regardless, and WUWENAI has decided to take on the role of the asphalt paver.