Ant Group's subsidiary, Ant LingBot, has released six open-source embodied AI models. The company is conducting parallel research in VLA and world models, but faces challenges related to data scarcity and ecosystem competition.
LingBot's Creation and Strategy
Ant Group established Ant LingBot as a wholly-owned subsidiary in Shanghai in December 2024. This structure functions as a division dedicated to implementing physical AI and employs over 90% staff members with a master's or doctoral degree. Under the leadership of CEO Zhu Xin and Chief Scientist Shen Yujun, LingBot unveiled six open-source embodied AI models at the WAIC 2026 conference. These models cover areas such as vision, video, spatial perception, manipulation, world models, and world action models.
Dual Technological Approach
LingBot's technological strategy features a dual direction. The VLA (vision-language-action) route, presented by LingBot-VLA 2.0, aligns with the widely accepted perception-to-action architecture used by Figure and the Google DeepMind RT series. The second route, based on world models and presented by LingBot-VA 2.0, focuses on predicting changes in the world before actions are generated. LingBot-VA 2.0 is positioned as the industry's first natively embodied world action model, trained from scratch on an autoregressive architecture. It provides real-time inference at a frequency of 150 Hz on a single GPU. The model's design incorporates semantic vision-action tokenization, strict causal pre-training, an MoE architecture, and enhanced asynchronous inference, allowing the robot to predict future states during current actions, responding in 6.7 milliseconds, which is significantly faster than human blinking (300–400 milliseconds).
Ecosystem and Partnerships
The open-source strategy is a key element of LingBot's ambition to build an ecosystem. Since technical standards in embodied AI are still forming, LingBot has provided full weights, code, post-training toolkits, and benchmark standards. The models are compatible with various robot configurations, enabling different hardware platforms to utilize the same 'brain.' More than a dozen robot manufacturers have entered into partnerships, including Unitree Robotics, Xinghaitu, and Leju Robotics. For instance, Leju Robotics' KUAVO 4 Pro successfully adapted LingBot-VLA in 95 real-world manipulation scenarios. Furthermore, LingBot developed its own service robot, Robbyant R1, intended for home use, elderly care, and healthcare, although shipment volumes remain limited.
Data Challenges and Competition
The most significant limitation for LingBot is the data problem. The embodied AI industry faces a fundamental data gap: while large language models can be trained on trillions of tokens of human language, physical interaction data must be collected through actual robot operation. LingBot employs a 'borrow eggs to fry an omelet' model, obtaining real robot data from ecosystem partners like Unitree, Leju, and Xinghaitu instead of operating its own large robot fleet. This dependency creates a strategic vulnerability, as partners are simultaneously developing their own AI capabilities. LingBot is the only major subsidiary of a Chinese tech firm that is simultaneously developing foundational brain models, world models, and hardware deployment. This broad strategic scope causes resource contention within Ant Group and competition with other embodied AI initiatives affiliated with Alibaba. LingBot is testing whether the fintech giant can successfully develop a robotics business, and whether open world action models and data borrowing practices can overcome embodied AI limitations. Unlike Alibaba, Tencent, Huawei, JD.com, and Meituan, which adhere to more focused strategies, LingBot demonstrates a unique full-stack approach—from brain models to world models and hardware deployment.