Robotera, which has support from Tsinghua University, announced that its world-class model VPP2 now ranks first in RoboDojo. This simulation benchmark is curated by MMLab from the University of Hong Kong and involves nearly 20 academic institutions.
According to QbitAI data, VPP2 demonstrated an average success rate of 32.26% and an average score of 39.26 across 42 tasks involving two RoboDojo manipulators. These tasks test aspects such as generalization, precise manipulation, long-horizon tasks, memory, and open-vocabulary instructions. The model also secured first place in the generalization, precision, and memory categories.
Robotera emphasizes that achieving these metrics was done without using additional training data or add-ons, such as agent-based recursive self-improvement, relying solely on the standard dataset. The strongest control model mentioned in their research—GPT-6-Astra after additional training—showed results of 22.48% and 28.97. The ranking aligns with the figures presented in the scientific paper. Previously, the company Simate claimed success with its Simate-beta system in RoboDojo.
The VPP2 model aims to address a known weakness of world models: if the video prediction is incorrect, the action follows that prediction. Robotera's training process is divided into three stages. The first stage is event-level video pre-training based on the open-source Wan2.1-I2V-14B from Alibaba. This trains the model to predict the entire manipulation process from start to finish, using detailed annotations from robot, human, and general video materials.
Next, the prediction is shortened to fixed eight-second segments and distilled into a single step that takes about 0.12 seconds. The third stage involves using Action DiT with 0.9 billion parameters, which learns to transform predictions into motion, while the video model remains frozen and adapts only through LoRA. The total latency for the action block is approximately 0.22 seconds.
When tested on the real dual-manipulator ALOHA robot, VPP2 showed an average of 58.5% across 10 task types without prior training, including picking, placing, stacking, folding, and pouring. This is higher than the 40% achieved by pi0.5 from Physical Intelligence, and the model was best in nine of these tasks. It scored 45.0% on LIBERO-Pro and 63.9% on LIBERO-OOD. Adding a vision and language model as a high-level planner increased the results of selected RoboDojo long-horizon tasks from 27.6% to 57.6%.
Robotera already uses robots in over 10 logistics centers across five provinces and cities in partnership with companies such as China Post and SF Express, although this does not imply full-scale deployment of VPP2. The model code is available open-source on GitHub.
