The AgiBot WITA-Omni Full-Modal model achieved a high score of 85.21 in the DailyOmni benchmark. Furthermore, the model demonstrated supremacy in six out of eight metrics, surpassing systems such as Alibaba Qwen, Google Gemini, ByteDance Doubao, and NVIDIA.
Testing and Model Capabilities
The benchmark assesses the model's ability to simultaneously process audio and visual signals in real-world conditions. Among the tested parameters were audiovisual data consistency (86.13), temporal reasoning about events (84.64), complex cross-modal reasoning (85.06), as well as high results in tasks involving 30 and 60-second video comprehension.
WITA-Omni Architectural Innovation
Unlike general-purpose multimodal models optimized for interacting with digital content, WITA-Omni was originally developed for embodied interaction. This means that robots utilizing this model must understand, react to, and act within a physical environment in real time.
The innovation lies in the Thinker-Talker-Actor paradigm, which replaces the traditional sequential pipeline. Standard robot interaction processes perform speech recognition, semantic understanding, response generation, speech synthesis, and action output as separate, sequential stages, leading to accumulated latency. WITA-Omni, however, runs the Thinker multimodal reasoning core, the Talker speech streaming generation, and the Actor action and expression modules in parallel on a single timeline.
Synchronization Mechanisms
The Actor module takes as input not only the semantic state from the Thinker but also latent audio variables from the Talker. This allows gestures and facial expressions to synchronize with the rhythm of speech, pauses, accents, and emotional tone, eliminating discrepancies between what the robot says and how it moves. The architecture extends the Thinker-Talker paradigm, previously created for conversational AI, by adding the Actor module as a full output channel alongside speech.
The Thinker integrates textual, visual, audio, and video inputs through specialized encoders and lightweight projection layers into a unified representation space for decision-making regarding interaction and cross-modal reasoning. The Talker generates natural speech conditioned on the hidden states of the Thinker. The Actor controls the action head and expression head, using both the semantic state of the Thinker and the audio features of the Talker. The synchronization mechanism between the streams ensures that the robot's movements react to changes in speech speed, shifts in expression, and that physical actions maintain temporal coherence with verbal communication.
Industry Positioning
This result follows the achievement of the AgiBot world model Genie Envisioner-Sim 2.0, which previously took first place at WorldArena. These two achievements form a three-pillar structure of AI capabilities: WITA-Omni for interaction, the world model for understanding the physical world, and the robot hardware for acting within it. The victory in the ranking confirms that companies focused on embodied AI can achieve superior overall capabilities by optimizing for real-world constraints rather than adapting general-purpose models.
As robots transition from controlled factory environments to unstructured social spaces, the ability to synchronously process multiple sensory streams and respond with coordinated verbal and physical actions becomes a key differentiator. AgiBot defines synchronous multimodal interaction as a new competitive dimension in the embodied AI industry, separate from competition based purely on model quality dominating general AI.