XPeng has unveiled the most significant update to its second-generation Vision-Language-Action (VLA) model, with the key change being the artificial intelligence's ability to understand temporal aspects. At a presentation focused on physical AI under the theme of 'TIME,' the company announced that its base model has evolved from static three-dimensional spatial perception to dynamic four-dimensional spacetime understanding with the release of new software XOS 6.3.0 on the G9L vehicle.
This breakthrough is based on three technologies. Firstly, Infini-VLA, a large-horizon architecture, can remember a road scene for up to 30 seconds, integrating continuous driving events into a unified temporal understanding. Secondly, Streaming Inference technology transforms the model's operation from discrete reasoning to continuous parallel processing, reducing end-to-end latency by 300%. Thirdly, the X-Foresight world prediction model anticipates the behavior of surrounding road participants up to six seconds in advance, thereby improving reaction to sudden lane changes and abrupt braking.
On-device parameters have increased by 3.5 times—more than 15 times more than in core VLA models—and comprehensive multidimensional safety has improved by 20 times. Furthermore, XPeng introduced the Master Agent—a universal 'brain' not tied to a specific vehicle and built on the multimodal Omni model. This agent integrates VLA perception with the VLM cabin, automatically discerning user intent and coordinating autonomous driving, chassis, and cabin systems.
New voice-activated functions, such as requesting a stop, effectively bring L4-level capabilities inherent to robotaxis into mass production. This update comes as XPeng intensifies its focus on physical AI. Recently, its robotics business closed its first private funding round exceeding $900 million at a valuation above $6.3 billion—a record for a single round in China's embodied intelligence sector. Investors included IDG, Gaorong, Tencent, and Alibaba.
For XPeng, the ultimate goal is to create a unified physical world model that treats driving, robots, and cabin intelligence as one continuous reasoning task that accounts for time, rather than as separate isolated tasks. By embedding time into the model, the company is betting that the next stage of autonomy will resemble less simple perception and more understanding of the world as it develops.
