Alibaba's Qwen3.8-Max and ZTE's Nebula Share First Place in SuperCLUE Embodied Brain Ranking
Read more
Pandaily
pandaily.com

Alibaba's Qwen3.8-Max and ZTE's Nebula Share First Place in SuperCLUE Embodied Brain Ranking

The Chinese AI evaluation group SuperCLUE has published the September edition of its EmbodiedCLUE-VLA ranking, which assesses the ability of AI models to act as the cognitive core of a robot, planning and reasoning while performing physical tasks.

Among Chinese developments, Alibaba's Qwen3.8-Max-0902 and ZTE's Nebula-EmbodiedBrain secured the first place, achieving identical overall scores of 77.48.

The testing covers four key areas: basic perception, which is divided into temporal, object, and spatial perception; visual reasoning, including mathematical, temporal, spatial, and logical reasoning; interaction and planning, consisting of task planning and trajectory planning; and embodied safety, which checks adherence to ethical norms, physical safety, and privacy protection.

SuperCLUE considers models whose scores differ by one point as a tie and includes international models only for reference outside the main ranking.

How the Leaders Achieved the Ranking

The two leaders achieved the same overall result through different paths. The Qwen3.8-Max-0902 model showed the best results among Chinese models in basic perception (83.33) and visual reasoning (84.72). Meanwhile, Nebula-EmbodiedBrain scored 97.92 in embodied safety, surpassing the Alibaba model, which scored 89.58.

Both models received an identical score of 39.29 in the interaction and planning category, but the ZTE model demonstrated greater strength in task planning (64.29 versus 28.57 for the Alibaba model), while the Alibaba model was stronger in trajectory planning (50 versus 14.29).

According to the report from the Chinese tech publication MyDrivers, the Alibaba model features the ability for spatio-temporal memory, allowing it to remember unfinished work after a robot task interruption. The ZTE version, conversely, is oriented towards on-device deployment and adaptation to the robot's hardware, making it more suitable for operation on physical machines. ZTE previously released EmbodiedBrain 1.0—a vision and language model for embodied task planning, with 7B and 32B weights published on Hugging Face.

Following in the ranking are open-source models with 10 billion parameters or less. Xiaomi's MiMo-Embodied-7B took second place among Chinese models with a score of 56.95 and topped a separate SuperCLUE ranking for models under 10B. It was followed by Alibaba's RynnBrain1.1-9B with a score of 50.33, then BAAI's RoboBrain2.5-8B-NV with 46.36, and Tencent's HY-Embodied-0.5, a 4B model that scored 29.14. Both leaders in this group are closed models available via API.

Planning remains the weakest area among Chinese developers. No Chinese model scored above 40 in the interaction and planning category, and four open entries scored 14.29 or lower. Among comparison models, GPT-6 Astra scored 86.75 overall, and Gemini-3.8-Flash scored 82.12, while Gemini-Robotics-ER-2-Preview, focused on robotics, received 70.86.

MyDrivers noted that as more companies create embodied brains, rankings based on different test sets often yield contradictory results, and the capabilities of these models will largely determine the scope of practical work future robots can perform. SuperCLUE previously published editions of this ranking in January and February 2026.

Similar stories

Alibaba's XekRung Model Takes First Place in CyberGym Ranking with 88.9% Score
Read more
pandaily.com

Alibaba's XekRung Model Takes First Place in CyberGym Ranking with 88.9% Score

Alibaba's cybersecurity large language model, XekRung, has secured the top spot in the CyberGym ranking, which serves as a benchmark for vulnerability reproduction and is supported by researchers from the University of California, Berkeley. This version, designated as XekRung-1.5-27B-Preview and fine-tuned based on Alibaba's open-source model Qwen3.8-27B, demonstrated a success rate of 88.9% as of September 13th. Chinese tech media reported this result on September 28th, noting it is the first time a model from a Chinese developer has led this ranking.

The CyberGym test includes 1507 real vulnerabilities from 188 open-source projects that were initially discovered by Google's OSS-Fuzz program. In the first level of the task, the model receives a vulnerability description and unpatched code, and must then create an input file in the form of a proof-of-concept that triggers the error in the vulnerable version but not in the patched one.

In the current ranking, the XekRung model surpasses Google's Gemini 3.8 Flash Cyber with a score of 86.3%, OpenAI's GPT-5.5-Cyber with 85.6%, Zhipu AI's GLM-5.3 with 84.5%, and DeepSeek-V4-Pro with 83.3%. Supporters of this benchmark note that evaluations are provided by separate teams, runs are stochastic, and small differences in scores may not reflect significant differences in capabilities.

The model was developed in Alibaba's AGI Security Lab under the guidance of Huang Luntao. According to the team's report, the base model Qwen3.8-27B achieved 54.51% under the same conditions, representing an increase of 34.39 percentage points after additional training. The lab explains this success through supervised fine-tuning on data anonymized from Alibaba's own security operations, reinforcement learning of agents within a general instrumental framework, reward collection based on build, failure, and PoC validation, as well as reusing failed attempts as corrective pairs for training. It is emphasized that no CyberGym tasks, patches, or reference PoCs were used during the training process. The evaluation was conducted on locally deployed FP8 inference with a 256K context window, without pre-installed fuzzing frameworks, and with limited network access according to the list permitted in the benchmark.

Alibaba also emphasizes efficiency, stating that the 27-billion parameter model is 1/27th or 1/370th the size compared to comparable cyber models, which will reduce computational costs for automated vulnerability sorting. The weights of XekRung have not yet been published, and Alibaba has not provided a timeline for external access, although Huang stated that the lab's security technology will be offered more broadly, and future versions will target adversarial intelligence, self-evolution, and agent tasks. A technical report on the 8-billion parameter version built on Qwen was previously released in May.

Popular