Alibaba's cybersecurity large language model, XekRung, has secured the top spot in the CyberGym ranking, which serves as a benchmark for vulnerability reproduction and is supported by researchers from the University of California, Berkeley. This version, designated as XekRung-1.5-27B-Preview and fine-tuned based on Alibaba's open-source model Qwen3.8-27B, demonstrated a success rate of 88.9% as of September 13th. Chinese tech media reported this result on September 28th, noting it is the first time a model from a Chinese developer has led this ranking.
The CyberGym test includes 1507 real vulnerabilities from 188 open-source projects that were initially discovered by Google's OSS-Fuzz program. In the first level of the task, the model receives a vulnerability description and unpatched code, and must then create an input file in the form of a proof-of-concept that triggers the error in the vulnerable version but not in the patched one.
In the current ranking, the XekRung model surpasses Google's Gemini 3.8 Flash Cyber with a score of 86.3%, OpenAI's GPT-5.5-Cyber with 85.6%, Zhipu AI's GLM-5.3 with 84.5%, and DeepSeek-V4-Pro with 83.3%. Supporters of this benchmark note that evaluations are provided by separate teams, runs are stochastic, and small differences in scores may not reflect significant differences in capabilities.
The model was developed in Alibaba's AGI Security Lab under the guidance of Huang Luntao. According to the team's report, the base model Qwen3.8-27B achieved 54.51% under the same conditions, representing an increase of 34.39 percentage points after additional training. The lab explains this success through supervised fine-tuning on data anonymized from Alibaba's own security operations, reinforcement learning of agents within a general instrumental framework, reward collection based on build, failure, and PoC validation, as well as reusing failed attempts as corrective pairs for training. It is emphasized that no CyberGym tasks, patches, or reference PoCs were used during the training process. The evaluation was conducted on locally deployed FP8 inference with a 256K context window, without pre-installed fuzzing frameworks, and with limited network access according to the list permitted in the benchmark.
Alibaba also emphasizes efficiency, stating that the 27-billion parameter model is 1/27th or 1/370th the size compared to comparable cyber models, which will reduce computational costs for automated vulnerability sorting. The weights of XekRung have not yet been published, and Alibaba has not provided a timeline for external access, although Huang stated that the lab's security technology will be offered more broadly, and future versions will target adversarial intelligence, self-evolution, and agent tasks. A technical report on the 8-billion parameter version built on Qwen was previously released in May.

