Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack
Read more
Pandaily
pandaily.com

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack

Xiaomi has made the weights, technical documentation, and reinforcement learning (RL) training resources for its MiMo-V2.6 series publicly available. According to English notes from Xiaomi dated September 22, the MiMo-V2.6-Pro and MiMo-V2.6-Flash repositories were published on the Hugging Face platform along with the RL code and over 7,000 test environments.

This release involves providing weights and code, which differs from previous coverage of the RL training process via live streaming. The branding remains unchanged: Xiaomi / MiMo.

The Pro and Flash models are positioned as native multimodal models with a claimed context window of one million tokens. Model cards and supplementary reports in English indicate that the Pro model features a sparse mixture-of-experts architecture with a total parameter count of approximately 1.02 trillion, actively utilizing about 42 billion parameters. The Flash model is rated at 309 billion total parameters with an activity level of around 15 billion.

Xiaomi reports that each model underwent 30 RL steps across approximately 750,000 trajectories in less than six days. The process utilized task mixing, including coding, general agent work, visual tasks, and cybersecurity. In each update, approximately 1,568 queries and 16 runs were used, with a data volume per step of 3.5–3.7 billion tokens.

According to the company, performance gains in training tasks were observed at approximately 25% for Flash and 12% for Pro. Furthermore, improvements were recorded in DeepSWE v1.1: from 48.8 to approximately 65.7 for Flash and from 58.4 to approximately 72.6 for Pro; however, these figures are vendor-provided data and require third-party verification.

The available assets go beyond mere checkpoints. Xiaomi has provided a comprehensive RL framework built on verl and related agent mechanisms, over 7,000 classified environments covering software development, vulnerability reproduction, intellectual labor, and web design. A Distill-Qwen-9B starting point is also available for community RL experiments, along with lightweight components for multichannel training. The API cost for hosted Pro and Flash versions is stated to be comparable to MiMo-V2.5, while the UltraSpeed tier promises up to 20 times higher throughput than standard Pro while maintaining quality.

Nevertheless, engineers still face high maintenance costs for large MoEs, and Xiaomi's claims regarding artificial intelligence analysis and agent benchmarks require external reproduction. The main news is the complete MiMo-V2.6 package with open weights, RL code, and environments that laboratories can study, rather than just a graphical representation in a live stream.

Similar stories

Xiaomi begins mass production of XRing O3 chip, with O100 and D100 planned for release in 2027
Read more
pandaily.com

Xiaomi begins mass production of XRing O3 chip, with O100 and D100 planned for release in 2027

Xiaomi has announced that its XRing O3 System-on-Chip has entered the mass production and commercial use phase. This is part of a broader roadmap that integrates on-device artificial intelligence with dedicated accelerators and intelligent control chips.

O3 is one of three chips developed by Xiaomi: a flagship mobile SoC already supplied in devices; the AI accelerator O100; and the D100 intelligent driving chip, which is at a later stage of commercialization.

According to company statements and coverage of the design presentation in August and the production update in September, O3 is already being supplied in consumer gadgets. Meanwhile, O100 and D100 have completed development and are scheduled for commercial implementation in the first half of 2027.

Xiaomi views this trio of components as complementary computing tiers: high-efficiency silicon for applications, high-throughput edge AI acceleration, and high-performance AI for vehicles, rather than just a processor for a single seasonal smartphone release.

Regarding the model lineup, Xiaomi's MiMo family now includes variants supporting different languages, multimodality, and speech. The on-device MiMo version is starting to appear in consumer terminals, allowing inferences to be performed on the device's hardware instead of sending every request to the cloud.

This combination is significant for the chip strategy: O3 provides the foundation for today's supplied SoC, while O100 is described as an adjacent memory AI accelerator designed to support larger on-device MiMo workloads across phones, PCs, robots, and other peripheral formats after reaching commercial level.

The D100 chip extends this internal silicon push into the computational power of autonomous transport, supporting large local models that must remain responsive without constant network connectivity.

Xiaomi also disclosed significant multi-year investments in this initiative—about 105.5 billion yuan in core technologies over five years, along with a ten-year plan for chips valued at around 50 billion yuan and over 20 billion yuan already invested. The company reported that devices equipped with the previous XRing O1 chip exceeded one million units sold.

These figures indicate that the mass production of O3 is a productization step following the design launch: it requires going through stages of yield ramp-up, software stack deployment, and supply chain tuning, including support for higher-bandwidth mobile memory, before the self-developed SoC becomes a reproducible series part for multiple SKUs.

Thus, the immediate signal is strategic rather than a smartphone review: XRing O3 is available in SoC supply volume with on-device MiMo functionality, while O100 and D100 are on track for commercial release in the first half of 2027, aiming to expand Xiaomi's internal AI computing capabilities in devices and transportation as the overall silicon program matures from flagship SoCs to specialized peripheral and automotive accelerators.

Popular