DeepSeek releases TileLang, DeepGEMM, and DeepEP versions for Huawei's Ascend platform, introducing SuperPoD Flex
Read more
Pandaily
pandaily.com

DeepSeek releases TileLang, DeepGEMM, and DeepEP versions for Huawei's Ascend platform, introducing SuperPoD Flex

On September 30th, DeepSeek released a set of infrastructure components for Huawei's Ascend computing platform. These components include support for the TileLang programming core for Ascend, a library of compute kernels, and a distributed communication library.

DeepSeek stated that each of these components aligns with libraries previously released by the company for NVIDIA GPUs, allowing developers to use the same tools on both platforms. This release goes beyond simple model adaptation, affecting the software underpinning DeepSeek's proprietary training and inference processes.

TileLang in Focus

TileLang is a key element. DeepSeek asserts that this language simplifies coding compared to CUDA while retaining the ability to utilize chip-specific features. Most operators required for training the V4 series have already been implemented. The version for Ascend wraps low-level Ascend C instructions, meaning every TileLang operator used by DeepSeek in training now has a high-performance implementation for Ascend. The company hopes this port will serve as a model for ecosystem software across various AI chips.

For matrix multiplication, DeepGEMM-Ascend has been introduced, which is API-compatible with the original DeepGEMM and supports BF16, FP8, and FP4 formats on Ascend 950 devices. TileKernels provides a unified Python interface for both Huawei's GPUs and NPUs, while FlashMLA handles sparse attention, and DeepSelect provides top-k sampling for attention indexing and sampling. DeepEP-Ascend manages all-to-all inter-process communication for mixture-of-experts models on Ascend 950DT NPUs.

The GitHub page for the Ascend 950DT NPU indicates an FP8 dispatch bandwidth ranging from 373 to 375 GB/s and a combined bandwidth of 345 to 347 GB/s at EP8, which DeepSeek claims approaches hardware limits. These figures were obtained using a test firmware package; DeepSeek directs users to Huawei's commercial release scheduled for around October 15th.

Support from Huawei and Testing Results

Huawei, to whom DeepSeek expressed gratitude for unlimited support, announced that it provided the jointly defined SuperPoD Flex supernode, along with the UBL128 network topology. This topology enables a scalable single-level network for 128 cards with a throughput of 3.2 Tbps and a two-level scaling network up to 256 thousand cards, as well as the ASC-COMM library for custom communication operators. Huawei published a joint paper in its CANN community as recipes detailing low-latency inference with many experts, deployment on a single card and node, large-scale training, long-context KV cache pooling, and reinforcement learning agent training.

In offline tests with an EP32 configuration and a context of 128 thousand tokens, Huawei reported that DeepSeek-V4.1-Flash achieved a speed of 2469 output tokens per second per card at 5 ms per output token, and 5102 tokens per second at 10 ms, excluding framework load balancing and service scheduling. The announcement did not specify whether V4 had completed large-scale training on Ascend.

Similar stories

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack
Read more
pandaily.com

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack

Xiaomi has made the weights, technical documentation, and reinforcement learning (RL) training resources for its MiMo-V2.6 series publicly available. According to English notes from Xiaomi dated September 22, the MiMo-V2.6-Pro and MiMo-V2.6-Flash repositories were published on the Hugging Face platform along with the RL code and over 7,000 test environments.

This release involves providing weights and code, which differs from previous coverage of the RL training process via live streaming. The branding remains unchanged: Xiaomi / MiMo.

The Pro and Flash models are positioned as native multimodal models with a claimed context window of one million tokens. Model cards and supplementary reports in English indicate that the Pro model features a sparse mixture-of-experts architecture with a total parameter count of approximately 1.02 trillion, actively utilizing about 42 billion parameters. The Flash model is rated at 309 billion total parameters with an activity level of around 15 billion.

Xiaomi reports that each model underwent 30 RL steps across approximately 750,000 trajectories in less than six days. The process utilized task mixing, including coding, general agent work, visual tasks, and cybersecurity. In each update, approximately 1,568 queries and 16 runs were used, with a data volume per step of 3.5–3.7 billion tokens.

According to the company, performance gains in training tasks were observed at approximately 25% for Flash and 12% for Pro. Furthermore, improvements were recorded in DeepSWE v1.1: from 48.8 to approximately 65.7 for Flash and from 58.4 to approximately 72.6 for Pro; however, these figures are vendor-provided data and require third-party verification.

The available assets go beyond mere checkpoints. Xiaomi has provided a comprehensive RL framework built on verl and related agent mechanisms, over 7,000 classified environments covering software development, vulnerability reproduction, intellectual labor, and web design. A Distill-Qwen-9B starting point is also available for community RL experiments, along with lightweight components for multichannel training. The API cost for hosted Pro and Flash versions is stated to be comparable to MiMo-V2.5, while the UltraSpeed tier promises up to 20 times higher throughput than standard Pro while maintaining quality.

Nevertheless, engineers still face high maintenance costs for large MoEs, and Xiaomi's claims regarding artificial intelligence analysis and agent benchmarks require external reproduction. The main news is the complete MiMo-V2.6 package with open weights, RL code, and environments that laboratories can study, rather than just a graphical representation in a live stream.

Popular