On September 30th, DeepSeek released a set of infrastructure components for Huawei's Ascend computing platform. These components include support for the TileLang programming core for Ascend, a library of compute kernels, and a distributed communication library.
DeepSeek stated that each of these components aligns with libraries previously released by the company for NVIDIA GPUs, allowing developers to use the same tools on both platforms. This release goes beyond simple model adaptation, affecting the software underpinning DeepSeek's proprietary training and inference processes.
TileLang in Focus
TileLang is a key element. DeepSeek asserts that this language simplifies coding compared to CUDA while retaining the ability to utilize chip-specific features. Most operators required for training the V4 series have already been implemented. The version for Ascend wraps low-level Ascend C instructions, meaning every TileLang operator used by DeepSeek in training now has a high-performance implementation for Ascend. The company hopes this port will serve as a model for ecosystem software across various AI chips.
For matrix multiplication, DeepGEMM-Ascend has been introduced, which is API-compatible with the original DeepGEMM and supports BF16, FP8, and FP4 formats on Ascend 950 devices. TileKernels provides a unified Python interface for both Huawei's GPUs and NPUs, while FlashMLA handles sparse attention, and DeepSelect provides top-k sampling for attention indexing and sampling. DeepEP-Ascend manages all-to-all inter-process communication for mixture-of-experts models on Ascend 950DT NPUs.
The GitHub page for the Ascend 950DT NPU indicates an FP8 dispatch bandwidth ranging from 373 to 375 GB/s and a combined bandwidth of 345 to 347 GB/s at EP8, which DeepSeek claims approaches hardware limits. These figures were obtained using a test firmware package; DeepSeek directs users to Huawei's commercial release scheduled for around October 15th.
Support from Huawei and Testing Results
Huawei, to whom DeepSeek expressed gratitude for unlimited support, announced that it provided the jointly defined SuperPoD Flex supernode, along with the UBL128 network topology. This topology enables a scalable single-level network for 128 cards with a throughput of 3.2 Tbps and a two-level scaling network up to 256 thousand cards, as well as the ASC-COMM library for custom communication operators. Huawei published a joint paper in its CANN community as recipes detailing low-latency inference with many experts, deployment on a single card and node, large-scale training, long-context KV cache pooling, and reinforcement learning agent training.
In offline tests with an EP32 configuration and a context of 128 thousand tokens, Huawei reported that DeepSeek-V4.1-Flash achieved a speed of 2469 output tokens per second per card at 5 ms per output token, and 5102 tokens per second at 10 ms, excluding framework load balancing and service scheduling. The announcement did not specify whether V4 had completed large-scale training on Ascend.

