DeepSeek Huawei Ascend Tools: What TileLang Changes for AI Developers
DeepSeek and Huawei have open-sourced Ascend 950 programming tools covering compute kernels, distributed communication and TileLang support.
On this page
DeepSeek and Huawei have released a set of open-source programming tools that let developers target Huawei's Ascend AI accelerators without rebuilding every low-level component from scratch. The September 30 release includes DeepGEMM-Ascend, DeepEP-Ascend and new Ascend 950 support in TileLang, turning what was mainly a hardware-alternative story into a programming-stack story. The interesting part is not that these tools make Ascend identical to Nvidia hardware; it is that they give developers more of the software layer needed to write, optimize and move large-model workloads onto it.
DeepSeek's Ascend release moves the compatibility problem into software
AI accelerators are only useful to developers when the software stack can expose their hardware efficiently. Huawei's Ascend processors have their own compiler, libraries and programming interfaces, so code written around Nvidia's CUDA ecosystem cannot simply be moved across and expected to perform the same way. DeepSeek's new repositories address several of the low-level pieces that matter when running modern mixture-of-experts models, including matrix multiplication and communication between accelerator devices.
That distinction is important because a programming platform is more than a driver. A driver allows software to communicate with hardware, while libraries and kernel languages determine how developers express operations such as matrix multiplication, attention and expert routing. DeepSeek is contributing at those higher layers, where much of the performance work for large AI models actually happens. The result is a stack that can reuse more of the programming patterns already developed around DeepSeek's own models rather than starting with an entirely different implementation for Ascend.
DeepGEMM-Ascend ports one of DeepSeek's core kernel libraries
DeepGEMM-Ascend is the clearest example of that approach. DeepGEMM is a kernel library for high-performance matrix multiplication, commonly abbreviated as GEMM, and the new Ascend version is designed to be API-compatible with the original DeepGEMM. The project supports BF16, FP8 and FP4 matrix operations, mixture-of-experts operations and multi-query attention logits on Ascend 950 hardware. That means developers familiar with the upstream library can keep much of the same programming interface while changing the underlying accelerator target.
The implementation is not simply a collection of renamed functions. DeepGEMM-Ascend wraps Ascend's matrix multiply-add primitives and hides details such as its fractal data layouts, alignment requirements and address calculations. It then applies Ascend-specific techniques including sparse data loading and coroutine-based pipelining. The practical goal is to let developers write relatively compact kernels while still reaching close to the hardware's available performance, although the repository's own performance claims are implementation measurements rather than proof that every workload will behave the same way.
DeepEP-Ascend handles the part that happens between accelerators
Matrix multiplication is only one side of large-model execution. When a model is spread across multiple accelerators, those devices also need to exchange data efficiently. DeepEP-Ascend addresses that problem by providing communication primitives for machine-learning training and inference, including expert-parallel all-to-all operations used to distribute and combine work in mixture-of-experts models. It also exposes work-in-progress support for pipeline, context and data parallel communication and remote memory access.
The difference between the two projects is therefore straightforward: DeepGEMM-Ascend focuses primarily on computation inside kernels, while DeepEP-Ascend focuses on communication between pieces of a distributed workload. A large model needs both. Making one operation fast while leaving data movement inefficient can simply move the bottleneck somewhere else, particularly when the model is distributed across many accelerators.
TileLang gives developers a higher-level way to target Ascend
TileLang sits one layer above those individual kernels. It is a domain-specific language with Python-like syntax that lets developers describe high-performance GPU, central processing unit and neural processing unit kernels while a compiler handles much of the translation into hardware-specific instructions. On September 30, the main TileLang project added official Ascend 950 support with native code generation, automatic scheduling and synchronization, plus vector-programming support.
That matters because writing directly against an accelerator's lowest-level instructions can produce excellent performance but requires developers to understand a large amount of hardware-specific detail. TileLang attempts to keep the programming model more portable while still exposing the controls needed for performance tuning. Its existing Ascend work already includes examples for operations such as matrix multiplication and FlashAttention, and the project maintains a separate Ascend adapter built around the same Pythonic programming approach and the Tensor Virtual Machine compiler infrastructure.
The new stack is built around Ascend 950, not every Huawei accelerator
The timing of the release is significant because DeepSeek's new DeepGEMM implementation specifically identifies the Ascend 950 series as its validated target. Its requirements include Huawei's CANN 9.20 toolkit, the torch_npu PyTorch backend, Python 3.10 or newer and a C++20-capable compiler. In other words, the release is immediately useful to developers with compatible Ascend hardware, but it is not a drop-in software package that can be tested on an ordinary Nvidia or consumer graphics card.
The installation requirements also reveal how much of the stack is still hardware-specific. DeepGEMM-Ascend relies on Huawei's Ascend toolkit and PyTorch integration, while the kernels use compiler and runtime components from that environment. That is normal for an accelerator backend, but it means portability has limits even when the public API resembles the upstream project. Developers can reuse application-level logic more easily than they can reuse every performance assumption underneath it.
DeepSeek and Huawei also tested a 128-chip system
DeepSeek said it and Huawei jointly advanced a supernode solution built around 128 Ascend 950 chips, optimizing both computation and communication. A supernode in this context is a tightly connected group of accelerators designed to operate together as one large computing system. The significance of the 128-chip figure is not simply the number of processors; it demonstrates why communication software such as DeepEP-Ascend matters alongside compute kernels such as DeepGEMM-Ascend.
Reuters reported that the collaboration is part of a broader effort by Chinese technology companies to build alternatives to Nvidia's software ecosystem. That context explains why DeepSeek's choice to open-source these components matters beyond its own models. A hardware platform becomes easier to adopt when researchers and developers can inspect, modify and reuse the software required to exploit it. But open source alone does not establish that Ascend will match CUDA's ecosystem size, documentation quality or hardware coverage.
What DeepSeek's CUDA comparison does and does not mean
DeepSeek has described TileLang as a simpler programming model than Nvidia's CUDA and argued that a high-level language is necessary for a more independent accelerator software ecosystem. CUDA is Nvidia's programming platform and development environment for its GPUs, so the comparison is really about developer abstraction rather than claiming that TileLang and CUDA are technically identical. TileLang already supports multiple hardware paths, while its Ascend backend translates the same general programming idea toward Huawei's architecture.
There is a useful distinction here for developers evaluating the release. A common language or similar API can reduce migration work, but it does not eliminate hardware-specific tuning. The DeepGEMM-Ascend source still contains Ascend-specific handling for scaling formats and memory layouts, while DeepEP-Ascend relies on Huawei's HCCL and HCOMM communication stack. Portability therefore exists at selected programming interfaces, not as a promise that one kernel will perform identically across every accelerator.
The more important change is the growing open-source layer around Ascend
DeepSeek's release is more consequential as a collection than as three isolated repositories. DeepGEMM-Ascend addresses compute kernels, DeepEP-Ascend addresses distributed communication, and TileLang provides a higher-level way to develop kernels for Ascend 950. Together they cover several layers that a developer encounters when moving a modern AI workload away from Nvidia hardware. That does not remove the remaining dependency on Huawei's compiler, runtime and accelerator stack, but it reduces the amount of specialized software that each model team has to create independently.
There is already evidence that this approach can be used for real model operators rather than only toy examples. The TileLang-Ascend project documents DeepSeek V4 kernels and provides development material for operations including sparse attention and other model-specific workloads, while Huawei's own Ascend developer material describes integrating TileLang-generated operators into DeepSeek V4 inference. Those projects are still specialized engineering infrastructure, not a universal replacement for established GPU development stacks.
What developers should watch next
The next useful test is whether more model projects adopt these interfaces and whether the same kernels remain competitive as workloads change. DeepGEMM-Ascend currently targets a defined set of operations and Ascend 950 hardware, while DeepEP-Ascend still marks some communication capabilities as work in progress. TileLang's broader multi-backend design gives it a potential path toward portability, but real adoption will depend on compiler quality, debugging tools, documentation, hardware availability and performance across complete applications rather than individual kernels.
For programmers, the practical development is therefore not simply that DeepSeek has released another AI library. It is that more of the code required to run sophisticated models on Huawei accelerators is becoming inspectable and reusable. If that layer keeps expanding, future model ports may require fewer Huawei-specific rewrites and more changes at the backend level. The repositories released on September 30 are an early but concrete step in that direction.
Written by


