跳到正文

deepseek-ai

DeepGEMM-Ascend

DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

DeepGEMM Ascend

DeepGEMM Ascend is a port of DeepGEMM to the HUAWEI Ascend platform. It is fully API-compatible with DeepGEMM and supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE. On Ascend platforms, users can simply install the package and use the same APIs and development workflow as DeepGEMM on other supported platforms.

DeepGEMM Ascend provides a lightweight abstraction over the Ascend MAD (matrix multiply-add) primitives, hiding much of the complexity associated with fractal layouts, alignment constraints, address calculations, and verbose low-level parameters. This enables GEMM kernels to remain both concise and efficient. DeepGEMM Ascend makes extensive use of Ascend-specific optimization techniques, such as sparse data loading and coroutine-based pipelining, to approach the performance limits of the Ascend hardware. These implementations can also serve as references for extreme performance optimization on the Ascend platform.

Despite its lightweight codebase, DeepGEMM Ascend can achieve peak hardware performance across a wide range of matrix shapes.

DeepGEMM Ascend 是华为昇腾平台上的 DeepGEMM 实现。它完全兼容 DeepGEMM 的 API,支持 BF16、FP8、FP4 GEMM、MQA logits 和 MegaMoE 算子。DeepGEMM Ascend 和 DeepGEMM 使用相同的包名,用户只需安装当前包,即可沿用其他平台上 DeepGEMM 的 API 和开发流程。

DeepGEMM Ascend 对昇腾平台的矩阵乘加原语(MAD)提供了一个轻量的抽象,能够隐藏分形矩阵布局、对齐约束、地址计算、参数转换等细节,让 GEMM kernel 的实现保持简洁高效。DeepGEMM Ascend 广泛采用昇腾平台特有的优化技术,包括稀疏数据加载、基于协程的流水线等,以接近硬件的性能极限。这些实现也可以作为昇腾平台极致性能优化的参考。

DeepGEMM Ascend 代码轻量,并且在多种矩阵形状上能达到硬件极限性能。

News

  • 2026.09.30: Initial release of DeepGEMM Ascend, with support for Ascend 950 devices. The kernels are designed to achieve near-peak hardware performance.

Quick Start

Requirements

  • HUAWEI Ascend NPU (developed and validated on the Ascend 950 series)
  • CANN 9.20 toolkit providing bin/bisheng, bin/ld.lld
  • Torch NPU package torch_npu
  • Python 3.10 or higher
  • Compilers and standard libraries with C++20 “ support
  • tilelang, used by the HC prenorm kernel (declared as a package dependency)
  • tree-sitter and tree-sitter-cpp, used to generate Python type stubs when building from source

Development

# Submodule must be cloned
git clone --recursive https://github.com/deepseek-ai/DeepGEMM-Ascend.git
cd DeepGEMM-Ascend

# Link some essential includes and build the C++ extension
cat develop.sh
./develop.sh

Installation

pip install . --no-build-isolation

Interfaces

Kernel Interface

Please refer to DeepGEMM’s interfaces for the kernel APIs.

[!NOTE] The scaling factor format on Ascend differs from NVIDIA’s: each pair of UE8M0 scaling factors along the K dimension is packed into an int16, and the packed values are stored in MN-major order for optimal hardware efficiency.

Utilities

The library provides some utility functions besides the above kernels:

  • deep_gemm.set_num_sms / get_num_sms: set/get the number of AI cores the kernels may use
  • deep_gemm.set_npu_arch / get_npu_arch: override/query the dav-* NPU architecture used for JIT compilation, 0 queries the device
  • deep_gemm.set_mk_alignment_for_contiguous_layout / get_mk_alignment_for_contiguous_layout: set/get the group-level M/K alignment for contiguous layout
  • deep_gemm.get_theoretical_mk_alignment_for_contiguous_layout: get the theoretical minimum M/K alignment
  • deep_gemm.use_deterministic_algorithms / get_deterministic_algorithms: enable/disable deterministic algorithms
  • deep_gemm.transform_sf_into_required_layout: transform scaling factors into the required layout
  • deep_gemm.transform_k_grouped_sf_into_required_layout: transform K-grouped scaling factors into the required layout
  • deep_gemm.get_paged_mqa_logits_metadata: build the scheduling metadata for the paged MQA kernels
  • deep_gemm.aclnn_fp8_fp4_gemm_{nt, nn, tn, tt} and deep_gemm.aclnn_bf16_gemm_{nt, nn, tn, tt}: ACLNN reference GEMMs used to cross-check the kernels in tests

Environment Variables

Ascend home path is given by ASCEND_HOME_PATH or ASCEND_TOOLKIT_HOME.

Each DG_JIT_* variable falls back to the corresponding global DJ_JIT_* variable when unset, and is snapshotted when the JIT runtime is first created.

  • General
    • DG_JIT_DEBUG: 0 or 1, enable JIT debugging features, including compiler commands, load-time reporting, and the selected config per shape; 0 by default
    • DG_PRINT_CONFIGS: 0 or 1, print the selected config for each shape, 0 by default
  • JIT cache
    • DG_JIT_CACHE_DIR: string, cache directory (or a :-separated list of directories) for compiled kernels; lookup searches all paths front-to-back (first hit wins) and a cache miss compiles into the first path, $HOME/.dj by default
  • Compiler output and artifacts
    • DG_JIT_PRINT_COMPILER_COMMAND: 0 or 1, print compiler and disassembler commands, 0 by default
    • DG_JIT_PRINT_LOAD_TIME: 0 or 1, print kernel load time, 0 by default
    • DG_JIT_KERNEL_DEBUG_INFO: 0 or 1, add Bisheng kernel debug information, 0 by default
  • Debugging
    • DG_JIT_LAUNCH_TIMEOUT: integer, Ascend kernel launch timeout in seconds; 300 by default, 0 disables it
    • ASCEND_LAUNCH_BLOCKING: synchronize the device after every launch, turning asynchronous launch errors into in-line failures

Performance

Measured on Ascend 950DT (CANN 9.20) using bench_msprof with cold L2. Shapes and cases follow the DeepGEMM test suite and cover inference and training workloads of the DeepSeek model series; see tests/ for reproduction.

Dense GEMM

Dense GEMM reaches up to 99.8% of the hardware limit across types:

TypeMNKLatency (us)Computation (TFLOPS)Hardware limit (TFLOPS)Utilization (%)
FP4×FP44096716816384565.61701173098.3
FP8×FP440967168163841117.886186599.5
FP8×FP840967168163841117.686186599.5
BF16×BF1640967168163842229.943143299.8

FP8 GEMM for inference:

MNKLatency (us)Computation (TFLOPS)Memory bandwidth (GB/s)
1282112716815.02581205
128576716810.1105564
12824576153620.64682068
1283276851210.44151829
12871681638454.85492456
1284096716819.03971797
1287168204810.13731670
409621127168156.9790319
4096576716854.6619689
4096245761536362.2854137
409632768512164.3837129
40967168163841117.6861186
409640967168282.9850234
409671682048143.4838181

M-Grouped GEMM

Grouped GEMM for DeepSeek MoE experts (FP8×FP4, BF16 output). #Groups is the number of experts, and M per group the average tokens each receives.

#GroupsM per groupNKLatency (us)Computation (TFLOPS)Memory bandwidth (GB/s)
48192614471683722.4818104
48192716830721778.085698
48192409640961356.1855148
4819240962048680.6852148
84096614471683689.4831136
84096716830721778.9862130
84096409640961356.5861180
8409640962048681.0858179

MQA Logits

MQA scoring for the DeepSeek Lightning Indexer. The kernel is FIX-pipe bound rather than compute bound, saturating the FIX pipe at 99% utilization.

TypeFormat#Q Tokens#K Tokens#HeadsDimLatency (us)Computation (TFLOPS)Memory bandwidth (GB/s)
PrefillFP84096819232128302.6681283
PrefillFP44096819232128277.4743277
DecodeFP8256819232128150.95652006
DecodeFP4256819232128124.27081349

MegaMoE

Mega MoE fuses EP dispatch, two grouped GEMMs, SwiGLU, and combine. Benchmarked over EP8 with top-k=6 (each token routed to 6 experts) and one shared expert; all values are averaged across 8 ranks.

#ExpertsHiddenIntermediateTokensLatency (us)Computation (TFLOPS)Memory bandwidth (GB/s)Communication bandwidth (GB/s)
965120230464138.7228.51850.037.5
9651202304256225.5562.51257.491.6
967168307264222.2266.42136.932.8
9671683072256359.2659.21425.480.5
3845120230440962567.7790.3567.5128.8
38451202304163849768.6831.0325.0135.2
3847168307240964600.7823.4531.3100.6
384716830721638417904.2846.3269.3103.3

HC Prenorm GEMM

Prenorm GEMM for the DeepSeek mHC (Manifold-Constrained Hyper-Connections) module, which nearly saturates the HBM write bandwidth.

MNKLatency (us)Computation (TFLOPS)Memory bandwidth (GB/s)
1324286726.73519
13724286729.6201110
512242867215.4462082
4096242867272.7783276
81922428672136.7823463

Contributors

  • Project Leads: Kexing Zhou, Zhean Xu, Chenggang Zhao
  • GEMM Kernels: Kexing Zhou, Zhean Xu, Yunfan Xiao, Yuhao Meng
  • MQA Logits: Anyi Xu, Zhean Xu, Kaifeng Chen
  • mHC Kernel: Yuxuan Zhou, Chenggang Zhao, Ruifan Xu, Huanqi Cao, Chenhao Xu
  • MegaMoE: Zhean Xu
  • SF Layout Kernels: Guanglin Li
  • Infrastructure: Kexing Zhou, Kuai Yu

Acknowledgements

DeepGEMM-Ascend follows the design of the upstream DeepGEMM project, which is inspired by CUTLASS. The JIT runtime is provided by DeepJIT. The mHC kernel is backed by Tilelang. We sincerely thank the developers of these projects for their contributions, and gratefully acknowledge Huawei for its technical support and engineering expertise throughout the development of DeepGEMM-Ascend.

License

This code repository is released under the MIT License.

Citation

If you use DeepGEMM-Ascend in your work, please cite:

@misc{deepgemm_ascend2026,
  title     = {DeepGEMM-Ascend: Clean and Efficient BLAS Kernel Library on Ascend NPU},
  author    = {Kexing Zhou and Zhean Xu and Anyi Xu and Chenggang Zhao and
               Yuxuan Zhou and Yunfan Xiao and Guanglin Li and Kaifeng Chen and Yuhao Meng and
               Huanqi Cao and Ruifan Xu and Chenhao Xu and Kuai Yu},
  year      = {2026},
  publisher = {GitHub},
  url       = {https://github.com/deepseek-ai/DeepGEMM-Ascend}
}

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。