deepseek-ai
DeepSelect
DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers
Documentation snapshot
README 快照
翻译暂时拿不到。
机器翻译的项目简介,仅供参考。原文在下方,也可以直接用浏览器自带的整页翻译 (Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」)。
下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。
本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。
DeepSelect
DeepSelect is a high performance implementation of the TopK kernel used in DeepSeek Sparse Attention (DSA) (which is used in DeepSeek V3.2, DeepSeek V4, and DeepSeek V4.1 models) and the sampler. It achieves 2 ~ 20x speedup compared to vanilla torch.topk.
News
- 2026.09.10: We’ve released a brief analysis of the algorithm and its implementation: English | 中文
- 2026.09.10: We’ve released DeepSelect v1.0.0
Supported Cases
TopK workloads vary widely, and the fastest algorithm & implementation highly depends on the input dtype, batch_size, vocab_size, and topk. This repository only focuses on the following cases:
Lightning Indexer Scenario
This scenario covers:
- Input dtype:
torch.bfloat16 batch_size: $1 \sim +\infty$ (both large and small batch sizes are optimized)vocab_size: $1 \sim +\infty$ (both large and small vocabularies are optimized)topk: small (must be $\le 4096$; larger values are not supported)
Recommendations:
- Disable
sorted_indexunless the output has to be ordered by index or by value; enabling either one costs performance. - Set
return_value=Falsewhen the values are not needed. This skips the value output and is faster.
Sampling Scenario
This scenario covers:
- Input dtype:
torch.float32 batch_size: $1 \sim +\infty$vocab_size: around 128Ktopk: small (must be $\le 4096$; larger values are not supported)
Performance
Measured with the benchmark in tests/test.py
(python3 tests/test.py --perf-only), which reports the ratio against torch.topk
on the same input. The metric is effective memory bandwidth: TopK does no
floating-point math, so a FLOP rate would not be meaningful here.
Lightning Indexer Scenario
bfloat16, topk = 512, one subplot per batch size, on a shared 0 - 7 TB/s axis.
图片:DeepSelect vs torch.topk, bfloat16 Lightning Indexer
Sampling Scenario
float32, vocab_size = 129280, topk = 512.
图片:DeepSelect vs torch.topk, float32 Sampling
Installation
git clone https://github.com/deepseek-ai/DeepSelect.git
cd DeepSelect
git submodule update --init --recursive
pip install -v .
Usage
import torch
import deep_select
# input: (batch_size, vocab_size), torch.bfloat16 or torch.float32.
# Its row stride must be a multiple of `deep_select.get_stride_requirement()[0]` bytes, and its last dimension must be contiguous.
batch_size, vocab_size, topk = 4, 204800, 1024
x = torch.randn(batch_size, vocab_size, dtype=torch.bfloat16, device="cuda")
values, indices = deep_select.topk(
x,
topk,
sorted_index=True, # return each row's indices in ascending order
indices_type=torch.int32, # torch.int32 or torch.int64
return_value=True, # False skips the value output (~10% faster)
)
# values: (batch_size, topk) of x.dtype
# indices: (batch_size, topk) of indices_type
The row stride of the input tensor (x) must be aligned to deep_select.get_stride_requirement()[0] bytes. For unaligned inputs, padding is necessary.
Both outputs are allocated by the call, and their strides are aligned to deep_select.get_stride_requirement()[1] bytes (so they may be non-contiguous). Pass output_idx= to write indices into a buffer you own, and that buffer must satisfy the same stride requirement.
For the full signature, see deep_select/interface.py.
Variable-length rows
end sets a per-row upper bound (exclusive). Rows shorter than topk are padded with
value_oob_fill_value / idx_oob_fill_value:
batch_size, vocab_size = 2, 129280 # 129280 is a multiple of 256, so float32 is fine
x = torch.randn(batch_size, vocab_size, dtype=torch.float32, device="cuda")
end = torch.tensor([129280, 100000], dtype=torch.int32, device="cuda") # (batch_size,)
values, indices = deep_select.topk(x, 1000, end=end, sorted=True,
indices_type=torch.int64)
NaN handling
NaN checking is always on. With the default abort_when_nan_found=True the kernel invokes trap() and aborts. Rows whose length is <= topk are never NaN-checked.
Citation
@misc{deepselect2026,
title={DeepSelect: High-Performance TopK Kernels for DeepSeek Sparse Attention and Sampling},
author={Yi Qian and Shengyu Liu and Yichen Li},
year={2026},
publisher = {GitHub},
howpublished = {\url{https://github.com/deepseek-ai/DeepSelect}},
} Official distribution
获取与安装
暂未发现可确认的官方软件包地址
当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。
本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。
Before installing
使用前核验
本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。