跳到正文

tonyd2wild

DeepSeek-V4.1-Flash-vLLM-DGX-Spark

DeepSeek-V4.1-Flash (552B MoE, MXFP4 experts, 1M ctx) on four NVIDIA DGX Sparks with vLLM TP4: Engram-on-disk patch, sm121 kernel build, launchers, measured numbers

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

DeepSeek-V4.1-Flash on four NVIDIA DGX Sparks (vLLM, TP4, DSpark, CUDA graphs)

Status (2026-09-10): serving (boot 10).

  • The model dropped at about 2 AM ET, and this stack started serving it at 9:17 AM ET the same day.
  • One stream, decode tok/s (after the first token): code 73.8, tables 55.4, JSON 52.1, math 50.9, reasoning 37.8, narrative 24.9, prose 24.4.
    • Counting ran at 92.2 tok/s back to back. It was 62.2 in the bench run, which hit a GPU slow-state phase (issue #1).
  • Concurrency: six streams give 131.9 tok/s aggregate across the 8 categories (boot 9: 98.0).
    • Peak aggregates: counting 259.9 (C5), code 225.5 (C6), tables 215.1 (C5), math 182.7 (C6).
  • Prefill: 902-1,539 tok/s cold. A 93K-token prompt takes 78 s. TTFT on short prompts is 0.3-0.5 s.
  • Context: 300K max context with a 1,070,168-token KV pool (3.57x at 300K).
  • Vision and tool calling are on: up to 4 images per request, tool calls including parallel calls and a full round trip. 7/7 end-to-end checks pass (below).
    • 1M max context was proven on a separate boot, with a 1,078,380-token DSpark KV pool.
  • Nothing on this page is a projection.

Vision and tool calling (on in the serving config)

Both are live on boot 10 and tested end to end with tools/vision_tools_demo.py; the output is in results/boot10/vision-tools.txt. The test images are generated in the script, so each expected answer is known exactly.

testresult
V1: one image, three color stripes, name them left to rightPASS: “red, green, blue” (0.6 s)
V2: two images in one messagePASS: first=red, second=blue (1.2 s)
V3: 2x2 grid, color of the top-right squarePASS: “Green” (0.6 s)
T1: tool call with argumentsPASS: get_weather {"city": "Paris", "unit": "c"} (1.4 s)
T2: full round trip (the tool result goes back, the model answers from it)PASS: “18°C with light rain” (2.2 s)
T3: parallel calls in one turnPASS: get_weather for Tokyo and Berlin (2.9 s)
T4: tool_choice forcing a named functionPASS: get_time {"city": "Sydney"} (1.6 s)

How it is switched on (launch/boot10-go.sh):

  • TEXT_ONLY=0: the vision encoder loads (no --language-model-only). It adds about 0.22 GiB per rank.
  • --limit-mm-per-prompt {"image":4} --mm-processor-cache-gb 1: up to 4 images per request.
  • PARSERS=1: --tool-call-parser deepseek_v41 --enable-auto-tool-choice --reasoning-parser deepseek_v41.
  • Thinking is off by default; turn it on per request with "chat_template_kwargs": {"thinking": true}.
  • Watch the defaults. launch/dsv41-tp4.sh on its own is text-only with tools off (TEXT_ONLY=1, PARSERS=0). Set both as boot10-go.sh does.

Example requests (OpenAI-compatible API; the served model name is deepseek-v4.1-flash):

# image (base64 data URL; up to 4 images per message)
curl -s http://:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "deepseek-v4.1-flash", "max_tokens": 200,
  "messages": [{"role": "user", "content": [
    {"type": "text", "text": "What is in this image?"},
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,"}}]}]}'
# tool call
curl -s http://:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "deepseek-v4.1-flash", "tool_choice": "auto",
  "messages": [{"role": "user", "content": "What is the weather in Paris in celsius?"}],
  "tools": [{"type": "function", "function": {"name": "get_weather", "description": "Get the current weather for a city",
    "parameters": {"type": "object", "properties": {"city": {"type": "string"}, "unit": {"type": "string", "enum": ["c", "f"]}},
    "required": ["city"]}}}]}'

Benchmark (boot 10, the serving config)

4x DGX Spark TP4, DSpark k=5, FULL_AND_PIECEWISE CUDA graphs, Engram rows node-local on every rank, tools and vision on, 300K context, gmu 0.80.

How it was measured:

  • Fixed prompt set bench/prompts-v1.json, identical on every boot: 8 categories plus a counting ceiling.
  • Streaming, temperature 0, thinking off, after a warmup. Short prompts (about 30-120 tokens), 150-256 token budgets.
  • At concurrency C, C streams are released together. There is one batch per cell.
  • Token counts come from the server’s usage block, never from stream chunks (DSpark packs several tokens per chunk).
  • Decode = tokens after the first / time after the first token, per stream. Aggregate = all streams’ tokens / batch wall time. TTFT = first token delta.
  • Raw output and more tables: results/boot10/ (report, made with tools/bench_report.py).
  • One caveat: the GB10 GPU slow state (issue #1) can move a single cell by up to about 1.5x. Counting at C1 is the clearest case (62.2 here, 92.2 back to back).

Throughput by concurrency (mean of the 8 categories; the counting ceiling is excluded):

Caggregate tok/sper-stream decode tok/smean TTFT (s)boot 9 aggregate
C137.9543.120.44133.88
C264.3037.430.44158.99
C378.7030.620.79367.56
C485.7224.660.57985.50
C5114.2027.010.58591.96
C6131.8625.350.50297.99

Decode: per-stream tok/s after the first token

categoryC1C2C3C4C5C6
code73.845.033.834.344.141.5
JSON52.130.420.722.726.827.5
math50.959.354.431.428.934.3
reasoning37.842.436.226.820.824.6
tables (format)55.459.541.535.553.135.3
summary25.722.422.413.715.916.1
prose24.423.619.418.414.813.8
narrative24.916.916.514.711.89.8
counting (ceiling)62.258.943.058.357.233.4

Aggregate throughput: tok/s across all streams (wall time, TTFT included)

categoryC1C2C3C4C5C6
code66.581.192.4123.8196.7225.5
JSON43.749.251.976.394.7130.4
math45.6105.5148.4110.2127.9182.7
reasoning35.075.396.597.094.1133.6
tables (format)44.094.3101.7112.0215.1182.1
summary21.937.638.243.164.073.5
prose22.941.555.269.967.974.8
narrative23.830.045.253.453.152.4
counting (ceiling)57.3109.0118.6208.8259.9184.9

TTFT: mean time to first token (s)

categoryC1C2C3C4C5C6
code0.310.440.490.510.400.42
JSON0.330.500.550.570.450.42
math0.470.350.380.460.620.42
reasoning0.440.350.420.550.570.41
tables (format)0.510.400.570.590.470.47
summary (~300-token passage)0.790.773.181.281.331.02
prose0.380.390.310.320.360.48
narrative0.290.320.440.350.490.37
counting (ceiling)0.340.350.460.400.350.53

Prefill: cold, unique prompt, 1-token reply (TTFT here is the whole prefill)

targetprompt tokensTTFT (s)prefill tok/s
2K2,9503.27902
8K11,59211.291,026
32K46,81030.431,539
64K93,33578.171,194

Decode per stream, boot 9 → boot 10 (same prompts; boot 10 adds node-local Engram rows):

categoryC1C2C3C4C5C6
code52.4 → 73.860.4 → 45.033.4 → 33.831.3 → 34.328.5 → 44.126.3 → 41.5
JSON38.6 → 52.137.1 → 30.420.0 → 20.720.8 → 22.718.8 → 26.816.2 → 27.5
math46.9 → 50.946.2 → 59.346.2 → 54.449.1 → 31.427.6 → 28.925.8 → 34.3
reasoning39.0 → 37.827.2 → 42.429.1 → 36.220.9 → 26.825.4 → 20.817.1 → 24.6
tables (format)71.4 → 55.450.2 → 59.538.9 → 41.535.4 → 35.536.3 → 53.130.8 → 35.3
summary24.7 → 25.717.4 → 22.413.6 → 22.411.5 → 13.711.4 → 15.910.9 → 16.1
prose23.3 → 24.416.1 → 23.616.1 → 19.417.6 → 18.412.0 → 14.814.0 → 13.8
narrative17.4 → 24.918.7 → 16.911.2 → 16.511.6 → 14.710.2 → 11.89.8 → 9.8
counting (ceiling)77.2 → 62.256.0 → 58.939.3 → 43.040.2 → 58.336.5 → 57.238.1 → 33.4

Step time, fast vs slow GPU state (streamed; every speculative step is timed from the stream; results/boot10/idletest.txt):

requestms/steptokens/stepdecode tok/s
count 1-100, first request after about 13 min idle94 (the whole request)5.8561.7
count, back to back635.8592.2
code, back to back70.54.9770.2
count, after 45 s idle635.8592.1
code, back to back675.3077.3

DSpark acceptance (vLLM SpecDecoding metrics across the boot 10 bench, 35 ten-second windows): mean acceptance length 3.57 tokens per step, range 1.88-5.92.

  • Counting, code and tables sit near the maximum of 6.
  • Prose and narrative stay near 2, which is why per-stream speed spans 10-92 tok/s by content.

What each fix bought (counting prompt, one stream, measured):

stacktok/s
eager, no speculation (boot 6)5.1
eager + DSpark k=5 (boot 7)19.5-22.1
DSpark + CUDA graphs + Engram rows staged before the forward (boot 8)41.5 (count-to-100 check)
+ GPU clock latch cleared on two nodes (boot 9)60.8 (count-to-100 check); 77.2 bench ceiling
+ node-local Engram rows, GPU clocks locked (boot 10)84.9 (count-to-100 check, same method); 92.2 decode back to back

Memory (boot 10):

value
Weights per rank, with the DSpark draft layers and the vision encoder81.58 GiB
CUDA graphs1.99 GiB target + 0.55 GiB draft
KV5.15 GiB = 1,070,168 tokens (3.57x at 300K)
Node-local Engram rows per worker48 GB on disk (sparse copy of the rank’s rows)

Earlier boots:

  • Boot 9: C1 per-stream 39.2, code 52.4 (results/boot9/).
  • Boot 8, before the GPU clock fix: C1 per-stream 24.2, code 32.9 (results/boot8/).
  • 1M max context (boot 7): served at --max-model-len 1048576 with DSpark: KV 1,078,380 tokens (5.21 GiB, 1.03x at 1M), eager, gmu 0.80.

The model and the problem

deepseek-ai/DeepSeek-V4.1-Flash:

  • 552B-backbone MoE (769B counting the Engram tables), 16B active decode / 8B prefill, 1M context.
  • MXFP4 experts, MXFP8 dense, 510 GB on disk.

It does not fit four GB10s as shipped. The 296 GB of experts split four ways is fine. The problem is the two Engram n-gram tables (203 GB FP8): vLLM keeps them in host memory, and on a DGX Spark host memory is the GPU pool.

This repo:

  • keeps the Engram tables on disk;
  • fixes what that and the GB10’s SM 12.1 break in day-0 vLLM;
  • gets CUDA graphs working around a host-side lookup.

Full recipe: docs/RECIPE.md

What had to be fixed

In boot order. Details in docs/.

  1. Engram on disk (patch/engram.py, patch/weight_utils.py): the two 101 GB tables stay in the safetensors files, and rows are read on demand. Without this it does not fit.
  2. Runtime JIT wedge (boot 3). A FlashInfer MXFP8 GEMM compiled at runtime with 22 parallel jobs and exhausted host memory on all four nodes, and the watchdogs reset them. Fix: kernels prebuilt in the image (build/build_overlay5.sh) plus MAX_JOBS=2. docs
  3. Engram rank offset (found in audit). The disk reader ignored each rank’s row offset, so ranks 1-3 read rank 0’s rows, silently. docs
  4. No common block size for 64 (boot 4). vLLM picked the smallest listed block size, and the V4 indexer backend refused it. Fix: --block-size 128. docs
  5. DeepGEMM block_kv == 32 or 64 (boot 5). The ratio-1 indexer cache had 128 states per block. Fix: SM12x indexer pages of 64 states. docs
  6. Relaunch race. A new worker joined the still-live old head’s rendezvous on the same port. Fix: tools/launch.sh stops every node first.
  7. persistent_topk on GB10 (boot 7, long context). It oversubscribes the 48 SMs, and its fallback needs 128 KB of shared memory per block (GB10 has 99 KB). Fix: top_k_per_row_decode, which is also 1.6-3.6x faster there. results
  8. Eager mode was the throughput ceiling. About 200 ms per step, host-bound; the GPUs sat near idle. Fix: the Engram lookup moves out of the forward into prepare_inputs (patch/model_state.py), with all rows read in parallel. The whole decode step is then captured as a CUDA graph, with exact capture sizes so DSpark batches are never padded (FlashInfer #5015).
  9. GPU clock latch (2 of 4 Sparks). Reddie and Asusi sat at 630-950 MHz with no visible cause. Every TP step waited for them. Fix: unplug the adapter for 30-60 s; a reboot does not clear it. Result: count 41.5 to 60.8 tok/s, code 32.9 to 57.1. docs
  10. Engram rows over NFS (boot 10).
    • Problem: the workers read their Engram rows from the head’s NFS export. That took 5.9-7.8 ms per step against 2.8 ms on the head, and every step waited for the slowest rank.
    • Fix: tools/engram_local.py copies each worker’s rows to local NVMe (47 GiB per worker, a sparse copy at the original offsets, verified row by row).
    • patch/engram.py reads the copy only when its recorded row range covers the rank, and falls back to NFS otherwise. launch/dsv41-tp4.sh mounts it with ENGRAM_LOCAL=1.
    • The same boot fixed boot_dsv41.sh, which was not forwarding the patch folder to each node (now PATCH_NAME).

Boot log

bootchangeoutcome
1overlay1, eager, text-only, Engram on diskKV 1,989,514. Died in decode warmup: FlashInfer 0.6.18 has no SM120 sparse-MLA decode kernel for page_block_size=32.
2overlay2 + SWA overrideDied at KV init: No common block size for 32.
3overlay3 (FlashInfer 0.7.0rc1) + SM12x page patchesWedged all 4 nodes (runtime JIT compile exhausted host memory).
4overlay5 (kernels prebuilt)No runtime JIT; KV 2,026,695. Died at KV init: No common block size for 64.
5+ --block-size 128, Engram offset fix, 300KKV 2,346,690. Died in decode warmup: DeepGEMM block_kv == 32 or block_kv == 64.
6+ indexer pages of 64 statesServed. Eager, no speculation: 5.1 tok/s (count). Smoke clean, greedy reference 8/8, garble gate 30/30.
7+ DSpark k=5 at 1M max contextServed at 1M. KV 1,078,380. DSpark eager 19.5-22.1 tok/s. A 32K request then killed it in persistent_topk (fix 7).
8+ CUDA graphs, Engram staged before the forward, top-k fix, gmu 0.78, 300KServed. Count check 41.5 tok/s, code 32.9 (two GPUs clock-latched, found later).
9+ tools, vision, gmu 0.80; Reddie and Asusi power-cycled to clear a GPU clock latchServed. KV 1,032,963 (3.44x at 300K). Count-to-100 60.8 tok/s, code 57.1, count-to-300 68.2. Tool call and image OK.
10+ node-local Engram rows on the 3 workers (one change); GPU clocks locked by the owner before launchServing. KV 1,070,168 (3.57x at 300K). Count-to-100 84.9 tok/s, code 65.3 (end to end). Bench: C1 code 73.8, C6 aggregate 131.9. Tool call and image OK. Restored 2026-09-11 after a worker was powered off by accident (tools/restore_boot10.sh, 11 min to serving). Same config; KV 1,250,201 this time (graph capture took 1.29 GiB vs 1.99). Warm: count 90-92 tok/s, code 72-73.

Known limits and next steps

  • GB10 GPU slow state (issue #1; help wanted).
    • The GPU switches between a fast and a slow state that nvidia-smi does not show. A decode-shaped GEMV runs at 70 vs 230 GB/s.
    • In serving that is 63 vs 94 ms per step. A long idle reliably leaves it slow; it also flips during work, less often.
    • It reproduces with a plain PyTorch script (tools/gpuflip.py), so it is not the engine. Network, thermals, CPU placement, ASPM and KV caching are ruled out.
  • Vision and tools.
    • Both on: up to 4 images per request, with the deepseek_v41 tool and reasoning parsers.
    • Thinking is off by default; a request can turn it on with "chat_template_kwargs": {"thinking": true}.
    • FlashInfer #4973 (vision on SM120) did not reproduce in the 3 image tests, one of which sends two images in one message. Heavier image traffic and large photos are untested.
  • Adaptive verification is off. It pads speculative batches, and padded batches can hang SM120 sparse MLA (FlashInfer #5015, open).
  • Step time in the fast state: 63 ms on counting and 67-71 ms on code.
    • Measured parts: 88 all-reduces per step cost about 5 ms (tools/nccl_lat.py, stable across the four Sparks). The Engram staging still runs before each forward (2-3 ms with local rows).
    • The rest is the MoE and attention kernels, the next lever.
  • 1M long context after the top-k fix. The top-k fix is validated on GB10 at row widths up to 300,000. A 1M-token request has not been re-run on the fixed stack.
  • Host hardening (system settings, not applied here): see docs/RECIPE.md step 7.

Repo layout

pathwhat
patch/The exact files bind-mounted over vLLM (md5s in patch/README.md), plus one folder per fix with its diff and test.
build/Image chain: overlay1 (branch + sm121 extension), overlay3 (FlashInfer 0.7.0rc1), overlay4/5 (prebuilt kernels).
launch/dsv41-tp4.sh , boot_dsv41.sh (worker-first fan-out), and one bootN-go.sh per boot. boot10-go.sh is the serving config.
tools/Launch wrapper, pre-launch steps, boot poll, post-serve checks, bench report, engram_local.py (node-local Engram rows), gpuflip.py / flipsum.py (GPU slow-state probe), nccl_lat.py (4-node all-reduce check), idletest.py (per-step timing after idle vs back to back), vision_tools_demo.py (vision and tool-calling checks).
bench/Fixed prompt set v1, C1-C6 bench, long-context needle test.
docs/Recipe and one post-mortem per failure.
results/Head logs of every boot, proofs, bench output, pre-launch GPU and network checks.

Fleet:

  • Reddie is the head: model on local NVMe, exported over NFS.
  • Asusi, Bluey and Spark4 are workers, each with a local copy of its Engram rows.
  • ConnectX-7 RoCE fabric, 192.168.192.0/24.

Sister repos: DeepSeek-V4-Flash-Vision-Exp (vLLM, 2x/4x Spark) · DeepSeek-V4-Flash-Vision (SGLang, 2x Spark)

Credits:

  • The vLLM team, for the day-0 dsv41-feat branch.
  • Kai, for the first Engram-on-disk patch and the SM12x page-size patches.
  • The Engram-on-disk idea follows our own PLE-on-disk patch for Qwen3.8-Flash-Next.
  • Prior art acknowledged at the idea level: vLLM PR #54129 (VLLM_PLE_MMAP). No code was copied.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。