跳到正文

MiaAI-Lab

Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold

Qwen3.8 Flash Next on one DGX Spark (TensorFold)

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

Qwen3.8 Flash Next on one DGX Spark (TensorFold)

by Mia’a AI Lab

Serve Qwen3.8 Flash Next from a single NVIDIA DGX Spark (GB10, 128 GB) through an OpenAI-compatible API, with 5 concurrent requests at the full 262,144-token context and image and video input. It runs TensorFold v0.3.6.3 in NVIDIA’s PyTorch container, plus a small set of patches that make prompt processing about 1.7x faster without changing a single output token, and that give the model its own vision tower on CUDA.

  • Checkpoint: Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP (MLX 4-bit, group size 32, with the MTP draft head)
  • API model id: Qwen3.8-Flash-Next
  • KV pool: 1,310,720 tokens (5 streams x 262,144, int8 KV cache, ~23.4 GiB), 25% more than 4 streams
  • Images and videos in chat messages (image_url / video_url parts), see Images and video
  • One command: ./start.sh sets everything up on the first run and starts the server; ./stop.sh stops it

Performance

One DGX Spark, int8 KV cache, n-gram tables read from SSD and MTP drafting, measured through the OpenAI API. The 5 concurrent requests row is from the current default (5 streams x 262,144 tokens); the other rows and the prefill table were measured with 4 streams x 262,144 tokens.

Decode, prose

Concurrent requestsAggregatePer requestTime to first token
162.4 tok/s62.4 tok/s152 ms
290.5 tok/s46.3 tok/s257 ms
4106.7 tok/s28.9 tok/s436 ms
5119.3 tok/s27.0 tok/s528 ms

Prefill

PromptTokensPrefill speedTime to first token
8k8,2292,503 tok/s3.29 s
16k16,4252,520 tok/s6.52 s
32k32,8062,499 tok/s13.13 s
64k65,5752,414 tok/s27.17 s
128k131,1102,200 tok/s59.60 s

Against unpatched TensorFold v0.3.6.2 with the same settings, prefill went from ~1,350-1,490 tok/s to ~2,340-2,480 tok/s (3k-50k-token prompts), a ~195k-token prompt from ~208 s to ~97 s, and single-request decode rose ~4%. Every reply stayed byte-identical. The prefill table used 4,096-row prompt chunks, which VISION=0 keeps; with image input on (the default), chunks are 2,048 rows to make room for the vision tower: prompts of 12k-150k tokens took 4-5% longer in our runs (e.g. 149k tokens in 74.0 s instead of 70.5 s), and a ~195k-token prompt ~102 s. Decode is unchanged.

Requirements

  • A DGX Spark (or another GB10 system with 128 GB unified memory) with nothing else large on the GPU: the default setting needs ~115 GiB free when the server starts (see KV pool and memory).
  • Docker with the NVIDIA container runtime, and your user in the docker group.
  • ~160 GB free disk on a fresh machine: ~125 GB for the checkpoint download under ~/.cache/huggingface (~114 GB) and ~35 GB for the image under Docker’s root (~24 GB); scripts/prepare.sh checks both.
  • Optional: the hf CLI on the host (faster, resumable download) and a Hugging Face token in ~/.cache/huggingface/token or HF_TOKEN.

Quick start

git clone https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold.git
cd Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
./start.sh

That is all. The first run sets everything up (see below): it pulls the prebuilt image (~11 GB) and downloads the ~106 GiB checkpoint, then compiles the CUDA kernels for the GB10 (a few minutes, once). Later starts take ~2.5 minutes to load the weights. start.sh shows each step, the server’s log and the loading progress, runs a smoke test, prints Qwen3.8-Flash-Next is now LIVE! on port 8888 with the endpoint, and returns you to the shell.

curl -s http://:8888/v1/models

curl -s http://:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "Qwen3.8-Flash-Next",
  "messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
  "max_tokens": 1000
}'

Any OpenAI client works with base_url = "http://:8888/v1" and the model Qwen3.8-Flash-Next. Streaming, tool calls (typed parameters, e.g. arrays come back as JSON arrays), reasoning content, images and videos are supported. The model thinks before it answers (reasoning_content), so give replies enough max_tokens.

./start.sh restart                            # restart it, e.g. after changing a setting
./stop.sh                                     # stop the server and free the GPU memory
docker logs -f qwen38-flash-next-tf           # server log
curl -s http://:8888/health    # busy flag and live token totals

Images and video

The model’s own vision tower (27 layers, 0.84 GiB, from the same checkpoint) turns images and video frames into tokens, placed with Qwen’s 3-D rotary positions, the way the reference implementation does. Send them as OpenAI-style content parts in a user message:

IMG=$(base64 -w0 photo.jpg)
curl -s http://:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "Qwen3.8-Flash-Next",
  "messages": [{"role": "user", "content": [
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,'"$IMG"'"}},
    {"type": "text", "text": "What is in this picture?"}]}],
  "max_tokens": 2000
}'

A video is a video_url part ({"type": "video_url", "video_url": {"url": "data:video/mp4;base64,..."}}); with OpenAI’s Python client, pass the same dicts in messages.

ImagesVideos
FormatsJPEG, PNG, WebPMP4, WebM, MOV, MKV (anything FFmpeg decodes)
Per requestup to 50 (all of a chat’s turns count), 10 MB each, 64 MB in allup to 2, 64 MB each, 96 MB in all, up to an hour of footage
Tokensup to 16,384 for all images, at most 4,096 an image (50 images: ~320 each; "detail": "low": 256 an image)2 frames a second (at most 256 frames, spread over the whole video), each pair of frames one timestamped block; up to 16,384 tokens a request (TENSORFOLD_VIDEO_TOKENS)

By default only data URLs are accepted; VISION_URLS=1 also lets the server fetch public https:// URLs. Image and video prompts are not kept for prefix reuse, so each turn of a chat with images processes them again. A request body can be up to 96 MiB (base64 makes data URLs a third larger than the files). Text requests are unaffected: their replies stay byte-identical with vision on. VISION=0 ./start.sh restart serves text only.

Other languages

Recommended only for replies mostly in Chinese or Japanese; leave it off otherwise.

MTP drafts may only propose tokens from a list, and TensorFold’s list (79,591 tokens) is English and code: it holds 50 Chinese characters and 433 Cyrillic tokens. Replies in other languages still come out right (every token is checked against the full vocabulary), but fewer drafts are accepted, so they decode slower. A second image adds a language’s tokens to that list (patch patches/languages/0010). It is opt-in: put the language in a .env file next to start.sh and restart.

echo 'DRAFT_LANGUAGE=zh' >> .env     # or ja; several: zh,ja
./start.sh restart                   # switches to the language image (pulled or built the first time)

Remove the line (or leave it empty) and ./start.sh restart to go back to the default image. The output is byte-identical with either image; only speed changes. Measured on one Spark (one stream, recipe sampling, seed 1234, one boot per arm, 2026-09-29):

Replies inDefault imageLanguage imageChange
Chinese, thinking off / on35.5 / 38.4 tok/s45.9 / 50.6 tok/s+29% / +32%
Japanese, thinking off / on39.5 / 46.0 tok/s46.9 / 49.2 tok/s+19% / +7%

The larger list makes every draft step a little slower, which is why it does not pay off for English or code. DRAFT_LANGUAGE also accepts ru, de, fr and pt, but those have not been measured to help, so they are not recommended. The language token lists come from the vLLM recipe’s language draft vocabularies; see CREDITS.md.

What start.sh and scripts/prepare.sh do

./start.sh works in five steps, each shown as it runs:

  1. Setup: runs scripts/prepare.sh whenever the setup is not ready: on the first run, after the patches change, or with another model or image. It compares what prepare.sh last left ready with the current settings, so later starts skip it instantly.
  2. Checks: the arguments (with TensorFold’s own parser, in a throwaway container), the previous server, the port and the free memory.
  3. Launch: tensorfold serve with the settings from scripts/config.sh.
  4. Loading: the server’s log as it comes, and every 15 s the elapsed time and how much of the startup estimate is on the GPU. If the server stops, the last log lines and the reason are shown.
  5. Smoke test: one chat completion, then the LIVE message and the endpoint.

If the server is already running, ./start.sh says so and leaves it alone; ./start.sh restart stops it and starts it again. It stops the server only after the setup and the argument check pass, so a typo leaves the running server alone and the server is down only while it restarts. Stopping cuts off requests still running (stop.sh warns when there are any). Extra arguments go to tensorfold serve after the defaults, so they win (./start.sh restart --context 131072); ./start.sh --help lists the options. FOREGROUND=1 ./start.sh stays attached to the server’s log and exits with its exit code (for a systemd unit).

scripts/prepare.sh does the one-time setup, and is safe to re-run (each step skips work already done):

  1. Preflight: Docker, the NVIDIA runtime, disk space.
  2. The image tensorfold-qwen38:v0.3.6.3: TensorFold v0.3.6.3 with every patches/*.patch applied, plus transformers (the vision tower) and PyAV (video decoding), on NVIDIA’s PyTorch container (nvcr.io/nvidia/pytorch:26.07-py3). With DRAFT_LANGUAGE set it is tensorfold-qwen38:v0.3.6.3-languages instead, which also applies patches/languages/*.patch. It first tries the matching prebuilt image from GitHub Container Registry (ghcr.io/miaai-lab/qwen3.8-flash-next-single-dgx-spark-tensorfold:v0.3.6.3-, ~11 GB; :latest is the default image, :languages the language image); if that tag is not there (e.g. after you change patches/), or with PULL=0, it builds the image locally instead (a few minutes).
  3. Downloads the checkpoint into ~/.cache/huggingface (resumable).
  4. Verifies the checkpoint with tensorfold info.

Run it yourself to download ahead of time or to rebuild the image from scratch:

scripts/prepare.sh             # set up without starting the server
scripts/prepare.sh --rebuild   # rebuild the image from scratch
PREPARE=1 ./start.sh restart   # force prepare.sh, then restart; PREPARE=0 skips the check

After changing patches/, scripts/publish-image.sh pushes the new image to GitHub Container Registry (latest and v0.3.6.3-), and DRAFT_LANGUAGE=zh scripts/publish-image.sh the language image (languages and its own v0.3.6.3-).

KV pool and memory

TensorFold gives every stream its own cache for a full window, so the KV pool is streams x window:

Default
Streams (PARALLEL)5
Window per stream (CONTEXT, the model’s native maximum)262,144 tokens
KV pool1,310,720 tokens (4 streams: 1,048,576)
KV precision (KV_DTYPE)int8 (an fp16 scale per 32 values)
Memory a stream, allocated (server log)4,799 MiB: the KV cache, the sparse-attention index and the stream’s own buffers
Memory for the pool, allocated~23.4 GiB (5 x 4,799 MiB)

The server reports these at every start: 5 streams of 262144 prompt/reply tokens (4799 MiB a stream) and startup estimate 102.50 GiB within 103.26 GiB (the budget varies a little from start to start).

Where the memory goes at the default setting (TensorFold’s startup estimate):

GiB
Model weights (the 29.8 GiB of n-gram tables stay on the SSD with PLE_ON_SSD=1)75.2
Vision tower (VISION=1)0.84
Stream caches (5 x 4.49, the context-sized part)22.5
Fixed buffers (DeltaNet states, decode windows, 2,048-row prompt-chunk scratch, 8 saved prompt states)4.0
Startup estimate102.5

The vision tower’s scratch (~0.8 GiB at most, measured on a 4,096-token image and a 256-frame video) is taken only while an image or video encodes and is handed back right after; startup reserves none for it (TENSORFOLD_VISION_WORKSPACE_MIB). With VISION=0 the tower is not loaded and the chunk scratch is 4,096 rows (estimate 102.6 GiB).

TensorFold’s budget is the free memory at start (MemAvailable) minus a host reserve of a tenth of RAM (12.2 GiB), so ~103-104 GiB on an otherwise idle Spark. The reserve covers what the estimate leaves out (CUDA context, workspaces, the Python process) and the host itself: on the Spark’s unified memory, running out tends to freeze the machine rather than fail an allocation. At the default setting (vision on) the host kept at least 8.3 GiB free through a 195k-token prompt with image, video and text requests running alongside (9.7 GiB with VISION=0 under a 195k-token prompt and 5 concurrent long requests).

Other settings that fit the same budget (TensorFold’s own estimate):

SettingKV poolEstimateNote
PARALLEL=4 (int8)1,048,57697.7 GiBmore headroom
PARALLEL=5 (int8, default)1,310,720102.5 GiB
PARALLEL=6 CONTEXT=220000 (int8)1,320,000~103 GiBshorter windows, one more stream
PARALLEL=6 KV_DTYPE=int41,572,86497.6 GiBint4 changes outputs slightly; quality not measured here
PARALLEL=8 KV_DTYPE=int4 CONTEXT=2500002,000,000~103 GiBtight
PARALLEL=3 KV_DTYPE=bf16786,432102.0 GiBfull-precision KV

A setting that does not fit is refused at startup, before any weights load, with a message naming a window that fits.

Configuration

Every setting lives in scripts/config.sh and can be overridden from the environment (PARALLEL=4 ./start.sh), in a .env file next to start.sh (KEY=value lines, e.g. PARALLEL=4; the environment wins over it), or with tensorfold serve flags (./start.sh --context 131072).

VariableDefaultMeaning
PARALLEL5requests decoded together (streams)
CONTEXT262144prompt + reply window per stream
KV_DTYPEint8bf16, int8 or int4 KV cache
PLE_ON_SSD1read the 29.8 GiB n-gram tables from SSD instead of RAM, leaving that memory to the KV cache
VISION1image and video input (--vision); 0 serves text only
VISION_URLS01 also accepts public https:// image and video URLs (default: data URLs only)
DRAFT_LANGUAGEemptyzh or ja: serve the language image, for replies mostly in that language (Other languages)
MTP_DRAFTS / MTP_CONFIDENCE6 / 0.60at most 6 MTP drafts a round; a chain stops before a draft under 60%
TEMPERATURE / TOP_P / TOP_K1.0 / 0.95 / 20default sampling (Qwen’s thinking-mode values); a request’s own values win
THINKING1open a think block by default; 0 answers directly unless a request asks to think
SERVED_NAMEQwen3.8-Flash-Nextthe model id in /v1/models and in replies
PORT / HOST8888 / 0.0.0.0where the API listens
TENSORFOLD_PREFILL_ROWS2048 (4096 with VISION=0)rows per prompt chunk (patch 0006); 4,096 is 2-5% faster from 3k tokens and takes 0.94 GiB more
TENSORFOLD_MTP_COPY1prompt-lookup drafts for text that repeats the prompt (patch 0007; needs PARALLEL >= 2); 0 turns them off
TENSORFOLD_MAX_IMAGES / TENSORFOLD_IMAGE_TOKENS50 / 16384images a request may carry and the tokens they share, each at most 4,096 (patch 0009)
TENSORFOLD_VIDEO_TOKENS16384a request’s video token budget
TENSORFOLD_VISION_WORKSPACE_MIB0what startup reserves for the vision tower’s scratch
PREPAREautostart.sh runs scripts/prepare.sh when needed; 1 always, 0 never
PULL1prepare.sh tries the prebuilt image first; 0 always builds locally
STOP_TIMEOUT30seconds stop.sh gives the server to shut down before removing it

Any TENSORFOLD_* variable in the environment is passed into the container (TENSORFOLD_NO_UPDATE_CHECK=1, the default, stops TensorFold asking GitHub for a newer release at each start). Less common settings are described in scripts/config.sh: MODEL_ID, TF_VERSION, TF_REPO, BASE_IMAGE (the patches are made for TensorFold v0.3.6.3; after changing any of these run scripts/prepare.sh --rebuild), IMAGE, CONTAINER_NAME, GHCR_IMAGE, HF_CACHE (default $HF_HOME or ~/.cache/huggingface), KERNEL_CACHE, MIN_FREE_GB, IMAGE_FREE_GB. start.sh also takes FOREGROUND=1, WAIT_TIMEOUT (seconds, default 1800) and HF_HUB_OFFLINE=0 (let the server reach Hugging Face; by default it serves from the local cache only).

Thinking and sampling

By default the model thinks before it answers, with Qwen’s recommended thinking-mode sampling: temperature 1.0, top_p 0.95, top_k 20. TensorFold has no min_p, presence penalty or repetition penalty, which is the same as min_p 0.0, presence_penalty 0.0 and repetition_penalty 1.0; requests that send those fields are served as if they had not. Per request:

  • temperature, top_p, top_k and seed override the defaults (temperature: 0 decodes greedily).
  • "chat_template_kwargs": {"enable_thinking": false} answers without thinking, and "chat_template_kwargs": {"reasoning_effort": "low"} (or "xhigh") sets Qwen’s reasoning effort; without it the template’s default (medium) applies. A top-level OpenAI-style reasoning_effort field is ignored.
  • The reasoning comes back in reasoning_content, the answer in content.

What the patches change

scripts/prepare.sh bakes every patches/*.patch into the image (unified diffs against TensorFold’s site-packages, applied with patch -p0), and start.sh rebuilds or re-pulls the image by itself when the patches change.

PatchChangeEffect
0001-cuda-live-token-counters/health reports live token totalsmonitoring (upstream #79)
0002-flash-next-ssd-read-aheada prompt chunk’s n-gram rows are read from SSD while the GPU processes the previous chunkmulti-chunk prefill +50%
0003-flash-next-ssd-native-readerthose reads run on a C++ thread pool outside the Python GILshort prompts’ time to first token -35%, 3k-12k prefill +10-40% on top, decode +4%
0004-flash-next-qsa-tiled-selectthe sparse-attention block selection no longer spills registers past 128k tokens149k-token prompts 25% faster; decode at 149k context +19%
0005-cuda-stream-draft-statsdrafted / accepted counts in concurrent requests’ statsobservability
0006-flash-next-prefill-rowsconfigurable prompt chunk size (port of #40)+2-5% at 4,096 rows
0007-flash-next-copy-draftsdrafts copied from earlier text when the reply repeats the prompt+6% on quoting and editing replies
0008-flash-next-visionimage and video input for Flash Next on CUDA: the Qwen3.5 vision tower, interleaved 3-D rotary positions in the attention and sparse-attention kernels, video frames in timestamped blocks--vision (TensorFold’s own --vision covers only the dense 27B)
0009-flash-next-many-imagesup to 50 images a request sharing 16,384 tokens (4,096 at most an image), encoded by the vision tower in bounded runs; request bodies up to 96 MiBmany-image chats; one image is encoded exactly as before
languages/0010-flash-next-draft-languagesonly in the opt-in language image (DRAFT_LANGUAGE): Chinese and Japanese (also Russian, German, French, Portuguese) tokens added to the list MTP drafts fromChinese +29-32%, Japanese +7-19% decode (Other languages)

Typed tool-call parameters (this recipe’s former patch 0001, #75) are part of TensorFold v0.3.6.3.

Outputs are unchanged. Every speed patch changes speed only: drafts are verified against the model’s own keyed samples, and the prefill changes read the same bytes and select the same attention blocks. This was checked by comparing reply hashes (sampled and greedy, prompts up to 149k tokens) against unpatched TensorFold, with vision on and off, and with a ~195k-token needle-in-a-haystack test. Text rows take exactly the rotary path they always did; on image and video prompts, drafted replies equal the serial reference too. Any request can also be sent with "draft": false to get TensorFold’s serial, one-token-at-a-time reference.

Checks

The scripts in tools/ talk to the running server (API_URL, default http://127.0.0.1:8888; or just PORT), from this machine or another one (API_URL=http://:8888 tools/bench.py):

ScriptWhat it does
tools/bench.py [label]prefill at ~0.85k / 3.2k / 12.6k / 50k tokens (fresh random prompts) and a short decode check
tools/needle.pyhides a passphrase in a ~195k-token prompt and checks the model returns it
tools/toolcheck.pymakes a tool call with an array parameter and checks it comes back as a JSON array
tools/visioncheck.pysends a drawn image (a red circle and a blue square) and checks the model names both

Repository layout

start.sh      set up (first run) and start the server
stop.sh       stop it
scripts/      prepare.sh (image + checkpoint), config.sh (all settings), publish-image.sh (push the image to GHCR),
              banner.sh (start.sh's banner)
patches/      patches baked into the image; patches/languages/ only into the language image (DRAFT_LANGUAGE)
tools/        benchmark and checks
.github/      issue and pull request templates, GitHub Sponsors
CREDITS.md    who and what this builds on

License

MIT, see LICENSE, which also carries TensorFold’s MIT notice for the patches. The model weights, downloaded from Hugging Face and not part of this repository, are under the Qwen Community License 1.0.

Third-party software in the image. The prebuilt image (and the one scripts/prepare.sh builds) is based on NVIDIA’s PyTorch container nvcr.io/nvidia/pytorch:26.07-py3, redistributed as a value-added runtime image. The NVIDIA software in it is governed by the NVIDIA Software License Agreement and the Product-Specific Terms for NVIDIA AI Products, which the container prints at every start (it shows in start.sh’s output); by pulling or running the image you accept them. The image also contains Hugging Face transformers (Apache 2.0) and PyAV (BSD) with its FFmpeg libraries (LGPL). The MIT license above covers this repository’s scripts and patches only.

Credits

Built on TensorFold by Ash Hart (ashhart), Qwen3.8 Flash Next by Qwen, and Vontra’s MLX 4-bit checkpoint, with a prompt-chunk change by MovieMaker93 (TensorFold #40). The full list, including the runtime stack and licenses, is in CREDITS.md.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。