incoai
splash
A local inference engine for Apple silicon, built around the model.
Documentation snapshot
README 快照
翻译暂时拿不到。
机器翻译的项目简介,仅供参考。原文在下方,也可以直接用浏览器自带的整页翻译 (Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」)。
下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。
本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。
Splash
图片:CI 图片:License 图片:Platform
A local inference engine for Apple silicon, built around the model.
Splash serves a small set of models to coding agents and to any OpenAI or Anthropic compatible client, on one Mac. On a 48 GB M5 Pro it decodes Qwen3.8-27B at 2× the speed of the next-fastest engine and, with a 32K context cached, returns the first token in 282 ms. Its kernels, draft model, and memory plan are specialized for each model it serves. That is why it is fast, and why there is nothing to configure.
Quick start
Apple M3 or newer, macOS 26.4 or later, Homebrew, and 36 GB of unified memory (48 GB or more recommended).
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.8-27B-Splash
The first run downloads and verifies the model package, checks available
memory, and starts serving on 127.0.0.1:8000.
Once it prints Ready, leave this terminal open. Open
in your browser, or run an installed coding agent from another terminal:
splash opencode # or: splash claude / splash codex / splash hermes
Press Ctrl+C in the server terminal to stop Splash.
Use the API
Splash speaks OpenAI Chat Completions (/v1/chat/completions), OpenAI Responses
(/v1/responses), and Anthropic Messages (/v1/messages), all with streaming,
tool calls, JSON Schema output, images, and inline PDFs. /tokenize and
/apply-template return token IDs and the rendered prompt without running the
model.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "incoai/Qwen3.8-27B-Splash",
"messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}]
}'
model is optional. If you set it, it must match the package you served.
Reasoning is on by default. "reasoning_effort": "none" turns it off, and
Qwen3.8-27B also takes low, medium, and xhigh.
Models
Package (--model) | Contents | Download |
|---|---|---|
incoai/Qwen3.8-27B-Splash | Qwen3.8-27B, 4-bit, with its DFlash 2 draft | 17.4 GB |
incoai/Qwen3.6-35B-A3B-Splash | Qwen3.6-35B-A3B, 4-bit, with its DFlash 2 draft | 20.9 GB |
--model takes any owner/repo that holds a Splash package, a format
DEVELOPMENT.md describes. Plain MLX or
Transformers checkpoints do not work. Private repositories need HF_TOKEN.
Packages download into the Hugging Face cache, and brew upgrade splash keeps
them, along with model links and agent sessions.
To download new models to another disk, set HF_HUB_CACHE before the first run:
HF_HUB_CACHE=/Volumes/Models/huggingface splash serve --model incoai/Qwen3.8-27B-Splash
Model links and agent sessions stay in Splash’s data directory. Existing models are not moved.
Settings
There is no config file. The server binds 127.0.0.1:8000 by default.
Context supports up to the model’s native 256K window; usable capacity
depends on available memory.
splash serve accepts these optional flags:
--port: local HTTP port. Defaults toSPLASH_PORTor8000.--max-memory: ceiling on Metal allocations, e.g.28G. Default: auto.--max-context: context limit, up to256K, e.g.100K. Default: auto.--max-image-pixels: maximum resized pixels per image. Default: 4,194,304.--allowed-host: extra HTTPHostname to accept, for a proxy. Repeatable.--api-key: require this key on API requests, as a bearer token orx-api-key. Defaults toSPLASH_API_KEY.--no-webui: turn off the chat page.
Set SPLASH_PORT in both the server and agent shells to use another port.
Separate ports allow separate servers; their memory limits are independent.
If the model does not fit in the memory available, startup prints a memory budget breakdown and stops.
Authentication is off by default. Set SPLASH_API_KEY in the shell that runs
splash serve and in the shell that runs an agent, and both sides use it.
Health and readiness probes stay public.
- Experimental cache offloading: PR #3
adds SSD offloading for KV cache and GDN states. Build from that branch and
set
--max-cache-disk 8Gto enable it. This helps preserve reusable prefixes when RAM is limited, reducing repeated prefill.
Performance
Measured on an M5 Pro (16-core GPU, 48 GB): selected SPEED-Bench coding prompts over HTTP, a 1,024-token output limit, reasoning on (medium for the 27B). The ratio in each cell is against the next-fastest engine we measured.
| Metric | Qwen3.6-35B-A3B | Qwen3.8-27B |
|---|---|---|
| Decode · short prompt | 210 tok/s (1.7×) | 74 tok/s (2.0×) |
| Prefill · 32K prompt | 2,011 tok/s (1.3×) | 363 tok/s (1.2×) |
| Cached time to first token · 32K replay | 123 ms (6.6×) | 282 ms (7.3×) |
| Aggregate decode · 4 concurrent short prompts | 357 tok/s (2.0×) | 170 tok/s (3.9×) |
Splash led on every measure at every prompt length we tested, and the lead grows with load: 3.8× at four concurrent 32K requests on the 35B. The launch post has the method and the full comparison against oMLX, Lily, uzu, and Ollama.
Design
The runtime, scheduler, cache, and API are shared. Everything else is rebuilt per model:
- A draft trained for the model. Speculative decoding is the decode path in Splash, not an option. Every model ships with its own DFlash 2 draft, and one pass of the target verifies a block of tokens in parallel.
- Kernels compiled for exact shapes. Fused Metal kernels, written and tuned by our in-house kernel agents for the model’s dimensions, read weights packed for them and mapped zero-copy from disk. They ship precompiled: no Xcode, no compiler toolchain, nothing tuned on your machine.
- A memory plan computed for this machine. Context, KV capacity, and batch limits are worked out at startup from the memory Metal recommends, less the weights, the draft, and each request’s state.
The launch post covers the design in depth.
More
- DEVELOPMENT.md: building from source, tests, model packages, and release packaging.
- Apache-2.0, see LICENSE. Model weights keep their own licenses.
Official distribution
获取与安装
暂未发现可确认的官方软件包地址
当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。
本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。
Before installing
使用前核验
本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。