antirez/ds4Public

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

AI summary: Lightweight native inference engine optimized for DeepSeek V4 Flash on consumer hardware.

Stars
20.8K
+212 today
Forks
1.9K
Watchers
163
Open issues
125
Open PRs
225
Contributors
~40
Commits
383
Branches
4

CMITCreated May 6, 2026Last push 2d ago+1.4K stars this week+1.5K this month

Star history

since May 3, 2026
010K20KMay 2026Jun 2026Jul 2026Aug 2026
20.8K stars as of Aug 7, 2026, tracked back to May 3, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulMonWedFri2025-08-03: 0 commits2025-08-04: 0 commits2025-08-05: 0 commits2025-08-06: 0 commits2025-08-07: 0 commits2025-08-08: 0 commits2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 9 commits2026-05-08: 3 commits2026-05-09: 11 commits2026-05-10: 8 commits2026-05-11: 13 commits2026-05-12: 23 commits2026-05-13: 11 commits2026-05-14: 15 commits2026-05-15: 10 commits2026-05-16: 12 commits2026-05-17: 4 commits2026-05-18: 3 commits2026-05-19: 2 commits2026-05-20: 21 commits2026-05-21: 2 commits2026-05-22: 0 commits2026-05-23: 26 commits2026-05-24: 22 commits2026-05-25: 10 commits2026-05-26: 2 commits2026-05-27: 5 commits2026-05-28: 4 commits2026-05-29: 8 commits2026-05-30: 8 commits2026-05-31: 3 commits2026-06-01: 2 commits2026-06-02: 2 commits2026-06-03: 0 commits2026-06-04: 2 commits2026-06-05: 6 commits2026-06-06: 7 commits2026-06-07: 9 commits2026-06-08: 9 commits2026-06-09: 4 commits2026-06-10: 2 commits2026-06-11: 2 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 5 commits2026-06-15: 2 commits2026-06-16: 5 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 10 commits2026-07-19: 0 commits2026-07-20: 4 commits2026-07-21: 0 commits2026-07-22: 9 commits2026-07-23: 1 commit2026-07-24: 0 commits2026-07-25: 9 commits2026-07-26: 14 commits2026-07-27: 1 commit2026-07-28: 1 commit2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits
341 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    20,836 stars

  • Breakout launch

    20,836 stars in 93 days

  • Rising fast

    +1,377 stars this week

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    6 trending appearances

What ds4 does

DwarfStar (ds4) is a highly specialized, self-contained native inference engine designed specifically for running DeepSeek V4 Flash (and GLM 5.2) on consumer machines like MacBooks and high-memory workstations. Built on top of GGML, it intentionally avoids being a general-purpose runner, instead tightly coupling model loading, prompt rendering, KV caching, and a built-in HTTP server. It supports Metal, CUDA, and ROCm backends, and includes SSD streaming to allow massive models to run on systems with limited RAM.

AI engineers and power users wanting to run specific cutting-edge models (DeepSeek V4) locally on high-end consumer hardware.

  • Feature: Runs DeepSeek V4 Flash and GLM 5.2 natively on consumer hardware.
  • Feature: Supports Metal (Apple Silicon), NVIDIA CUDA, and ROCm backends.
  • Feature: Implements SSD streaming to run large models on machines with insufficient RAM.
  • Feature: Integrates an HTTP server, tool calling, and prompt rendering in a single self-contained binary.
  • Feature: Includes multi-GPU support for distributed inference on systems like DGX Spark.

Where teams use it

Local High-End Inference

Developers with 96GB+ Macs can run DeepSeek V4 Flash locally for private coding assistance without cloud APIs.

Memory-Constrained Execution

Users with standard laptops can leverage SSD streaming to run models that technically exceed their available RAM.

Multi-GPU Deployments

Researchers can deploy the engine on multi-GPU CUDA workstations using the integrated ds4-server for shared network access.

Getting started: See repository documentation for build instructions based on your hardware backend.

README

main branch

DwarfStar logo

DwarfStar is a small native inference engine optimized first for DeepSeek V4 Flash. It also supports GLM 5.2 and, on very high-memory machines, DeepSeek V4 PRO. It is self-contained and deliberately narrow, not a general GGUF runner. Model loading, prompt rendering, tool calls, KV state, the HTTP server, and the coding agent are built and tested together. The repository also includes tools and data for GGUF, imatrix, quality, and speed.

Supported backends:

  • Metal, the primary target, on Macs with 96 GB or more. Smaller machines can use SSD streaming.
  • NVIDIA CUDA, including multi-GPU systems and DGX Spark.
  • ROCm on Strix Halo systems such as the Framework Desktop.

This project would not exist without llama.cpp and GGML, make sure to read the acknowledgements section, a big thank you to Georgi Gerganov and all the other contributors.

Model support is intentionally opportunistic. The project follows the best open weights for useful local machine sizes, especially 128 GB laptops and 512 GB workstations. A model may be removed when a better replacement arrives.

So, what can I do with this software?

  • You can run a very capable models in your consumer hardware, a MacBook, a DGX Spark, or a Strix Halo for example. Even if you have not enough RAM, with SSD streaming, you can run it at a decent speed.
  • Using the CUDA multi-GPU support and with ds4-server micro batching of decoding and generation, you can turn a server with old-ish CUDA cards (Ada Lovelace architecture), no longer supported for new models by vLLM, into a multi-user LLM server for your company. We tested this setup with 8xL40S NVIDIA cards and multiple sessions with very good results. 120 t/s aggreated generation, 2000 t/s prefill.
  • Using two MacBook M5 Max / M3 Ultra RDMA, you can run 4 bit DeepSeek Flash or GLM 5.2 with tensor parallelism.
  • You can also use pipeline paralellism to glue together multiple systems to sum their RAM and run larger models.

Motivations

  • Capable open-weight models now fit on high-end personal machines.
  • DeepSeek V4 Flash and PRO, GLM 5.2, tolerate aggressive routed-expert quantization.
  • Compressed KV caches and fast local SSDs make long contexts practical.
  • The idea of an inference system specialized for a few models.

AI full disclosure

  • This software is developed with strong assistance from GPT 5.5, 5.6, Claude Fable and with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you. The acknowledgement below is equally important: this would not exist without llama.cpp and GGML, largely written by hand.

Acknowledgements to llama.cpp and GGML

ds4.c does not link against GGML, but it exists thanks to the path opened by the llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won engineering knowledge developed there. We are thankful and indebted to llama.cpp and its contributors. Their implementation, kernels, tests, and design choices were an essential reference while building this DeepSeek V4 specific inference path. Some source-level pieces are retained or adapted here under the MIT license: GGUF quant layouts and tables, CPU quant/dot logic, and certain kernels. For this reason, and because we are genuinely grateful, we keep the GGML authors copyright notice in our LICENSE file.

Status

The software is currently very fast changing. Consider it beta quality. Before each release, a big QA run is executed, however instabilities are definitely possible.

More Documentation

If you are looking for very specific things, we have other sub-README files. Otherwise for normal usage keep reading the next sections.

Model Weights

This implementation only works with the DeepSeek V4 and GLM 5.2 GGUFs listed below. It is not a general GGUF loader, and arbitrary GGUF files will not have the tensor layout, quantization mix, metadata, or optional MTP state expected by the engine. The 2 bit quantizations provided here are verified to be actually high quality: they behave well, work under coding agents, call tools in a reliable way.

The 2 bit quants use a very asymmetrical quantization: only the routed MoE experts are quantized, up/gate at IQ2_XXS, down at Q2_K. They are the majority of all the model space: the other components (shared experts, projections, routing) are left untouched to guarantee quality.

Download one main model. Prefer the imatrix versions.

./download_model.sh q2-imatrix   # 96/128 GB RAM machines, imatrix-tuned q2
./download_model.sh q2-q4-imatrix  # 96/128 GB RAM machines, q2 with last 6 layers q4
./download_model.sh q4-imatrix   # >= 256 GB RAM machines, imatrix-tuned q4
./download_model.sh pro-q2-imatrix  # 512 GB RAM machines, PRO q2 imatrix quant

For the full PRO Q4 distributed run, download one half on each machine:

./download_model.sh pro-q4-layers00-30      # first half of PRO Q4 split
./download_model.sh pro-q4-layers31-output  # second half of PRO Q4 split

The script downloads from https://huggingface.co/antirez/deepseek-v4-gguf, stores files under ./gguf/, resumes partial downloads with curl -C -, and updates ./ds4flash.gguf to point at the selected main model. The pro-q4-layers00-30, pro-q4-layers31-output, and pro-q4-split targets download distributed PRO Q4 pieces and do not update ./ds4flash.gguf. Authentication is optional for public downloads, but --token TOKEN, HF_TOKEN, or the local Hugging Face token cache are used when present.

If you want to regenerate GGUF files or collect a new imatrix, see gguf-tools/README.md. Those tools are meant for offline model-building work and can take a long time on the full DeepSeek V4 Flash weights. Flash GGUF generation is supported by the local tools. PRO GGUF production currently still depends on the external llama.cpp-based workflow; native tooling can be added later.

./download_model.sh mtp fetches the optional speculative decoding support GGUF for Flash. It can be used with q2-imatrix, q2-q4-imatrix, and q4-imatrix, but must be enabled explicitly with --mtp. The current MTP/speculative decoding path is still experimental: it is correctness-gated and currently provides at most a slight speedup, not a meaningful generation-speed win.

GLM 5.2 support is limited to the GGUF files tested by this branch:

./download_model.sh glm-unsloth-q4  # Unsloth UD-Q4_K_XL, 11 shards
./download_model.sh glm-antirez-iq2xxs  # antirez routed IQ2_XXS single-file GGUF
./download_model.sh glm-antirez-q2  # antirez routed Q2_K single-file GGUF
./download_model.sh glm-antirez-q4  # antirez routed Q4_K single-file GGUF

The supported GLM layout keeps dense/model-control tensors in the existing Q8/F32 paths and supports routed expert gate/up tensors in Q2_K, Q4_K, or Q5_K; routed expert down tensors are supported in Q2_K, Q4_K, Q5_K, or Q6_K. Other GLM GGUF quant layouts should be treated as unsupported until they are added deliberately and scored against the official 100-case fixture.

These formats do not all support the same execution modes. The Q4 files work for normal Metal and CUDA inference. Two-Mac tensor parallelism currently requires an ownership-aware IQ2_XXS or Q2_K routed layout; a routed Q4 GLM must be rejected before evaluation.

GLM's MTP block is part of the main GGUF; it does not use the separate Flash MTP file. Ordinary decode remains the default. --glm-mtp enables experimental greedy speculation. --glm-mtp-timing also enables it and prints acceptance and timing counters:

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --glm-mtp-timing --temp 0

GLM inference uses the Metal, CUDA, or ROCm graph backend. Directional steering, --power below 100, an explicit --prefill-chunk, and the external --mtp file are not supported for GLM yet.

Then build:

make                  # macOS Metal
make cuda-spark       # Linux CUDA, DGX Spark / GB10
make cuda-generic     # Linux CUDA, other local CUDA GPUs
make strix-halo       # Linux ROCm, AMD Strix Halo
make cpu              # CPU-only diagnostics build

./ds4flash.gguf is the default model path used by both binaries. Pass -m to select another supported GGUF from ./gguf/. Run ./ds4 --help and ./ds4-server --help for the full flag list.

DSpark Speculative Decoding

DSpark is an auxiliary draft model released by DeepSeek for DeepSeek V4 Flash. It reads hidden states from the main model and proposes up to five future tokens. DwarfStar checks those proposals with the main Flash model and commits only the accepted prefix. The main model remains authoritative; a rejected or low-confidence suffix falls back to ordinary target decoding.

The possible gain is faster generation: when several proposed tokens are accepted, one target verification pass advances the stream by several tokens. It does not accelerate prefill, and the draft and verification work is not free. Predictable continuations, especially code, tend to benefit most; low-yield prompts can be no faster or even slower. DSpark is therefore still experimental and explicitly opt-in.

The released DSpark checkpoint is packaged here as a separate support GGUF of about 5.6 GiB. It is not a standalone model. Download it once:

./download_model.sh dspark-support

The same support file can be used with the Flash q2-imatrix, q2-q4-imatrix, and q4-imatrix models listed above. For now DeepSeek V4 PRO is not supported. On Metal, the main model may be resident or use --ssd-streaming; the support model still adds its own weights and runtime state to the memory requirement. DSpark replaces the legacy one-stage MTP support model for that run rather than stacking with it.

Run it with greedy decoding:

./ds4 -m ds4flash.gguf \
  --mtp gguf/DeepSeek-V4-Flash-DSpark-support.gguf \
  --dspark --temp 0

--mtp supplies the support GGUF, while --dspark selects the DSpark runtime. The default confidence threshold is 0.9; it prunes suffixes that are unlikely to repay their verification cost. --dspark-confidence 0 forces fixed five-token blocks and is intended for diagnostics. Sampled decoding does not use DSpark proposals. --quality and --dspark-strict also keep target-only decoding, which is useful for comparisons and correctness checks.

Speed

Warning: some of those numbers may no longer be updated, because of the optimization efforts that improved the runtime speed without updating the benchmark results.

These are single-run Metal CLI numbers with --ctx 32768, --nothink, greedy decoding, and -n 256. The short prompt is a normal small Italian story prompt. The long prompts exercise chunked prefill plus long-context decode. Q4 requires the larger-memory machine class, so M3 Max Q4 numbers are N/A.

Machine Quant Prompt Prefill Generation
MacBook Pro M3 Max, 128 GB q2 short 58.52 t/s 26.68 t/s
MacBook Pro M3 Max, 128 GB q2 11709 tokens 250.11 t/s 21.47 t/s
MacBook Pro M3 Max, 128 GB q4 short N/A N/A
MacBook Pro M3 Max, 128 GB q4 long N/A N/A
MacBook Pro M5 Max, 128 GB q2 short 87.25 t/s 34.27 t/s
MacBook Pro M5 Max, 128 GB q2 11707 tokens 463.44 t/s 25.90 t/s
Mac Studio M3 Ultra, 512 GB q2 short 84.43 t/s 36.86 t/s
Mac Studio M3 Ultra, 512 GB q2 11709 tokens 468.03 t/s 27.39 t/s
Mac Studio M3 Ultra, 512 GB q4 short 78.95 t/s 35.50 t/s
Mac Studio M3 Ultra, 512 GB q4 12018 tokens 448.82 t/s 26.62 t/s
Mac Studio M3 Ultra, 512 GB PRO q2 32768 tokens 138.82 t/s 9.56 t/s
DGX Spark GB10, 128 GB q2 7047 tokens 343.81 t/s 13.75 t/s

M3 Max t/s PRO model M3 Ultra t/s

Running models larger than RAM

The normal Metal path tries to make the model resident in GPU-addressable memory. This is the fastest path and should remain your default when the model fits. DwarfStar also has an SSD streaming capacity mode on Metal and for GLM 5.2 on ROCm. In this mode the non-routed model weights stay resident, while routed MoE experts are kept in an in-memory cache and loaded from the GGUF file on cache misses.

Streaming is not as fast as fitting the full model in RAM. It still needs memory for non-routed weights, KV cache, graph scratch, activations, and the routed expert cache. It is useful because routed experts dominate model size and modern Mac SSDs are fast enough to make cache misses tolerable. Long prefills can still be fast; generation is more sensitive to cache misses because every new token routes through experts again.

Start with the automatic cache budget:

./ds4 -m ./ds4flash.gguf --ssd-streaming

If startup reports that the expert cache is too large, or if you want to reserve more memory for context, set the routed expert cache explicitly:

./ds4 -m ./ds4flash.gguf --ssd-streaming --ssd-streaming-cache-experts 32GB

The 32GB value is a routed-expert memory budget, not a generic byte cache. DwarfStar first reserves headroom for the two full routed layers used by overlapped streaming prefill, then converts the remaining bytes to the number of dynamic cached experts that fit for the current GGUF. Explicit NGB budgets may also be capped after context/KV accounting so the backend working set stays out of the slow pressure zone. A plain number such as --ssd-streaming-cache-experts 4000 is different: it means exactly 4000 dynamic expert slots, with no extra accounting. Non-routed weights, KV cache, graph scratch, and activations need additional memory. The automatic cache budget takes 80% of the backend's recommended working set, subtracts non-routed weights, then applies the same routed-prefill headroom before sizing the dynamic cache. Leave the hot expert preload enabled for normal use; use --ssd-streaming-cold and --ssd-streaming-preload-experts N only for measurements.

Practical SSD streaming examples

On 64GB MacBooks, start with the 2-bit Flash GGUF and a moderate expert cache:

./download_model.sh q2-imatrix

./ds4 \
  -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB \
  --ctx 32768 \
  --nothink

On 128GB MacBooks, PRO q2 streaming is experimental but usable for inspection and occasional work when you accept slow generation. Start with --nothink:

./download_model.sh pro-q2-imatrix

./ds4 \
  -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
  --ssd-streaming \
  --ctx 32768 \
  --nothink

On an M5 Max with 128GB of RAM, a short PRO q2 streaming decode benchmark found the automatic budget best: it selected about 59GB of routed expert cache. Manual 64GB to 75GB caches were close on that machine. Prefer the automatic budget; if setting the cache manually on this class of machine, start around 48GB to 64GB, then increase only while the machine remains responsive and the startup log shows the requested dynamic cache. Once the machine is stable, re-enable thinking with a conservative generation limit:

./ds4 \
  -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
  --ssd-streaming \
  --ctx 32768 \
  --think \
  --tokens 1500

GLM 5.2 uses the same option. Its streaming path keeps the largest full-layer prefix that fits resident, then uses the remaining budget for a dynamic expert cache. Start with the automatic budget:

./ds4 \
  -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --ssd-streaming \
  --ctx 32768

The important startup line is the cache report. Start conservative, then increase the cache if the machine has headroom.

On a 128GB Strix Halo, use the routed Q2_K model and a 4096-token context as the starting point. The automatic cache budget leaves room for the GLM graph and KV state:

./download_model.sh glm-antirez-q2
make strix-halo
./ds4 --rocm -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
  --ssd-streaming --ctx 4096

Distributed inference with pipeline parallelism

Pipeline parallelism lets DwarfStar run a model that is too large for one machine by splitting transformer layers across multiple machines. The main example is the full 4-bit Flash quant across two 128 GB MacBooks: each process maps only its own layer slice, activations are sent over TCP, and the coordinator keeps normal CLI/API behavior.

Pipeline parallelism can also speed up prefill by using multiple GPUs at the same time to process different micro-batches at different layers, like in an assembly line. Only prefill can be accelerated this way. Generation is purely autoregressive: each token must finish across the route before the next token can start. The model work is the same as a single process, plus coordination latency, so distributed generation is slower.

To build an initial mental model, here are the high level concepts:

  1. You put the GGUF on every machine, but each one loads just a subset. --layers controls which tensors are mapped, so a worker with --layers 20:output does not load the earlier layers.
  2. Layer ranges are inclusive: 10:20 means layers 10, 11, ..., 20. N:output means layer N through the final layer plus the output head.
  3. You assign one of the machines the role of coordinator, the others the roles of workers. Workers will connect to the coordinator and will tell they are there and which layers they are able to process.
  4. Each worker keeps its slice of the KV cache.
  5. Communication is worker-to-worker, there is no need to use the coordinator as relay, so if your coordinator is A, and you make a request, activations will flow in A -> B -> C -> back to A.

How it works and how to configure it

The prefill path is pipelined (this is why it can go faster than in a single machine). For large prompts the coordinator can run its slice on chunk N+1 while the worker is running its slice on chunk N. The distributed rows below were measured with two M5 Max 128 GB MacBooks connected by Thunderbolt 5, using the Q4 Flash GGUF and the default 4096-token distributed prefill chunk. The single-process column is a reference run with the Q2 GGUF on a single machine, so it actually is a bit faster since the routed MoEs are smaller.

Prompt Single-process reference Two MacBooks Speedup
9421 tokens 421.70 t/s 582.22 t/s 1.38x
28684 tokens 405.30 t/s 674.16 t/s 1.66x
63819 tokens 353.62 t/s 654.79 t/s 1.85x

Generation is different. It is strictly autoregressive: token N+1 cannot start until token N has produced logits and sampling has selected the next token. That means distributed generation cannot use the long prefill pipeline. It pays at least one cross-machine activation hop per generated token, so generation is slower than a single local process. On the same two-Mac Thunderbolt setup, a 12k-context control run with the 91 GB Flash quant went from 30.59 t/s single-process to 24.67 t/s distributed, a 19.4% loss. Distributed inference is therefore mainly for fitting larger models and speeding up long prefills, not for making decode faster.

Full DeepSeek V4 PRO Q4 on two Mac Studios

The full-size PRO Q4 GGUF can be run across two 512 GB Mac Studio M3 Ultra machines by giving the coordinator layers 0:30 and the worker 31:output. Use the split GGUF files so each side maps only the tensors it needs:

# Coordinator machine.
./download_model.sh pro-q4-layers00-30

# Worker machine.
./download_model.sh pro-q4-layers31-output

The two files are:

gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf
gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf

This is a capacity use case: each process maps only its own half of the model, while the worker owns the output head and returns logits.

The current PRO Q4 Metal path uses queue-resident exact expert tables for the large routed experts. This avoids the broad multi-GiB routed-tensor bindings that made early distributed PRO Q4 attempts either run very slowly or hit Metal memory accounting limits. In a short greedy smoke test over the direct 192.168.0.182 / 192.168.0.183 link, the model generated coherent text and measured 11.47 t/s generation after startup. Per-token telemetry was balanced: local layers were around 39-43 ms, remote layers around 44-49 ms, for total token times around 84-92 ms. Expect a slow startup while each side maps and makes its half of the model resident. Long-context PRO Q4 prefill and decode performance still needs separate benchmarking.

The measurements above use a Thunderbolt 5 cable. The implementation is plain TCP and also works over slower links, including WiFi, but fast Ethernet or Thunderbolt networking is strongly recommended. Slow links mostly hurt generation latency and short prefills; large prefills can still benefit when the layer split is balanced. In the normal performance path, the last worker owns the output head and returns logits directly.

Minimal two-host configuration:

# Machine A: coordinator, owns tokenization, sampling, the prompt, and layers 0..30.
./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf \
  --role coordinator \
  --layers 0:30 \
  --listen 169.254.43.68 1234

# Machine B: worker, connects to A and owns layers 31..output.
./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf \
  --role worker \
  --layers 31:output \
  --coordinator 169.254.43.68 1234

Normally the final worker should own the output head too, for example --layers 20:output. This avoids returning a full final hidden-state batch after prefill and lets the final worker produce the logits directly. On very slow or metered links, --layers 20:42 is also supported: the coordinator will load the output head and compute logits locally, trading extra coordinator work for smaller per-token replies.

Network Link Comparison

The table below shows the same two M5 Max hosts, the same 91 GB Flash quant, coordinator --layers 0:19, worker --layers 20:output, an 8192-token prompt from speed-bench/promessi_sposi.txt, and 128 generated tokens. WiFi and Internet numbers vary with local conditions, but the shape is the important part: high latency hurts generation directly, while lower bandwidth also pulls down long-prefill speed.

Link Addresses Ping avg Prefill Generation
Thunderbolt 5 169.254.43.68 -> 169.254.12.245 0.45 ms 582.99 t/s 25.09 t/s
WiFi 192.168.1.57 -> 192.168.1.95 77.20 ms 250.70 t/s 10.70 t/s
Internet / VPN 10.77.0.4 -> 10.77.0.3 152.10 ms 114.88 t/s 3.63 t/s

The Internet/VPN case is not meant to be a good interactive experience. It is still useful for collective testing: multiple people can temporarily combine machines to run a larger model that would not fit on any single host, accepting slow decode in exchange for being able to inspect the model at all.

Use the coordinator exactly like normal ./ds4: interactive chat, /read, and ordinary generation go through the same high-level session API. The same distributed options are also wired into ds4-agent, ds4-eval, and ds4-bench. For benchmarks, workers should already be running; ds4-bench waits until a complete route is available.

Useful tuning and diagnostics:

./ds4-bench \
  -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
  --prompt-file speed-bench/promessi_sposi.txt \
  --ctx-start 32768 \
  --ctx-max 65536 \
  --step-incr 32768 \
  --gen-tokens 0 \
  --role coordinator \
  --layers 0:19 \
  --listen 169.254.43.68 1234 \
  --debug

--debug on the coordinator prints route formation and per-hop telemetry: layer range, token span, local evaluation time, downstream wait time, socket send time, and input/output byte counts. This is the current profiling tool for deciding whether a split is balanced. --dist-prefill-window N controls how many prefill chunks may be in flight end-to-end; the default is conservative and bounded. --dist-prefill-chunk N exists for experiments, but the default 4096-token chunk is the canonical setting and should be used unless you are explicitly validating a different chunk size.

By default DwarfStar sends hidden-state activations as 32-bit floats. To reduce traffic, pass --dist-activation-bits 16 or --dist-activation-bits 8 on the coordinator. This changes only the transport format between machines, not the model weights or KV cache. 16-bit transport halves activation traffic and is the first option to try on Ethernet or WiFi. 8-bit transport is more aggressive and should be treated as an approximate/experimental mode unless you have validated the output for your use case. However experimentally reduction activation size didn't provide a significant improvement, so this option may be removed in the future.

If a worker disconnects, the coordinator removes that worker from the active route. The request already in flight can fail, and later calls report an incomplete route until a compatible worker reconnects and sends a new registration. For live sessions, the coordinator keeps the token history and can rebuild worker KV state by replaying the prefix when the route is available again. Workers also validate a rolling 64-bit token-prefix hash on every work item, so a restarted worker at position 0 cannot silently accept work for position N; it reports the mismatch and the coordinator replays the current transcript. Ctrl+C in the CLI and agent is cooperative: DwarfStar waits for the current distributed token or prefill chunk to drain before returning control, which avoids coordinator-caused KV splits. Saved agent/server sessions use the same KV file format as single-machine sessions: during save the coordinator fetches worker-owned layer tensors and serializes one normal payload; during load it splits that payload over the currently registered route.

Distributed protocol overview

At the protocol level there are two kinds of connections. Workers keep a control TCP connection open to the coordinator and send a HELLO with their model ID, model family, quant profile, layer slice, context capacity, and data port. The coordinator uses these registrations to build a route that covers all layers. Work then moves over low-latency TCP data connections: the coordinator computes the first slice, sends a WORK frame with session ID, token positions, rolling token-prefix hashes before and after the span, route information, and hidden-state payload, and each worker computes its slice. Middle workers can forward directly to the next worker. The final worker returns logits to the coordinator, or ACKs for non-final prefill chunks so the prefill pipeline can stay full. RESULT frames echo the request ID and the post-span hash. A worker status error is handled differently from a socket failure: KV/hash mismatch can be recovered by replaying the token history on the same route, while transport failure drops the route and waits for a replacement worker. For persistent KV, the coordinator opens worker data connections and sends snapshot save/load messages for each worker-owned layer range; the disk payload remains a single agent/server cache file. The protocol has no encryption or authentication, and is not release-stable yet; coordinator and workers should be built from the same commit and used on trusted machines and trusted networks.

Tensor Parallelism over RDMA

Tensor parallelism runs a single decode across two Macs connected with a Thunderbolt 5 cable, splitting the heavy per-layer work between the two GPUs and exchanging 16-24KB partial sums at synchronization gates inside the graph (RDMA over Thunderbolt when available, a dedicated TCP socket otherwise). Unlike the pipelined distributed mode above, both machines work on the same token at the same time, so it reduces per-token latency instead of just fitting a bigger model.

Each machine keeps one contiguous half of the routed experts resident. Dense, attention, shared-expert, embedding, and output weights remain replicated. This lets a model whose routed experts do not fit on one machine run fully resident across the pair; routed kernels never touch the peer's expert half.

Running GLM 5.2 across two 128 GB MacBooks

One-time setup per boot, on both machines:

# Let the GPU wire ~117 GB (default cap is ~75% of RAM; the resident
# expert shard needs ~97.5 GiB plus KV/scratch).
sudo sysctl iogpu.wired_limit_mb=120000

# RDMA over Thunderbolt needs an IPv4 address directly on the cabled
# member interface (the bridge IP does not count). Use the interface
# that is 'active' in ifconfig, e.g. en1 on one side and en6 on the
# other. Skip this if you are fine with the TCP fallback.
sudo ifconfig en1 inet 10.99.0.2/30 alias     # machine A
sudo ifconfig en6 inet 10.99.0.1/30 alias     # machine B

Check the verbs device before loading the model:

rdma_ctl status
ibv_devinfo -v

The device must be active and expose the IPv4-mapped GID for the address above, for example ::ffff:10.99.0.2. A working IP ping does not prove that RDMA is active.

Both machines need the same tree, commit, and GGUF path. Tensor parallelism is always a 50/50 split with one worker, so do not pass --layers. Start the worker first; it retries while the coordinator loads. The worker must dial the address on the Thunderbolt member interface, not the bridge address:

MODEL=gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf

# Machine B: worker.
./ds4 -m "$MODEL" --tensor-parallel --role worker \
  --coordinator 10.99.0.2 9911 --transport rdma

# Machine A: coordinator.
./ds4 -m "$MODEL" --tensor-parallel --role coordinator \
  --listen 10.99.0.2 9911 --transport rdma -c 8192 \
  -p "Tell me something about the sea."

The active verbs device and IPv4-mapped GID are selected automatically. If that is ambiguous, add --rdma-device rdma_en6 --rdma-gid-index 1 on the worker and the matching rdma_en1 flags on the coordinator. Use --transport tcp on both sides to force TCP. Tensor parallel roles are currently exposed by the ds4 CLI, not by ds4-server or ds4-agent.

Startup takes about 9 seconds per machine: each rank pre-faults its ~100 GiB shard from SSD and pins it through a Metal residency set. DeepSeek V4 Flash works the same way with its own GGUF on both machines. DeepSeek gate vectors are 16 KB and ride as one RDMA message. GLM's 6144-wide 24 KB vectors are split into two ordered RDMA messages.

Measured on two M5 Max 128 GB MacBooks (GLM 5.2, IQ2_XXS, 188 GiB):

two Macs, tensor parallel one Mac, SSD streaming
decode ~16.8 t/s (15.4 at 4k context) ~4.8 t/s
prefill (4096 tokens) ~94 t/s ~3-5 t/s
residency fully memory-resident streams experts from SSD

Notes: the coordinator mirrors every prompt sync and eval to the worker, so both KV caches stay in lockstep; prompt processing splits both the routed-expert GEMMs (by expert ownership) and the attention heads (a contiguous half per machine) with one bulk partial-sum exchange per layer per stage (--tensor-parallel-token-prefill selects a slower token-by-token prefill that exactly matches the single-machine arithmetic). The split graph is deterministic, but its changed floating-point reduction order is not generally byte-identical to single-machine execution.

Tensor Parallelism across CUDA GPUs

On a single CUDA server, --cuda-tensor-parallel splits DeepSeek V4 Flash tensor and routed-expert work across an even number of GPUs. This is separate from the Mac-to-Mac mode above: it does not use --role, RDMA, or the distributed layer pipeline. GPU placement and memory budgets are selected with the normal --gpu-devices and --gpu-vram options.

The device order is significant. With N devices, the first N/2 logical tiers are contiguous layer-pipeline homes and the second N/2 tiers are their tensor-parallel partners. Specify all homes first and then all partners, with the closest P2P pair at matching positions. For example, the tested L40S host uses physical pairs (0,1), (2,3), (4,5), and (6,7), expressed as 0,2,4,6,1,3,5,7. Each pair stores a 50/50 split of the routed experts, and the vocabulary head is row-sharded across the participating output tiers. Those large tensors are not duplicated. Dense attention, router, and shared expert weights are replicated within each pair.

For maximum throughput on eight 48 GB L40S cards, use the imatrix Q4 model. Its routed Q4_K layout has the native grouped multi-session kernels; the Q2 model is the lower-memory choice (including tested four-card runs), but its unsupported grouped routed shapes use the exact fallback and have lower aggregate serving throughput. Download and build the L40S target with:

./download_model.sh q4-imatrix
make cuda CUDA_ARCH=sm_89

This is the interactive-agent setup used on the eight-L40S server:

MODEL=gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf

./ds4-agent --cuda --cuda-tensor-parallel \
  --gpu-vram auto \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --model "$MODEL" \
  --ctx 100000

For serving, keep multiple KV sessions resident so decode rows can be grouped across requests. The tested host is configured for up to 16 resident sessions:

./ds4-server --cuda --cuda-tensor-parallel \
  --gpu-vram auto \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --model "$MODEL" \
  --ctx 100000 \
  --batched-session 16 \
  --host 0.0.0.0

The equivalent local launchers are ./run-nvidia-tp-agent.sh and ./run-nvidia-tp-server.sh. The server launcher also enables the on-disk KV cache. Reduce the session count or context size if the requested resident KV caches do not fit after model loading. CUDA TP, half-resident expert ownership, output sharding, pipelined prefill, and compatible grouped decode are selected by --cuda-tensor-parallel; no DS4_CUDA_* environment tuning is required. Without an explicit --prefill-chunk, this mode uses 2048-token chunks so the tested 16-session, 100k-context layout retains enough VRAM for resident KV caches. An explicit --prefill-chunk remains an override for other topologies.

Any even card count that can hold the selected model and graph scratch is a valid topology. On this class of 48 GB card, the useful measured endpoints are Q2 on four cards (two pipeline stages) and Q4 on eight cards (four stages). For a four-card PIX-paired subset such as physical GPUs 0,1,4,5, the ordered list is 0,4,1,5. Two cards do not have enough memory for these Flash models.

This mode currently requires DeepSeek V4 Flash and an even multi-GPU placement. GLM 5.2 instead uses normal layer placement across the selected CUDA devices. DGX Spark is a single-GPU target and must not be started with --cuda-tensor-parallel.

Reducing heat, power usage and fan noise

Long local inference runs can keep the GPU busy for extended periods. If you care more about heat, fan noise, battery life on MacBooks, or reducing thermal stress on the hardware than about maximum throughput, use --power N.

--power 100 is the default and means full speed. Lower values ask DwarfStar to target that percentage of GPU usage: --power 70 targets about 70%, --power 50 targets about half usage, and so forth. DwarfStar does this by measuring GPU work time and inserting small sleeps between work units: during prefill it sleeps between layers, and during generation it sleeps between decoded tokens. This reduces sustained load without changing model output.

The option is available on the CLI, server, agent, eval, and benchmark tools for DeepSeek models. GLM 5.2 currently accepts only --power 100. For example:

./ds4 --power 50
./ds4-agent --power 70
./ds4-server --power 40 --ctx 100000

Native agent

DwarfStar features a native coding agent that works in a different way than most other systems: the inference is controlled from within the agent itself, without socket/API boundaries, so the session is represented by the on-disk KV cache itself. Moreover the tools and the system prompt are all designed vertically for DeepSeek v4 Flash and PRO. This provides a few advantages:

  • Low latency experience, bounded mainly by the prefill speed limits. Displaying of generated text, tool calling, start of a new session are always instantaneous.
  • Live progress bar during prefill time.
  • No DSML tool calling conversion, the tools are handled natively in the LLM format.
  • KV cache mismatch are impossible by construction, the current state is always the truth.
  • Everything is tuned for this model.
  • Ability to switch saved sessions with /list and /switch; full KV sessions resume without a prefill stage.

Agent sessions are stored in ~/.ds4/kvcache. Use /save to persist the current session, /list to show saved sessions sorted by recent update time, and /switch <sha> to resume one of them. The session ID is stable across future saves and is derived from the first user prompt and creation time. /del <sha> removes a saved session. /strip <sha> keeps the rendered conversation text and title but removes the heavy KV payload; switching to a stripped session rebuilds the KV cache by prefilling the saved text.

Use --chdir /path/to/ds4 when launching ds4-agent from another directory, so relative runtime files such as metal/*.metal resolve from the project tree.

However while the system already works, there is a lot of work to do in order to make it ready for prime time. When finally the agent will reach the wanted shape, we will likely split the server and the client creating a stateful session-based protocol that can recreate all that in a client-server way.

Benchmarking

ds4-bench measures instantaneous prefill and generation throughput at context frontiers instead of reporting one whole-run average. It loads the model once, walks a fixed token sequence to frontiers such as 2048, 4096, 6144, and uses incremental prefill so each row measures only the newly-added token interval. After each frontier it saves the live KV state to memory, generates a fixed greedy non-EOS probe, restores the memory snapshot, and continues prefill.

./ds4-bench \
  -m ds4flash.gguf \
  --prompt-file speed-bench/promessi_sposi.txt \
  --ctx-start 2048 \
  --ctx-max 65536 \
  --step-incr 2048 \
  --gen-tokens 128

The example file is a cleaned public-domain Project Gutenberg text of Alessandro Manzoni's I Promessi Sposi (ebook #45334), with the Gutenberg header and footer removed: https://www.gutenberg.org/ebooks/45334.

Use --step-incr N for different linear spacing, or --step-mul F for exponential sweeps. Output is CSV with one row per frontier: latest prefill interval tokens/sec, generation tokens/sec at that frontier, and kvcache_bytes.

Sessions prefill long prompts in 4096-token chunks by default. Use --prefill-chunk 2048, for example, to match the strict official-vector checkpoint path. Changing the chunk changes the KV checkpoint/logit path, so compare it as an explicit run configuration. Chunked Metal prefill reuses the same range-capable layer-major graph for each chunk, preserving absolute compressor/indexer boundaries while avoiding the old per-layer chunk dispatch path.

Capability Evaluation

ds4-eval is a small real-model integration benchmark. It is not a leaderboard runner and should not be reported as an official GPQA, SuperGPQA, AIME, or security benchmark score: the questions are an embedded 92-item subset chosen to make local regression testing useful and visually inspectable. The program loads the real GGUF, renders DeepSeek chat prompts, streams sampled tokens in a split-screen TUI, grades the final answer, and prints a per-question report with prompt tokens, generated tokens, pass/fail state, the model answer, and the correct answer.

./ds4-eval -m ds4flash.gguf --trace /tmp/ds4-eval.txt

The default run uses --tokens 16000, thinking mode enabled, and a soft/hard </think> budget cutoff so the model has room to produce a visible answer. ds4-eval sizes the context internally from the largest selected prompt plus the generation budget, and refuses runs that would need more than 1M context tokens. Press p to pause, q to exit and print the report, Up/Down to inspect or select another question, and Enter to run the selected question next. --plain disables the TUI.

Use --regrade-trace /path/to/trace.txt to replay the current answer extractor and scorer against a prior --trace file without loading the model or regenerating tokens. This is useful when auditing evaluator changes: it shows which cases changed, the old picked answer, the new picked answer, and a pass/fail summary.

For inference changes that can affect generation drift, keep this deterministic q1..q4 token-count gate in the test plan:

./ds4-eval \
  -m ds4flash.gguf \
  --plain \
  --questions 4 \
  --tokens 2048 \
  --temp 0 \
  --seed 1

The generated-token counts must stay aligned with the baseline:

Question Expected state Expected generated tokens Expected given/correct
1 PASSED 2048 B / B
2 PASSED 438 C / C
3 PASSED 666 70 / 70
4 FAILED 2048 A / C

The first 75 embedded questions are interleaved as 25 GPQA Diamond, 25 audited SuperGPQA, and 25 AIME 2025 problems. The final 17 are an audited COMPSEC subset of reduced single-function C/C++ vulnerability-localization questions. The model is asked for the single best source line, or the smallest exact line set only when the bug cannot be localized to one line; the scorer accepts small audited ranges only when adjacent lines are equivalent locations for the same bug. The order is intentionally progressive: early questions are useful smoke tests, while later questions are hard enough that a strong reasoning model should still miss some of them. The SuperGPQA slice is curated rather than blind: upstream rows with wrong keys, missing figures, or underspecified prompts are replaced with cleaner rows.

The set should be treated as a hard capability regression suite rather than a pass/fail unit test.

  • GPQA Diamond contributes graduate-level science questions with multiple-choice answers. DeepSeek's model card reports strong results on full GPQA Diamond in thinking mode, but individual items still require careful physics, chemistry, or biology reasoning and are easy to lose with a small prompt/rendering or sampling regression.
  • SuperGPQA contributes broad specialist knowledge and domain-transfer questions. The model-card SuperGPQA number is much lower than GPQA Diamond, so these items are expected to be uneven: some look mundane, others require niche professional knowledge or exact interpretation of a translated-style exam question.
  • AIME 2025 contributes exact-answer contest math. These are often the most unforgiving items in the set: no multiple-choice prior, no partial credit, and a single arithmetic or algebraic slip changes the grade.
  • COMPSEC contributes single-function C/C++ security reasoning items reduced from public CVE writeups. These are not exploit prompts: the task is to identify the best source line where the defensive code flaw is introduced, or return 0 for a safe function.

In practice this means ds4-eval should not be expected to produce a perfect 92/92 run. It is meant to answer a more useful engineering question: after a kernel, quantization, prompt-rendering, KV-cache, or tool-streaming change, does DeepSeek V4 Flash still solve a representative mix of hard science, broad knowledge, exact math, and security-code problems while using the same inference path users run?

CLI

One-shot prompt:

./ds4 -p "Explain Redis streams in one paragraph."

No -p starts the interactive prompt:

./ds4
ds4>

The interactive CLI is a real multi-turn chat. It keeps the rendered chat transcript and the live graph KV checkpoint, so each turn extends the previous conversation. Useful commands are /help, /think, /think-max, /nothink, /ctx N, /read FILE, and /quit. Ctrl+C interrupts the current generation and returns to ds4>.

The CLI defaults to thinking mode. Use /nothink or --nothink for direct answers. --mtp MTP.gguf --mtp-draft 2 enables the optional MTP speculative path; it is useful only for greedy decoding, currently uses a confidence gate (--mtp-margin) to avoid slow partial accepts, and should be treated as an experimental slight-speedup path.

Server

Start a local OpenAI/Anthropic-compatible server:

./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192

Use --chdir /path/to/ds4 when launching ds4-server from another directory, so relative runtime files such as metal/*.metal resolve from the project tree.

By default the server keeps one mutable backend/KV checkpoint in memory, so stateless clients that resend a longer version of the same prompt can reuse the shared prefix instead of pre-filling from token zero.

--batched-session N preallocates N independent resident KV sessions. Ready decode steps are evaluated together, while long prefills alternate in bounded chunks so one request does not block every decoder. Requests beyond N wait for a resident slot. If disk KV caching is enabled, an idle slot is persisted before reuse and can be restored when that conversation returns; an active request is never evicted. Choose N and --ctx so all resident KV allocations fit in GPU memory. Without this option, inference retains the original single-session behavior.

Batching is exact: when a native batched kernel is unavailable, DwarfStar runs the affected rows in a fixed order and returns the same full logits as separate session evaluations. The current backend behavior is:

Backend and model Session execution
Metal, resident DeepSeek Flash Native shared-expert and QKV batching from two rows upward when supported; ordered fallback otherwise.
Metal, GLM 5.2 Ordered exact fallback.
CUDA, DeepSeek Flash on a supported multi-GPU TP/EP layout Native decode and mixed prefill/decode, with exact fallbacks for unsupported kernel shapes.
CUDA single GPU, including DGX Spark Ordered exact fallback.

N resident sessions allocate N KV states, so a context size that fits once may not fit eight times. Native batching can improve aggregate throughput; an ordered fallback provides concurrency and fairness, but not the same speedup. MTP speculative decoding is disabled while native session batching is active.

Supported endpoints:

  • GET /v1/models
  • GET /v1/models/deepseek-v4-flash
  • GET /v1/models/deepseek-v4-pro
  • POST /v1/chat/completions
  • POST /v1/responses
  • POST /v1/completions
  • POST /v1/messages

The Flash and PRO model endpoints are compatibility aliases. They both report the model currently loaded from the GGUF passed with -m; the endpoint name does not select a different model.

/v1/chat/completions accepts the usual OpenAI-style messages, max_tokens/max_completion_tokens, temperature, top_p, top_k, min_p, seed, stream, stream_options.include_usage, tools, and tool_choice. Tool schemas are rendered into DeepSeek's DSML tool format, and generated DSML tool calls are mapped back to OpenAI tool calls.

/v1/responses accepts OpenAI Responses-style input, instructions, tools, tool_choice, max_output_tokens, temperature, top_p, stream, and reasoning. It is the preferred endpoint for Codex CLI. The server keeps Responses continuations bound to live state when possible, and can fall back to the same DSML rendering and KV prefix reuse used by chat completions.

/v1/messages is the Anthropic-compatible endpoint used by Claude Code style clients. It accepts system, messages, tools, tool_choice, max_tokens, temperature, top_p, top_k, stream, stop_sequences, and thinking controls. Tool uses are returned as Anthropic tool_use blocks.

Default sampled API generation uses temperature=1, top_p=1, and min_p=0.05, so the default filter is relative probability rather than nucleus mass. In thinking mode DwarfStar applies those fixed sampling defaults to any knob the request omits, matching DeepSeek's fixed-thinking API behavior, but sampling parameters set explicitly in the request always win: a temperature=0 request is greedy through the whole reasoning phase, so benchmark harnesses get deterministic thinking-mode output.

The chat, Responses, and Anthropic endpoints support SSE streaming. In thinking mode, reasoning is streamed in the native API shape instead of being mixed into final text. OpenAI chat streaming also streams tool calls as soon as the DSML invocation is recognized: the tool header is sent first, then parameter bytes are forwarded as tool_calls[].function.arguments deltas while generation continues. The Anthropic endpoint streams thinking and text live, then emits structured tool_use blocks when the generated tool block is complete. The Responses endpoint streams the Responses event lifecycle expected by Codex, including response.output_text.delta, function-call argument events, and terminal response.completed / response.incomplete / response.failed events.

For browser JavaScript clients served from another origin, start the server with --cors to emit Access-Control-Allow-* headers. This only changes HTTP headers; it does not expose the server on the LAN. Use --host 0.0.0.0 explicitly when remote machines should be able to connect.

Tool call handling and canonicalization

DeepSeek V4 emits tool calls as DSML text. Agent clients do not send that same text back on the next request: they send normalized OpenAI/Anthropic JSON tool-call objects. If the server re-rendered those objects slightly differently, the rendered byte prefix would no longer match the live KV checkpoint and the next turn would have to be rebuilt.

The first line of defense is exact replay. Every tool call gets an unguessable API tool ID, and the server remembers tool id -> exact sampled DSML block in a bounded in-memory map backed by radix trees. When the client later sends that tool ID back, the prompt renderer uses the exact DSML bytes the model sampled, not a freshly formatted approximation. This map can also be saved inside KV cache files, so exact replay survives server restarts for cached histories.

Canonicalization is only the backup path. If the exact DSML block is missing, or exact replay is disabled with --disable-exact-dsml-tool-replay, the server renders a deterministic DSML form from the JSON tool object. After a tool-call turn, it compares the live sampled token stream with the prompt that the next client request will render. If needed, it rewrites the live checkpoint, or falls back to an older disk KV snapshot and replays only the suffix. This keeps the model continuation aligned with the stateless API transcript.

During generation, the server also treats DSML syntax differently from payload. When the model is emitting stable protocol structure such as DSML tags, parameter headers, JSON punctuation, or closing markers, sampling is forced to temperature=0 so the tool call stays parseable. This greedy mode does not apply to argument payloads: string=true parameter bodies and JSON string values, including file contents and edit text, use the request's normal sampling settings. That separation is important: deterministic decoding is helpful for syntax, but can create repeated text when applied to long code or file bodies.

Minimal OpenAI example:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model":"deepseek-v4-flash",
    "messages":[{"role":"user","content":"List three Redis design principles."}],
    "stream":true
  }'

Agent Client Usage

ds4-server can be used by local coding agents that speak OpenAI-compatible chat completions. Start the server first, and set the client context limit no higher than the --ctx value you started the server with:

./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192

You can use larger context and larger cache if you wish. Full context of 1M tokens is going to use more or less 26GB of memory (compressed indexer alone will be like 22GB), so configure a context which makes sense in your system. With 128GB of RAM you would run the 2-bit quants, which are already 81GB, 26GB are going to be likely too much, so a context window of 100~300k tokens is wiser. However users reported being able to run 2bit quants with 250k ctx window in a Macs with just 96GB of system memory: make sure to kill processes that use too much memory, if you plan doing so ;)

The 384000 output limit below avoids token caps since the model is able to generate very long replies otherwise (up to 384k tokens). The server still stops when the configured context window is full.

For opencode, add a provider and agent entry to ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ds4": {
      "name": "ds4.c (local)",
      "npm": "@ai-sdk/openai-compatible",
      "options": {
        "baseURL": "http://127.0.0.1:8000/v1",
        "apiKey": "dsv4-local"
      },
      "models": {
        "deepseek-v4-flash": {
          "name": "DeepSeek V4 Flash (ds4.c local)",
          "limit": {
            "context": 100000,
            "output": 384000
          }
        }
      }
    }
  },
  "agent": {
    "ds4": {
      "description": "DeepSeek V4 Flash served by local ds4-server",
      "model": "ds4/deepseek-v4-flash",
      "temperature": 0
    }
  }
}

For Pi, add a provider to ~/.pi/agent/models.json:

{
  "providers": {
    "ds4": {
      "name": "ds4.c local",
      "baseUrl": "http://127.0.0.1:8000/v1",
      "api": "openai-completions",
      "apiKey": "dsv4-local",
      "compat": {
        "supportsStore": false,
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": true,
        "supportsUsageInStreaming": true,
        "maxTokensField": "max_tokens",
        "supportsStrictMode": false,
        "thinkingFormat": "deepseek",
        "requiresReasoningContentOnAssistantMessages": true
      },
      "models": [
        {
          "id": "deepseek-v4-flash",
          "name": "DeepSeek V4 Flash (ds4.c local)",
          "reasoning": true,
          "thinkingLevelMap": {
            "off": null,
            "minimal": "low",
            "low": "low",
            "medium": "medium",
            "high": "high",
            "xhigh": "xhigh"
          },
          "input": ["text"],
          "contextWindow": 100000,
          "maxTokens": 384000,
          "cost": {
            "input": 0,
            "output": 0,
            "cacheRead": 0,
            "cacheWrite": 0
          }
        }
      ]
    }
  }
}

Optionally make it the default Pi model in ~/.pi/agent/settings.json:

{
  "defaultProvider": "ds4",
  "defaultModel": "deepseek-v4-flash"
}

For Codex CLI, use the Responses wire API:

[model_providers.ds4]
name = (README truncated)

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions
+1.4M-1.4MWeek of 2026-05-03: +65,870 linesWeek of 2026-05-03: -485 linesWeek of 2026-05-10: +798,981 linesWeek of 2026-05-10: -284,717 linesWeek of 2026-05-17: +31,164 linesWeek of 2026-05-17: -16,688 linesWeek of 2026-05-24: +18,909 linesWeek of 2026-05-24: -2,192 linesWeek of 2026-05-31: +36,983 linesWeek of 2026-05-31: -4,076 linesWeek of 2026-06-07: +20,420 linesWeek of 2026-06-07: -538 linesWeek of 2026-06-14: +8,222 linesWeek of 2026-06-14: -1,098 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +1,366,507 linesWeek of 2026-07-12: -39,004 linesWeek of 2026-07-19: +2,548 linesWeek of 2026-07-19: -697 linesWeek of 2026-07-26: +1,561 linesWeek of 2026-07-26: -1,248 linesMay 3, 2026Jul 26, 2026
+2.4M lines added, -350.7K removed over the last year.

Commits per week

last 52 weeks
920Week of 2025-08-03: 0 commitsWeek of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 23 commitsWeek of 2026-05-10: 92 commitsWeek of 2026-05-17: 58 commitsWeek of 2026-05-24: 59 commitsWeek of 2026-05-31: 22 commitsWeek of 2026-06-07: 26 commitsWeek of 2026-06-14: 12 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 10 commitsWeek of 2026-07-19: 23 commitsWeek of 2026-07-26: 16 commitsAug 3, 2025Jul 26, 2026
341 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 4 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 1 commitsSun 9:00 — 4 commitsSun 10:00 — 6 commitsSun 11:00 — 5 commitsSun 12:00 — 2 commitsSun 13:00 — 6 commitsSun 14:00 — 3 commitsSun 15:00 — 2 commitsSun 16:00 — 3 commitsSun 17:00 — 0 commitsSun 18:00 — 7 commitsSun 19:00 — 7 commitsSun 20:00 — 2 commitsSun 21:00 — 3 commitsSun 22:00 — 7 commitsSun 23:00 — 3 commitsMon 0:00 — 1 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 1 commitsMon 7:00 — 6 commitsMon 8:00 — 5 commitsMon 9:00 — 5 commitsMon 10:00 — 3 commitsMon 11:00 — 4 commitsMon 12:00 — 0 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 5 commitsMon 16:00 — 5 commitsMon 17:00 — 0 commitsMon 18:00 — 3 commitsMon 19:00 — 2 commitsMon 20:00 — 0 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 1 commitsTue 0:00 — 1 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 1 commitsTue 4:00 — 0 commitsTue 5:00 — 1 commitsTue 6:00 — 1 commitsTue 7:00 — 1 commitsTue 8:00 — 3 commitsTue 9:00 — 2 commitsTue 10:00 — 3 commitsTue 11:00 — 8 commitsTue 12:00 — 0 commitsTue 13:00 — 4 commitsTue 14:00 — 0 commitsTue 15:00 — 2 commitsTue 16:00 — 4 commitsTue 17:00 — 3 commitsTue 18:00 — 1 commitsTue 19:00 — 0 commitsTue 20:00 — 2 commitsTue 21:00 — 0 commitsTue 22:00 — 1 commitsTue 23:00 — 1 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 1 commitsWed 3:00 — 1 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 2 commitsWed 7:00 — 0 commitsWed 8:00 — 4 commitsWed 9:00 — 2 commitsWed 10:00 — 0 commitsWed 11:00 — 2 commitsWed 12:00 — 5 commitsWed 13:00 — 3 commitsWed 14:00 — 2 commitsWed 15:00 — 2 commitsWed 16:00 — 4 commitsWed 17:00 — 2 commitsWed 18:00 — 0 commitsWed 19:00 — 5 commitsWed 20:00 — 5 commitsWed 21:00 — 2 commitsWed 22:00 — 0 commitsWed 23:00 — 4 commitsThu 0:00 — 2 commitsThu 1:00 — 1 commitsThu 2:00 — 0 commitsThu 3:00 — 2 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 2 commitsThu 9:00 — 2 commitsThu 10:00 — 3 commitsThu 11:00 — 1 commitsThu 12:00 — 4 commitsThu 13:00 — 1 commitsThu 14:00 — 0 commitsThu 15:00 — 1 commitsThu 16:00 — 3 commitsThu 17:00 — 2 commitsThu 18:00 — 3 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 2 commitsThu 22:00 — 2 commitsThu 23:00 — 3 commitsFri 0:00 — 2 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 3 commitsFri 9:00 — 1 commitsFri 10:00 — 1 commitsFri 11:00 — 3 commitsFri 12:00 — 1 commitsFri 13:00 — 3 commitsFri 14:00 — 1 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 0 commitsFri 18:00 — 2 commitsFri 19:00 — 5 commitsFri 20:00 — 2 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 3 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 1 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 3 commitsSat 9:00 — 3 commitsSat 10:00 — 11 commitsSat 11:00 — 10 commitsSat 12:00 — 7 commitsSat 13:00 — 6 commitsSat 14:00 — 1 commitsSat 15:00 — 8 commitsSat 16:00 — 0 commitsSat 17:00 — 3 commitsSat 18:00 — 13 commitsSat 19:00 — 3 commitsSat 20:00 — 2 commitsSat 21:00 — 2 commitsSat 22:00 — 5 commitsSat 23:00 — 5 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits248 (71%)
Community commits99 (29%)

347 commits in total over the last year.

DateListRankStars gained
May 15, 2026daily#20+64
May 12, 2026daily#24+44
May 11, 2026daily#11+145
May 10, 2026daily#1+415
May 9, 2026daily#3+268
May 8, 2026daily#8+112
  • Genymobile/scrcpy

    Display and control your Android device

    147.1K stars · C

  • microsoft/PowerToys

    Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

    137.6K stars · C

  • colbymchenry/codegraph

    Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

    65.2K stars · C

  • DeusData/codebase-memory-mcp

    High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

    37.8K stars · C

  • JustVugg/colibri

    Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

    23K stars · C

  • erincatto/box3d

    Box3D is a 3D physics engine for games

    5.9K stars · C