
From the creator of Redis: ds4 runs frontier open models locally
DwarfStar 4, written in C by Salvatore Sanfilippo, the creator of Redis, fits a 284-billion-parameter DeepSeek mixture-of-experts model into 96-128 GB machines via asymmetric 2-bit quantization and a disk-persisted KV cache.
DwarfStar 4 (ds4) is a local inference engine written in C by Salvatore Sanfilippo, the developer known as antirez who created Redis. The project ships under the MIT license on GitHub and runs frontier open-weight models on high-memory Apple Silicon Macs and on NVIDIA CUDA and AMD ROCm machines, without a cloud service in the loop.
ds4 is deliberately narrow rather than a universal GGUF runner: it supports a small set of model families. The list includes DeepSeek V4 Flash and V4.1 Flash, GLM 5.x, Qwen3.8 Flash Next and their vision variants; DeepSeek V4 PRO also runs on very high-memory systems. According to the project's blog, a September update of 165 commits added DeepSeek V4.1 Flash, Qwen3.8 Flash Next, broader vision support and model-specific server batching. On GitHub the project has drawn more than 22,000 stars.
Fitting a 284-billion-parameter model onto one machine
The headline model, DeepSeek V4 Flash, is a 284-billion-parameter mixture-of-experts design. ds4 places it on machines with 96 to 128 GB of unified memory through asymmetric 2-bit quantization: the routed experts that hold most of the parameters are compressed to Q2 formats, while shared experts and other decision-critical tensors keep high precision. The quantization is tuned with imatrix data and checked against reference outputs.
SSD streaming and a KV cache that survives restarts
Models larger than RAM stream their weights from SSD, so even a 128 GB machine can run builds that would not otherwise fit. The KV cache is treated as durable data: entries are keyed by a hash of the rendered prompt, stored on disk and survive restarts, so long agent sessions stop paying the full prefill cost again. Tool calls get a deterministic replay map that persists with the cache.
Hardware, interfaces and measured speed
ds4 offers three interfaces: the chat CLI, a server with OpenAI- and Anthropic-compatible endpoints, and a native coding agent. The server works with Claude Code, Codex CLI and OpenCode. On an M5 Max with 128 GB, the project reports 39.4 tokens/s of generation and 790 tokens/s of prefill at a 2,048-token context; a DGX Spark reports 18.1 and 825.8 for the same row. A CUDA micro-batching mode turns servers of older Ada Lovelace cards into multi-user LLM hosts: eight L40S GPUs measured 120 tokens/s of aggregate generation and 2,000 tokens/s of prefill.
The project credits llama.cpp and GGML as its foundation. Its thesis is the opposite of a general-purpose runner: a small, validated stack that makes frontier-class open weights usable on hardware people already own.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.