The Creator of Redis Built an Engine to Run Frontier Models at Home
Salvatore Sanfilippo, the creator of Redis, released DwarfStar 4 (ds4), an MIT-licensed local inference engine purpose-built for huge routed MoE models like DeepSeek V4 on a 128GB Mac, using asymmetric 2-bit quantization and disk-based KV caching.
Salvatore Sanfilippo, the creator of Redis, has released DwarfStar 4 (ds4), an MIT-licensed local inference engine with one narrow obsession: running frontier-scale routed-MoE models, DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next, on a 128GB machine on your desk. It is written in C with three backends (Metal, CUDA, ROCm), ships a CLI, OpenAI and Anthropic-compatible local API servers, and a coding agent in one stack, and the benchmarks on a 128GB M5 Max are the proof: 790 tokens per second prefill at 2,048-token context, and 398 tokens per second prefill at 65,536-token context, for a 284-billion-parameter model with 2-bit quantized experts.
The Two Ideas That Make It Work
ds4’s approach rests on two engineering decisions that generic local runners have not taken. First, asymmetric 2-bit quantization: instead of compressing the whole model uniformly, ds4 compresses the routed experts in MoE architectures aggressively while keeping the critical shared paths precise, with importance-matrix (imatrix) loading guiding what gets sacrificed. Second, disk-based KV caching: long conversation prefixes are saved to SSD and resumed by prompt hash, so restarting a session avoids full re-prefill entirely. That second one is what makes the engine practical for agents rather than chat toys; always-on agent architectures live and die on how much state they can afford to keep warm, and ds4 makes the disk do the work RAM cannot.
Not a GGUF Runner, on Purpose
The design philosophy is the opposite of the llama.cpp ecosystem’s, and Sanfilippo is explicit about it: ds4 is “not a generic GGUF runner,” it is “narrow on purpose,” supporting only a small set of model families whose layouts are validated end to end against official outputs. Generic GGUF files are explicitly unsupported. This is the Redis lesson applied to inference: correctness and speed come from owning the whole layout rather than being compatible with everything, and for the specific goal of running a frontier MoE on unified-memory hardware, validation against official outputs is worth more than broad format support. Hardware targets are Apple Silicon (64GB and up, to a 512GB Mac Studio for V4.1 at 4-bit), NVIDIA DGX Spark and CUDA Linux, and AMD Strix Halo via ROCm, which covers essentially the entire local-AI hardware landscape of 2026.
Why This Matters Beyond the Hacker News Points
The post hit the front page with strong engagement, and the significance is bigger than one clever engine. The capability to run DeepSeek V4-class models locally at usable speeds (27.6 tokens per second generation at 64k context on a laptop-class machine) is the privacy and sovereignty story of local AI finally arriving at the frontier tier: nothing leaves the machine, no API terms apply, and the model can be audited. Combined with Kolibri’s open-weight sovereignty play landing the same week, the local-first ecosystem now has both European jurisdiction-safe weights and an engine to run huge models on consumer hardware. Sanfilippo’s track record with performance-critical infrastructure is the reason to take the benchmarks seriously, and the MIT license means the techniques (asymmetric quantization, hash-resumable KV caching) are now public property.
What to Watch
Three things. First, independent benchmark replication: the numbers come from Sanfilippo’s own hardware, and the community will have M5 Max and DGX Spark results within days. Second, model-family expansion: which MoE layouts get validated next is the roadmap signal. Third, the quality question that 2-bit quantization always raises: ds4 validates outputs end to end against official releases, and published perplexity or task-quality comparisons between q2 ds4 and full-precision API outputs will show exactly how much capability the compression trades away. If the answer is “not much,” the calculus of who needs an API changes for a meaningful class of workloads.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
What Is Quantization in AI?
Quantization shrinks AI models by reducing numerical precision. Here's how it works, what formats exist, and how to choose the right tradeoff between size, speed, and quality.
DeepSeek's 242k-Star Agent Harness Reaches the Desktop
DeepSeek Harness v0.2 preview shipped September 29 with official macOS and Windows desktop apps, in-app plugin management without Node or pnpm, and scheduled automations, maturing the MIT-licensed framework that has drawn 242k GitHub stars.
Livenerf: A Pre-Registered Benchmark Built to Catch Silent Model Downgrades
A 780-point Hacker News project called Livenerf runs a daily 78-question panel against Opus 5.5 through Claude Code to detect quiet post-launch degradation, with pre-registered decision rules, a control arm, and a first verdict expected around October 24.
Alibaba Open-Sources RADAR, a Medical AI That Beat Most Radiologists on Abdominal CT Scans
Alibaba's Damo Academy open-sourced RADAR, a medical AI trained on 420,000+ CT exams that detects nearly 150 abdominal conditions with an AUC of 0.913, outperforming 23 of 26 radiologists in comparative testing.
Anthropic Names DeepSeek, Moonshot, and MiniMax in Industrial-Scale Claude Distillation Campaigns
Anthropic's September threat intelligence report details seven harm areas including industrial-scale distillation campaigns by DeepSeek, Moonshot AI, and MiniMax using roughly 24,000 fraudulent accounts to extract capabilities from over 16 million Claude exchanges.