Ai Engineering 4 min read

Strata Runs a 125B Model on a Gaming PC at 94 Tokens Per Second

Strata, an MIT-licensed inference engine that hit 746 points on Hacker News, runs Qwen3.8-Flash-Next (125B MoE) on consumer RTX and Radeon cards via expert offloading and speculative decoding, with one-click installers and localhost APIs.

The local-inference wave produced its second headline in three days: Strata, an MIT-licensed engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, on consumer gaming PCs, with one-click Windows and Linux installers and 746 points on Hacker News. The measured numbers: 94 tokens per second generation and 2,650 tokens per second prompt processing on an RTX 5070 with 12GB of VRAM (Q2_0 quantization), 60 tokens per second on an RX 9070 XT, and a projected 100 to 140 tokens per second on an RTX 3090’s 24GB. Where ds4 targeted Apple Silicon and DGX-class unified memory, Strata’s target is the machine most developers already own: a gaming PC with a mid-range GPU.

The Technique: Stop Pretending You Need All the Experts

Strata’s core insight is about what MoE architectures actually load. Qwen3.8-Flash-Next has 24,576 experts, but each token activates only 10, which means the overwhelming majority of parameters are idle for any given token. Strata exploits that gap with expert offloading: frequently used experts live in VRAM, the full set lives in system RAM with the CPU processing overflow, and an SSD holds the expert lookup table. Layered on top is speculative decoding, a small draft model proposes tokens and the big model verifies them in parallel, for a 1.6 to 1.8 times speedup, plus GGUF quantizations from Q2_0 to roughly 4-bit. Requirements are 12GB-plus of VRAM and 32GB of RAM (64 recommended), with 35 to 55GB loading into RAM at startup, which the README warns briefly freezes the PC.

The Honest Numbers Are the Credibility

The README publishes the awkward cases alongside the flattering ones, which is why the project earns its 11.9k stars. Long first prompts process at about one minute per 30,000 tokens. The largest quants mostly stream from SSD and drop to 7 to 8.5 tokens per second on 64GB machines. The coder variant is weaker on Chinese and CJK text. And the headline “RTX 4090 at 100T/s” from the viral title is explicitly a projection, not a measurement; the verified numbers are the 5070 and 9070 XT results above. That distinction matters less than it seems, since 94 tokens per second on a 12GB card is already conversational speed for a 125B model, but it is the kind of title-vs-table gap readers should check.

Local Frontier AI Now Has Two Working Architectures

The significance is what Strata and ds4 prove together. ds4’s approach keeps shared paths precise, crushes routed experts to 2-bit, and streams KV cache from SSD on Apple Silicon and DGX hardware. Strata keeps experts in RAM, offloads by frequency, and uses speculative decoding on commodity NVIDIA and AMD gaming hardware. Different hardware, different techniques, same conclusion: the barrier to running a frontier-class MoE locally is no longer VRAM capacity, it is engineering. Both are MIT licensed, both expose OpenAI-compatible local APIs, and both arrived within three days of each other. The 2023 stack diagram, where capable models meant API keys and per-token billing, is now optional for anyone with 32GB of RAM and a gaming GPU.

What to Watch

Three things. First, whether Strata and ds4 converge, since Strata’s disk-based expert lookup and ds4’s SSD-resident KV cache are complementary techniques pointing at the same hybrid RAM-SSD future for local inference. Second, output quality at Q2_0: the speed numbers are real, but 2-bit expert quantization on a 125B model trades capability, and per-task quality comparisons against full-precision API outputs are the missing table. Third, the ecosystem effect: Strata exposing an MCP server and Anthropic-compatible APIs locally means agent frameworks can treat a gaming PC as a model provider, and local inference plus documentation-based agent memory starts to look like a complete self-hosted agent stack.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

What Is Quantization in AI?

Quantization shrinks AI models by reducing numerical precision. Here's how it works, what formats exist, and how to choose the right tradeoff between size, speed, and quality.

Ai Engineering

The Creator of Redis Built an Engine to Run Frontier Models at Home

Salvatore Sanfilippo, the creator of Redis, released DwarfStar 4 (ds4), an MIT-licensed local inference engine purpose-built for huge routed MoE models like DeepSeek V4 on a 128GB Mac, using asymmetric 2-bit quantization and disk-based KV caching.

Ai Engineering

Livenerf: A Pre-Registered Benchmark Built to Catch Silent Model Downgrades

A 780-point Hacker News project called Livenerf runs a daily 78-question panel against Opus 5.5 through Claude Code to detect quiet post-launch degradation, with pre-registered decision rules, a control arm, and a first verdict expected around October 24.

Ai Engineering

Xiaomi's MiMo v2.6 Takes the Top Open-Weights Spot on the Intelligence Index

Xiaomi released MiMo v2.6 on September 21, a 1-trillion-parameter open-weights MoE model that debuts at number one among open models on Artificial Analysis' Intelligence Index, with an MIT license and aggressive API pricing.

Ai Engineering

Qwen Image 2.1 Ships Native Transparency and 10-Reference Editing in a 7B Open-Weights Model

Alibaba's Qwen team released Qwen Image 2.1, a 7B open-weight model that unifies image generation and editing, outputs native RGBA transparency at 2K, and runs on consumer GPUs. The catch is a research-only license.