Whistle: Speech-to-Text in 16.9 MB, Faster Than Whisper at 8.6 Times Smaller
Cactus Compute released Whistle on October 2, a 16.9MB open speech recognition model covering seven languages with word timestamps, that beats Whisper base on most benchmarks while decoding at 1,319 tokens per second on an Apple M4 Pro CPU.
Cactus Compute released Whistle on October 2, and the spec sheet reads like a dare: a complete speech recognition model in a single 16.9MB file, covering seven languages with automatic detection, word-level timestamps, speech embeddings, and an 11.1-millisecond time-to-first-token on an Apple M4 Pro CPU. That last number is 6.6 times faster than Whisper base at 8.6 times smaller, and on the LibriSpeech test sets, SPGISpeech, Earnings-22, and the FLEURS average, the 16.9MB model out-accurates the 145.3MB Whisper base. It hit 717 points on Hacker News within two days.
The Architecture: Small by Composition, Not by Trimming
Whistle is small because it was built small, not because a bigger model was pruned. The encoder runs a log-mel frontend through a convolution stem that compresses 3,000 frames to 375, then eight Simple Attention blocks with Monarch Hadamard MLPs. The decoder uses eight laddered blocks with GQA and gated cross-attention reading the encoder, with an 8,192-piece vocabulary. The most unusual feature is laddered depth: every decoder depth from two layers up was trained as its own model, so --audio-depth selects a smaller or larger variant at load time, one checkpoint serving an entire size spectrum from watch to phone. The engine is pure C++ with zero dependencies, prebuilt for 17 targets including Android, iOS, watchOS, RISC-V, MIPS, and browser WASM, and it shares its compute engine with Cactus’s small text model Needle, which means one binary can transcribe a clip and chain it straight into a tool call: “turn off the kitchen lights” becomes a set_lights function call with 0.94 confidence.
The Benchmark Discipline Is the Credibility
The evaluation is unusually careful for an open release: 86,174 utterances across the benchmarks, no test-set contamination (verified by audio checksums and speaker IDs), and losses disclosed alongside wins, with Whisper base still ahead on TED-LIUM, AMI, and MLS. Silence detection skips beam search entirely where no speech exists, and 5-beam search supports keyword biasing through an Aho-Corasick automaton, which is the feature that makes it usable as a wake-word-and-command engine rather than just a transcriber. Accuracy measured over 86,174 utterances with contamination checks is the difference between a demo and a component.
Why 16.9MB Is a Line
The significance is the hardware tier it opens. Whistle targets mobiles, wearables, robots, smart home, automotive, and microcontrollers, and at 16.9MB with a pure-CPU engine it runs where no cloud-dependent transcription can: offline, battery-powered, latency-critical. The local-inference wave this week has been about frontier models on gaming PCs; Whistle is the same wave at the other end of the size spectrum, and together they bracket the stack: 284B-parameter models on desks, speech recognition on watches and RISC-V boards, nothing in between requiring an API key. Privacy is the structural beneficiary: audio never leaves the device, which removes the data-flow questions that this month’s chatbot-tracking research put back on the agenda.
What to Watch
Three things. First, the license and weights, which the post describes as open with public source and weights on Hugging Face (Cactus-Compute/whistle) but does not pin to a specific license name; the Apache-versus-research distinction decided Qwen Image 2.1’s fate and will decide Whistle’s adoption. Second, multilingual expansion beyond the current seven languages, which determines whether this is a European-market tool or a global one. Third, the combined stack: Needle plus Whistle in one binary implies a complete voice-agent-on-microcontroller reference design, and the first product shipping that stack will mark the moment on-device AI stopped being a phone feature and became an everything feature.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Run Gemma 4 E2B on Raspberry Pi 5 with LiteRT
Deploy the Gemma 4 E2B model locally on a Raspberry Pi 5 using the LiteRT-LM runtime for real-time edge workflows and robotics.
Nothing OS 4.1 Adds On-Device Voice Dictation
Nothing released Essential Voice, an on-device AI dictation tool in Nothing OS 4.1 that removes filler words, translates languages, and applies formatting.
Google AI Edge Eloquent brings free offline dictation to iOS
Google's new AI Edge Eloquent app uses Gemma 4 models to offer high-quality, offline-first transcription and text polishing for free on iPhone.
Voxtral TTS: Mistral's Open-Source Answer to Voice Agents
Mistral’s reported Voxtral TTS release could help developers build low-latency, open-source voice apps and agents on edge devices.
Strata Runs a 125B Model on a Gaming PC at 94 Tokens Per Second
Strata, an MIT-licensed inference engine that hit 746 points on Hacker News, runs Qwen3.8-Flash-Next (125B MoE) on consumer RTX and Radeon cards via expert offloading and speculative decoding, with one-click installers and localhost APIs.