Ai Engineering 3 min read

Whistle: Speech-to-Text in 16.9 MB, Faster Than Whisper at 8.6 Times Smaller

Cactus Compute released Whistle on October 2, a 16.9MB open speech recognition model covering seven languages with word timestamps, that beats Whisper base on most benchmarks while decoding at 1,319 tokens per second on an Apple M4 Pro CPU.

Cactus Compute released Whistle on October 2, and the spec sheet reads like a dare: a complete speech recognition model in a single 16.9MB file, covering seven languages with automatic detection, word-level timestamps, speech embeddings, and an 11.1-millisecond time-to-first-token on an Apple M4 Pro CPU. That last number is 6.6 times faster than Whisper base at 8.6 times smaller, and on the LibriSpeech test sets, SPGISpeech, Earnings-22, and the FLEURS average, the 16.9MB model out-accurates the 145.3MB Whisper base. It hit 717 points on Hacker News within two days.

The Architecture: Small by Composition, Not by Trimming

Whistle is small because it was built small, not because a bigger model was pruned. The encoder runs a log-mel frontend through a convolution stem that compresses 3,000 frames to 375, then eight Simple Attention blocks with Monarch Hadamard MLPs. The decoder uses eight laddered blocks with GQA and gated cross-attention reading the encoder, with an 8,192-piece vocabulary. The most unusual feature is laddered depth: every decoder depth from two layers up was trained as its own model, so --audio-depth selects a smaller or larger variant at load time, one checkpoint serving an entire size spectrum from watch to phone. The engine is pure C++ with zero dependencies, prebuilt for 17 targets including Android, iOS, watchOS, RISC-V, MIPS, and browser WASM, and it shares its compute engine with Cactus’s small text model Needle, which means one binary can transcribe a clip and chain it straight into a tool call: “turn off the kitchen lights” becomes a set_lights function call with 0.94 confidence.

The Benchmark Discipline Is the Credibility

The evaluation is unusually careful for an open release: 86,174 utterances across the benchmarks, no test-set contamination (verified by audio checksums and speaker IDs), and losses disclosed alongside wins, with Whisper base still ahead on TED-LIUM, AMI, and MLS. Silence detection skips beam search entirely where no speech exists, and 5-beam search supports keyword biasing through an Aho-Corasick automaton, which is the feature that makes it usable as a wake-word-and-command engine rather than just a transcriber. Accuracy measured over 86,174 utterances with contamination checks is the difference between a demo and a component.

Why 16.9MB Is a Line

The significance is the hardware tier it opens. Whistle targets mobiles, wearables, robots, smart home, automotive, and microcontrollers, and at 16.9MB with a pure-CPU engine it runs where no cloud-dependent transcription can: offline, battery-powered, latency-critical. The local-inference wave this week has been about frontier models on gaming PCs; Whistle is the same wave at the other end of the size spectrum, and together they bracket the stack: 284B-parameter models on desks, speech recognition on watches and RISC-V boards, nothing in between requiring an API key. Privacy is the structural beneficiary: audio never leaves the device, which removes the data-flow questions that this month’s chatbot-tracking research put back on the agenda.

What to Watch

Three things. First, the license and weights, which the post describes as open with public source and weights on Hugging Face (Cactus-Compute/whistle) but does not pin to a specific license name; the Apache-versus-research distinction decided Qwen Image 2.1’s fate and will decide Whistle’s adoption. Second, multilingual expansion beyond the current seven languages, which determines whether this is a European-market tool or a global one. Third, the combined stack: Needle plus Whistle in one binary implies a complete voice-agent-on-microcontroller reference design, and the first product shipping that stack will mark the moment on-device AI stopped being a phone feature and became an everything feature.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading