Whistle — 16.9 MB on-device speech recognition
A tiny speech-to-text model (single 16.9 MB file) for phones, wearables, robots, and microcontrollers: transcription, word timestamps, and speech embeddings, all on-device with no GPU or dependencies.
Why it matters
Full ASR in a 16.9 MB file that shares the Needle engine's SIMD kernels and KV cache, with keyword biasing for names/products and benchmark WER measured over 86k utterances against Whisper and Moonshine.
What you could build with it
A hardware startup building a voice-controlled smart-home device could ship Whistle on a microcontroller-class chip, doing wake-word-adjacent keyword-biased transcription locally with no cloud round-trip, no per-minute API cost, and no privacy exposure.
Does it hold up?
Promising but early: the model card's benchmarks are thorough and self-scored against published Whisper/Moonshine numbers, and the Cactus engine has real adoption — but no third-party device-integration reports for Whistle itself yet.
Built with Whistle
- Cactus Compute releases a 16.9MB speech model for local CPUs (RuntimeWire deep dive)runtimewire · Third-party launch coverage: walks through Whistle's Needle-runtime integration, keyword biasing and word timestamps, with the release video from Cactus's X post.
- Cactus fits speech transcription into a 16.9 MB model (alextech.ai)alextech.ai · Writeup of the Whistle collection on HF: variable-depth decoder, loading alongside Needle for speech-to-tool-call pipelines, 11ms first-token latency claim.
- Whistle: This Tiny 16MB AI Model That Can Hear You (YouTube walkthrough)youtube · Video demo of Whistle's offline transcription on device, 7-language support, and use cases across wearables, robots and agents.
- cactus-compute/needlegithub · 13,419 ★ · The Needle engine repo (13k stars) that Whistle runs on — same binary, same .cact container; Whistle loads via the same needle runtime.
Learn more
First spotted on hf: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.