Whistle: Speech to Text in 16.9 MB
Cactus Compute unveils Whistle, a remarkably compact (16.9 MB) and fast speech-to-text model designed for on-device applications. Running solely on CPU with no dependencies, it delivers impressive accuracy and speed across seven languages, outperforming larger models in key metrics. Hacker News is abuzz with excitement over its potential for local AI solutions and the technical feat of its efficiency.
The Lowdown
Whistle is a new open speech recognition model developed by Jakub Mroz and Henry Ndubuaku of Cactus Compute, making significant strides in on-device AI. Its core appeal lies in its incredibly small footprint and high performance, enabling powerful speech-to-text capabilities without reliance on cloud services or dedicated GPUs.
- Miniature & Mighty: The entire model is contained within a single 16.9 MB file, running efficiently on a CPU with no external dependencies.
- Multilingual Support: It transcribes seven languages: English, German, French, Spanish, Italian, Dutch, and Polish, with automatic language detection.
- Core Functionality: Whistle provides transcription, precise word timestamps derived from the decoder's attention, and speech embeddings from the encoder output.
- Blazing Fast: It achieves an 11 ms time to first token and processes 1,319 tokens per second on a 10-second audio clip, significantly faster than competitors like Whisper base.
- Robust Architecture: Built on shared components with Cactus Compute's Needle text model, it utilizes Simple Attention blocks for the encoder and Laddered Simple Attention blocks with gated cross-attention for the decoder, employing a 5-beam search with keyword biasing.
- Seamless Integration: Designed to load alongside Needle, allowing a single binary to handle both speech input and tool calls, directly converting audio to structured JSON outputs.
- Broad Compatibility: Prebuilt binaries are available for seventeen targets, including various operating systems (macOS, Linux, Android, iOS, Windows) and architectures (ARM, RISC-V, MIPS, WASI).
This release positions Whistle as a leading solution for edge computing and local AI applications where size, speed, and privacy are paramount, offering a compelling alternative to larger, cloud-dependent speech recognition systems.
The Gossip
Accuracy Assessments & Accent Adaptability
Users shared varied experiences regarding Whistle's transcription accuracy. Many praised its performance for English, noting its ability to handle diverse accents well (e.g., Indian accent). However, others reported mixed results with non-English languages like Spanish, citing typos or mangled words, and some issues with specific English contexts or less common languages. The discussion also touched upon the challenging real-world scenario of transcribing impaired speech, suggesting that general STT models might not suffice for specialized dictation needs.
Comparative Capabilities & On-device Potentials
The conversation frequently drew comparisons between Whistle and other local or on-device speech-to-text solutions, such as Parakeet, Whisper, and FUTO. Commenters were particularly impressed by Whistle's minuscule size (16.9 MB) and superior speed benchmarks, seeing it as a significant step forward for efficient, local AI. Many envisioned its widespread use in mobile devices, wearables, and embedded systems, with some even speculating about its potential for being 'always on' from an L3 cache on a CPU.
Feature Gaps & Future Desires
While generally positive, the community identified several areas for improvement or desired features. Key requests included real-time streaming transcription (where text appears as one speaks), the removal or extension of the 30-second audio clip limit, and support for a wider array of languages beyond the current seven. There was also an underlying theme about the need for more specialized dictation models to handle nuanced or impaired speech patterns more effectively than general-purpose STT solutions.