2026-10-11 16:38 UTC

Cactus Compute claims its released Whistle ASR model โ€” 55M parameters in a 16.9MB file โ€” mostly beats Whisper base across seven languages at ~9x less size and 6x the speed, and independent adoption on budget phones, wearables, and microcontrollers versus a quiet fade decides whether ultra-small ASR becomes a practical edge-inference default.

state: corroboratedheat: mediumuncertainty: mediumnovelscott: highon-device-speech-recognition edge-inference model-compressionCactus Compute

What is this?

Cactus Compute โ€” the startup behind the Needle tool-calling models for phones, wearables, and microcontrollers โ€” released Whistle on Oct 2, 2026: a speech-to-text model shipped as a single 16.9MB file (55M parameters per the case; the snippets confirm the file size and architecture but the parameter count appears only in the case hypothesis). It runs on CPU in the same .cact container, Cactus Quants, and C++ engine as Needle, supports seven European languages (English, German, French, Spanish, Italian, Dutch, Polish), adds keyword biasing and decoder-native word timestamps, and deploys across 17 platforms including WebAssembly and WASI. The 'mostly beats Whisper base' claim is vendor-reported and internally consistent: the GitHub README lists Whistle ahead on LibriSpeech (4.31 vs 4.9 WER test-clean), SPGISpeech, Earnings-22, and the FLEURS average, with Whisper base ahead on TED-LIUM, AMI, and MLS โ€” but the snippets contain no independent verification and no evidence yet either way on adoption versus a quiet fade; that part of the hypothesis is forward-looking.

Why it matters to Scott

This is the third release in Cactus Compute's tracked lineage (Needle 2 โ†’ Needle 3 โ†’ Whistle) and it makes a direct claim against the ASR Scott actually runs: if the vendor benchmarks hold, Whistle beats the Whisper base already in his transcription fallback chains (bulk-transcribe) at ~1/9th the size on plain CPU, and gives the listen/ambient-conversation and mobile voice projects an on-device option their current Parakeet-on-Mac-mini and cloud-STT (Groq/Ultravox) setups lack โ€” dev:project.audio is a ready testbed. No page in his canon holds or challenges the ultra-small-ASR-beats-Whisper claim, so this is new information to test, not a receipt for or threat to an existing position.
dev:technology.whisperdev:technology.parakeet-mlxdev:project.bulk-transcribe-youtube-videos-from-playlistdev:project.audiodev:project.listendev:concept.hardware-aware-local-inferenceradar:cactus-needle3-on-device-automationradar:needle-2-edge-agent-modelradar:orukeet-multilingual-local-asrradar:concept.speech-to-textradar:concept.speech-recognitionradar:concept.edge-inferenceradar:concept.local-inference
queries asked of Scott's wikis
  • tiny model edge inference local-first strategy
  • local speech transcription pipeline tooling RAG
  • quantization 2-bit compression model quality tradeoff
  • voice input agent harness hands-free interface
  • on-device privacy vs cloud transcription tradeoff
  • Moonshine Whisper local ASR benchmarks comparison

Measured heat

now 0 pts/hpeak 88 pts/hcomments 0/hpeers p16momentum: steady3 platformsage 242h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-01 14:00โญ origin echo-reconstructedBlog: "Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It
Cactus Compute (post bylines: Jakub Mroz and Henry Ndubuaku) on blog (echo) ยท attributed from reddit.post.1wyemcb
โ€”
10-05 17:27first on r/LocalLLaMA ยท published ยท +99.5hWhistle: speech to text in a 16.9MB file
Henrie_the_dreamer
โ€”
10-08 16:59first on hacker news ยท published ยท +171.0hWhistle: Speech to Text in 16.9 MB
gmays
โ€”
10-05 17:27amplified on r/LocalLLaMAreddit.post.1wyemcb
Henrie_the_dreamer
peak 186 ยท 63 comments ยท 11% of case engagement
10-08 16:59amplified on hacker news ๐Ÿ‘‘hn.story.50008427
gmays
peak 898 ยท 174 comments ยท 89% of case engagement
10-05 20:20our radar first saw it ยท +102.3hdiscovery anchor: reddit.post.1wyemcbโ€”
pace: p76 vs 1188 stories at the 168h mark (now 242h old) โ€” ahead of aisle-six-curl-cves (1.0x), behind gitspawn-repository-agent-hijack (1.0x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditWhistle: speech to text in a 16.9MB file
LocalLLaMA
Henrie_the_dreamer18663
๐ŸŸง echo.blog โญBlog: "Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It Cactus Compute (post bylines: Jakub Mroz and Henry Ndubuaku)โ€”โ€”
๐ŸŸง hnWhistle: Speech to Text in 16.9 MBgmays898174

Interpretation history

Decision trace