2026-10-11 16:37 UTC

sanoTTS’s creator claims the released 294K-parameter multilingual speech stack fits in 337 KB and runs without an NPU on a $3 microcontroller, potentially making usable neural TTS practical on severely constrained edge hardware.

state: watchingheat: lowuncertainty: highconvergesscott: mediumlocal-tts edge-inference model-compression local-inferencesanoTTSAffectionate_Hat_585

What is this?

sanoTTS is an open-source GPLv3 family of tiny neural text-to-speech models released by the Ampixa/ampixa account, with project listings claiming fully local operation without an NPU and real-time inference on an approximately $3 ESP-class microcontroller. The release claims a 294K-parameter, 337 KB complete stack plus larger variants, multilingual support, and benchmark results competitive with models several times larger. The supplied snippets support the broad constrained-edge claim, but model-size descriptions vary—from 294K in the release claim to 745K–1.8M in the Hugging Face excerpt—and the performance and “smallest” assertions appear to be project-authored rather than independently verified.

Why it matters to Scott

The release converges with Scott’s hardware-aware local-inference work and directly extends his local speech-engine laboratory toward microcontroller-class deployment rather than workstation or mobile inference. If independently validated, the claimed 337 KB stack could materially change where he can deploy voice interfaces, but the inconsistent size descriptions and project-authored benchmarks make this an evaluation candidate rather than a settled breakthrough.
dev:concept.hardware-aware-local-inferencedev:project.audiodev:technology.kokoro-ttsradar:concept.edge-inferenceradar:concept.tiny-modelsradar:concept.local-audio-inferenceradar:264kb-microcontroller-diffusion
queries asked of Scott's wikis
  • tiny-model compression and capability thresholds
  • microcontroller inference without NPUs
  • local voice stack and offline assistants
  • edge inference economics versus cloud APIs
  • privacy and sovereignty in on-device speech
  • resource-constrained multimodal agent interfaces

Measured heat

now 0 pts/hpeak 30 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 906h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-03 22:01⭐ origin directly observedI released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size
Affectionate_Hat_585 on r/LocalLLaMA
—
09-20 08:30first on r/MachineLearning · published · +394.5hInside sanoTTS — a 294,279-parameter TTS system [P]
donttmesswithme
—
09-30 11:34first on r/LocalLLaMA · published · +637.6hOído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)
Significant-Price695
—
10-02 15:04first on hacker news · published · +689.0hShow HN: Ito – streaming neural TTS for the ESP32-S3 (4.9 MB, no NPU, no cloud)
dani-lokutor
—
09-03 22:01amplified on r/LocalLLaMA 👑reddit.post.1w6lmmg
Affectionate_Hat_585
peak 515 · 121 comments · 66% of case engagement
09-20 08:30amplified on r/MachineLearningreddit.post.1wlbhw8
donttmesswithme
peak 15 · 4 comments · 2% of case engagement
09-30 11:34amplified on r/LocalLLaMAreddit.post.1wu2jjy
Significant-Price695
peak 228 · 40 comments · 28% of case engagement
10-02 15:04amplified on hacker newshn.story.49934308
dani-lokutor
peak 2 · 2 comments · 1% of case engagement
10-05 13:38amplified on r/LocalLLaMAreddit.post.1wy8ske
Significant-Price695
peak 25 · 3 comments · 3% of case engagement
09-03 22:20our radar first saw it · +0.3hdiscovery anchor: reddit.post.1w6lmmg—
pace: p90 vs 519 stories at the 720h mark (now 906h old) — ahead of llama-cpp-hot-swappable-ple-memory (1.0x), behind debian-llm-usage-vote (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size
LocalLLaMA
Affectionate_Hat_585515121
🟠 redditInside sanoTTS — a 294,279-parameter TTS system [P]
MachineLearning
donttmesswithme144
🟠 redditOído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)
LocalLLaMA
Significant-Price69522740
🟧 hnShow HN: Ito – streaming neural TTS for the ESP32-S3 (4.9 MB, no NPU, no cloud)dani-lokutor22
🟠 redditItoTTS: two natural English voices in 4.89 MB for a $5 ESP32-S3
LocalLLaMA
Significant-Price695253

Interpretation history

Decision trace