2026-10-11 16:38 UTC

Lokutor claims its released Oído engine runs open-vocabulary English speech recognition entirely on a $5 ESP32-S3 (3.7% test-clean LibriSpeech WER, no cloud and no NPU — which it calls the most accurate published microcontroller result) despite real-time speed still being only emulator-estimated; on-silicon confirmation would make open-vocabulary ASR on microcontrollers a practical edge-inference capability for voice agents rather than a benchmark claim.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highon-device-speech-recognition microcontroller-inference local-models edge-inferenceLokutordani-lokutor

What is this?

Lokutor (dani-lokutor) released Oído, an open-vocabulary English speech recognition engine claiming 3.7% WER on LibriSpeech test-clean and 8.2% on test-other using an int8 Conformer-CTC model running entirely on a $5 ESP32-S3 microcontroller with no NPU and no cloud. The project states this is the most accurate published microcontroller ASR result. However, real-time latency is currently only emulator-estimated via QEMU instruction counts; on-silicon measurement on physical hardware remains the unresolved gate. Independently, the Babytalk project demonstrates offline STT+TTS on ESP32-S3/P4 using 4-bit quantized kernels and larger models, corroborating that the ESP32-S3/P4 hardware family can run open-vocabulary speech both ways. The core thesis — open-vocabulary ASR on microcontrollers as a practical edge-inference capability for voice agents — now has two independent implementations on the same hardware family, though Lokutor's specific accuracy and latency claims still await physical-board validation.

Why it matters to Scott

Lokutor's Oído engine independently validates Scott's hardware-aware local inference thesis — int8 Conformer-CTC on ESP32-S3 with no NPU, explicit SRAM pressure management, and accelerator-aware compilation — while the Babytalk replication corroborates the same hardware family running duplex STT+TTS. This directly bears on Scott's audio lab stack (Whisper, Kokoro, Parler-TTS, F5-TTS) and the sanoTTS MCU-TTS radar case as the ASR half of a duplex neural speech loop on microcontrollers. The unresolved on-silicon latency gate is a concrete near-term validation point for Scott's edge-inference economics.
dev:concept.hardware-aware-local-inferencedev:project.audiodev:technology.whisperdev:technology.kokoro-ttsdev:technology.parler-ttsdev:technology.f5-ttsdev:technology.parakeet-mlxip:framework.voice-ais-forkwork:concept.dynaquest-ai-dental-receptionistwork:project.leverageairadar:sanotts-microcontroller-ttsradar:cactus-whistle-edge-asrradar:audio-cpp-0-4-local-speech-validationradar:fusion-runtime-local-voice-stackradar:speakrail-local-full-duplex-voiceradar:parakeet-webgpu-browser-asr
queries asked of Scott's wikis
  • hardware-aware local inference thesis int8 quantization SRAM pressure
  • audio lab stack sanoTTS MCU-TTS duplex neural speech loop
  • open-weights model sovereignty local inference economics
  • edge inference voice agents microcontroller deployment patterns
  • quantization int8 vs int4 kernels ESP32-S3 P4 performance
  • local models agent memory voice interface architecture

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 290h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-29 14:00⭐ origin echo-reconstructedRepo claims open-vocabulary English ASR fitting in a $5 ESP32-S3: 3.7% test-clean / 8.2% test-other LibriSpeech WER with int8 Conformer-CTC,
Lokutor (posted to HN by dani-lokutor) on github (echo) · attributed from hn.story.49907387
—
09-30 11:30first on hacker news · published · +21.5hOído: Open-vocabulary speech recognition on a $5 ESP32-S3 (3.7% LibriSpeech WER)
dani-lokutor
—
09-30 11:30amplified on hacker newshn.story.49907387
dani-lokutor
peak 3 · 0 comments · 13% of case engagement
10-09 21:33amplified on hacker news 👑hn.story.50026819
tlack
peak 13 · 7 comments · 87% of case engagement
09-30 12:20our radar first saw it · +22.4hdiscovery anchor: hn.story.49907387—
pace: p22 vs 1188 stories at the 168h mark (now 290h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOído: Open-vocabulary speech recognition on a $5 ESP32-S3 (3.7% LibriSpeech WER)
Retrieved article excerpt

Open article · Retrieved 2026-09-30T12:26:38.462229+00:00

# Oído: speech recognition that fits in a $5 chip

*¡Oído!* is what cooks call out in a Spanish kitchen to confirm an order: *heard, got it*.

Speech-to-text for any English sentence, running entirely on an **ESP32-S3** (240 MHz dual-core Xtensa LX7, 8 MB PSRAM,
16 MB flash). No cloud, no command list, no neural accelerator. Built by [Lokutor](https://lokutor.com).
Model on Hugging Face: [lokutor-ai/oido-ctc-small-int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8).

> **Status (30 September 2026).** Every transcript below comes from the exact arithmetic of the on-chip engine: the host
> build is bit-identical to the firmware, and firmware transcripts under Espressif's QEMU emulator match it word for word.
> Real-time speed is **estimated** from exact emulator instruction counts. Measurements on physical boards follow in the
> next days and will be added here.

## Accuracy

Word error rate (%) on LibriSpeech, same text normalization for every system.

| System | Runs on | test-clean | test-other | Size |
| --- | --- | --- | --- | --- |
| **Oído**: NVIDIA Conformer-CTC Small, int8, greedy (this repo) | ESP32-S3 | **3.7** | **8.2** | 14.0 MB |
| Oído with NVIDIA Conformer-Transducer Small, int8 (weights not included, see below) | ESP32-S3 | 3.0 | 6.7 | 15.5 MB |
| Espressif MultiNet7 (ESP-SR benchmark; its API takes fixed command lists) | ESP32-S3 | 8.5 | 21.3 | 2.9 MB |
| Moonshine tiny, fp32 | laptop | 5.0 | 12.1 | 27 M params |
| Whisper tiny.en, fp32 | laptop | 6.3 | 15.9 | 39 M params |
| Vosk small (Kaldi) | laptop | 9.9 | 21.6 | 40 MB |

- On-chip rows use the full test sets. Laptop baselines use 500 evenly spaced utterances per set. MultiNet7 figures are from
  Espressif's ESP-SR benchmark page.
- The int8 engine is within 0.1 points of full precision: 3.70 / 8.23 on chip vs 3.68 / 8.11 for the original fp32 model.
- To our knowledge this is the most accurate LibriSpeech result published for any microcontroller. It is not the first
  open-vocabulary recognizer on one (Arm has shown Conformer models on Cortex-M55 + Ethos-U NPUs).

**Robustness** (300 LibriSpeech utterances under real DEMAND noise, babble and room reverb; `eval/make_robust.py`,
full numbers in [`results/robustness.json`](https://github.com/lokutor-ai/oido/blob/main/results/robustness.json)):

| Mean WER over 14 conditions | This repo, CTC int8 (chip) | Transducer int8 (chip) | Whisper tiny.en | Moonshine tiny | Vosk small |
| --- | --- | --- | --- | --- | --- |
|  | **8.4** | 6.7 | 12.1 | 12.2 | 21.7 |

For the transducer, car and kitchen noise at 5 dB SNR cost under 1 point, and living-room noise about 1.7. Four-talker babble at 5 dB and very
reverberant rooms are the hard cases.

## Speed and memory

|  |  |
| --- | --- |
| Flash | 14.0 MB model (int8); partition layouts for 16 MB modules in `esp32/firmware/partitions_*.csv` |
| PSRAM | 2.4 MB working memory peak for a 20 s utterance (measured in QEMU); the rest caches the most-reused weights |
| Compute | ~225 M instructions per second of audio across both cores, ~121 M on the dual-core critical path (exact, QEMU `-icount`) |
| Real-time factor | **estimated 0.7–0.95**: 1.3–1.6 cycles per instruction at 240 MHz, plus flash/PSRAM stalls. Not yet measured on silicon |
| Latency | Utterance mode. Text appears after a 0.8 s pause plus compute: about 3 s for a 2–4 s command |

## How it works

- **Front end:** log-mel features, then 2× (3×3, stride 2) convolution subsampling to 25 Hz.
- **Encoder:** 16 Conformer layers (d = 176, 4 heads, relative-position attention, conv kernel 31).
- **Decoding:** CTC over 1024 BPE tokens. The engine also supports an RNN-T head (LSTM 320 + joint network) and a GRU
  language model with CTC prefix beam search.

The engine (`esp32/components/tinyasr`) is new C written for the ESP32-S3's PIE vector unit:

- int8 matrix kernels on `EE.VMULAS.S8.ACCX` (16 MACs per instruction), plus int4 outer-product kernels;
- int8 relative-position attention with a lookup-table softmax;
- dual-core scheduling;
- tiling so that each weight is streamed from flash once per 64 frames (weight traffic cut from 18 to 7.5 MB/s);
- a VAD/AGC utterance segmenter;
- an optional SSD1306 OLED that shows the live transcript.

## Try it

**On a laptop, with the chip's exact arithmetic.** Needs Python with numpy, soundfile, sentencepiece and sounddevice.

```
cd esp32/host && make
./tasr_cli ../../models/nemo8.tnm recording.wav      # 16 kHz mono PCM16 wav
python live_demo.py                                   # microphone -> the firmware's VAD + engine, with ESP32 time estimates
```

**On a board.** ESP32-S3-DevKitC-1 **N16R8**, an INMP441 I2S microphone (SCK→GPIO4, WS→GPIO5, SD→GPIO6, L/R→GND),
and optionally a 0.96" SSD1306 OLED (SDA→GPIO8, SCL→GPIO9). Needs ESP-IDF v5.5.

```
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm              # live microphone
TASR_OLED=1 esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm  # + transcript on the OLED
TASR_MODE=file esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm clip.wav "reference"   # prints measured RTF
```

**In the emulator** (Espressif QEMU 9.x): runs the real firmware, then reports the transcript and instruction counts.

```
esp32/tools/emulate.sh clip.wav
```

**The transducer model.** Its weights are NVIDIA's, distributed on NGC under NVIDIA's terms, so they are not included
here. You can fetch and convert them yourself:

```
cd train && python fetch_nemo_small.py ../models/nemo_rnnt --transducer
python export_nemo.py ../models/nemo_rnnt ../models/rnnt8.tnm 8
```

## Repository

```
esp32/components/tinyasr/  on-chip engine: tasr_nemo.c (Conformer CTC/RNN-T), kernels.c (PIE SIMD), tinyasr_lm.c
                           (GRU LM + beam search), tasr_seg.c (VAD), tinyasr.c (streaming engine)
esp32/firmware/            ESP-IDF app: live I2S microphone or benchmark mode, OLED, partition layouts
esp32/host/                host build of the engine: tasr_cli, live_demo.py, seg_test, eval_engine.py, benchmark.py
esp32/tools/               flash.sh, emulate.sh, run_qemu.sh, bench_latency.py, mkimages.py
train/                     PyTorch port of NVIDIA's model (nemo_small.py, rnnt_small.py), exporters, GRU LM training
eval/                      WER normalization, robustness benchmark builder, laptop baselines
results/                   benchmark outputs behind the numbers above
models/                    nemo8.tnm (int8 Conformer-CTC Small) and its tokenizer
```

## Limitations

- English only.
- Text appears after each utterance, not word by word.
- Very noisy crowds and reverberant rooms remain hard.
- Speed is estimated until board measurements are published.
- Requires an ESP32-S3 with 16 MB flash and 8 MB octal PSRAM (N16R8).

## License

- **Code** is licensed under the **GNU GPL v3** ([`LICENSE`](https://github.com/lokutor-ai/oido/blob/main/LICENSE)).
- For products that cannot meet GPLv3 terms (for example, consumer devices that do not allow users to install modified
  firmware), Lokutor offers commercial licenses and support. See [`COMMERCIAL.md`](https://github.com/lokutor-ai/oido/blob/main/COMMERCIAL.md).
- **Model weights** in `models/` are derived from NVIDIA's `stt_en_conformer_ctc_small` and remain under
  **CC-BY-4.0**. See [`NOTICE`](https://github.com/lokutor-ai/oido/blob/main/NOTICE).

Lokutor also has Spanish and other-language models, an int4 profile with more compute headroom, and an on-device TTS for
the same chip. Contact us for these.
dani-lokutor30
🟧 echo.github ⭐Repo claims open-vocabulary English ASR fitting in a $5 ESP32-S3: 3.7% test-clean / 8.2% test-other LibriSpeech WER with int8 Conformer-CTC,Lokutor (posted to HN by dani-lokutor)——
🟧 hnShow HN: Babytalk: Offline speech to text and text to speech on ESP32tlack137

Interpretation history

Decision trace