2026-10-11 17:15 UTC

Speakrail's creator claims the released open-source full-duplex voice stack โ€” Voxtral Realtime with an 80ms turn-taking head, a microturn-tuned Gemma 4 12B, and Breeze TTS 2 on one RTX 4090 โ€” rivals GPT-Live (94.0 Full-Duplex-Bench conversational dynamics, ~0.7โ€“0.8s median reply latency) and becomes a widely adopted self-hosted alternative to hosted realtime voice APIs; independent replication of its results and real adoption confirm it.

state: seedheat: lowuncertainty: mediumconvergesscott: highvoice-agents local-inference full-duplex-voiceSpeakraildanil_rootint

What is this?

Speakrail is an open-source, fully-local full-duplex voice assistant from creator danil_rootint, chaining Mistral's Voxtral Realtime streaming ASR with an 80ms turn-taking head, a LoRA-tuned Gemma 4 12B, and Breeze TTS 2 on a single RTX 4090. The Voxtral component side checks out in the supplied coverage: Voxtral Mini 4B Realtime is a real Apache-2.0 streaming model (Feb 2026 arXiv paper) with user-tunable 80msโ€“2.4s delay, sub-second accuracy competitive with Whisper and ElevenLabs Scribe v2 Realtime, day-1 vLLM support, and even a pure-C inference port โ€” and its native audio frame is exactly 80ms (12.5Hz), matching the claimed turn-taking head granularity. The supplied results contain nothing about Speakrail itself, GPT-Live, Full-Duplex-Bench, Breeze TTS 2, or any adoption, so the project's headline claims (94.0 benchmark score, ~0.7โ€“0.8s median reply latency, 'rivals GPT-Live', wide adoption) are attested only by the creator's own release material.

Why it matters to Scott

An independent builder has arrived where the Voice AI's Fork argued the conversation layer would go โ€” a single-RTX-4090 full-duplex stack that commoditises exactly the human-sounding turn-taking that hosted realtime APIs price โ€” and its 'microturn-tuned' Gemma is a direct implementation of latency-renegotiation's time-buying microturns, giving the fork ebook a dated receipt while slotting straight into his Twilio-bridge-vs-hand-built-loop lab lineage, the gamepc 4090 zoo, and the audio TTS comparison bench. Per the grounding the GPT-Live-rivalling numbers are creator-attested only, so replication on his own hardware is the obvious test; this is a distinct sibling to radar:fusion-runtime-local-voice-stack (lineage, not a duplicate), and the Scott-side verdict is converges rather than known because his wikis hold the architecture and the prediction, not this empirical arrival.
ip:source.voice-ais-fork-ebookip:framework.voice-ais-forkip:framework.fast-slow-splitip:concept.latency-renegotiationip:framework.don-t-buy-software-build-aidev:project.twiliodev:project.gamepcdev:project.audioradar:fusion-runtime-local-voice-stackradar:openai-gpt-live-voice-architectureradar:concept.voice-agentsradar:concept.inference-economicsradar:breeze-tts-2-local-releaseradar:tavus-sparrow-2-conversational-flowradar:realtime-venus-open-av-interaction
queries asked of Scott's wikis
  • full-duplex turn-taking voice agent architecture
  • local inference vs hosted realtime voice API economics
  • open-weights self-hosted voice stack strategy
  • end-to-end voice latency budget speech-to-speech
  • voice agent dev project
  • LoRA fine-tuning conversational model behavior

Measured heat

now 0 pts/hpeak 5 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 144h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-05 17:43 (minted)โญ origin echo-reconstructedFull-duplex voice assistant on a single RTX 4090: Voxtral Mini 4B Realtime with an 80ms turn-taking head, Gemma 4 12B with a Speakrail LoRA
danil_rootint (Speakrail) on github (echo) ยท attributed from reddit.post.1wyc1sc ยท published time unknown
โ€”
10-05 15:50first on r/LocalLLaMA ยท published ยท lag ?Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090
danil_rootint
โ€”
10-05 15:50amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyc1sc
danil_rootint
peak 18 ยท 13 comments ยท 100% of case engagement
10-05 17:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wyc1scโ€”
pace: p58 vs 1247 stories at the 96h mark (now 144h old) โ€” ahead of ai-sre-arena-benchmark (1.0x), behind aa-agentperf-local-benchmark (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditSpeakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-10-05T17:32:26.486038+00:00

# Speakrail

A full-duplex voice assistant that runs on a single RTX 4090.

```
git clone https://github.com/speakrail/speakrail && cd speakrail && docker compose up -d
```

Then open **<http://localhost:8080>** and allow the microphone. Don't worry if it takes long to launch, the entire package is ~75 GB and it takes a long time to compile audio.cpp, prerender TTS cache and start vLLM.

agi.mp4

[
](https://private-user-images.githubusercontent.com/34777903/665824562-4d73f149-2d94-4bf6-b0a8-76be4b8a1840.mp4?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTEyMjE4NDUsIm5iZiI6MTc5MTIyMTU0NSwicGF0aCI6Ii8zNDc3NzkwMy82NjU4MjQ1NjItNGQ3M2YxNDktMmQ5NC00YmY2LWIwYTgtNzZiZTRiOGExODQwLm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjEwMDUlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYxMDA1VDE3MzIyNVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWNhODFhOGViNTM1ODNiYmJiMmExYzliM2YwY2NlNjc0M2I5ZmRiZDViZjI0MmEyYTUzMjk0MzEzMGZhMWE4MjImWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT12aWRlbyUyRm1wNCJ9.CV3pwVhoX0uLz6v9rv-42gvaw3AtmX7gOf_om0enpQo)

Speakrail listens while it talks. You can interrupt it, say "mm-hmm" without stopping it, pause mid-sentence without being cut off, and ask it to count your reps while you keep talking.

## How it works

```
mic โ”€โ–บ Voxtral Mini 4B Realtime (audio.cpp) + turn-taking head โ”€โ–บ words + "is the user done?" every 80 ms
                                                                         โ”‚
                       Gemma 4 12B + Speakrail LoRA โ—„โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                       one decision per event: speak / listen / interrupt / yield / interject
                                   โ”‚
                                   โ–ผ
                       Breeze TTS 2 โ”€โ–บ speaker
```

- **Turn-taking head:** a small head on Voxtral's hidden states tells, every 80 ms, whether you're speaking, finished, pausing mid-thought, or just backchanneling. [Model card](https://huggingface.co/speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead).
- **Turn-taking LLM:** Gemma 4 12B with a LoRA that reads your words as they arrive and makes every turn-taking decision itself, in one token per event (~60 ms). [Model card](https://huggingface.co/speakrail/gemma-4-12B-it-qat-Speakrail).
- **Peek and speculation:** when the head thinks you're done, the recognizer's last words are drained early and the reply starts before the turn is confirmed, so the answer is ready the moment you stop.
- **Listening notes:** while you talk, the base model takes notes on what you want, so long or corrected requests are answered correctly.
- **Tools:** weather, timers, notes, lists, calculator, unit conversion, optional web search and an optional larger model for hard questions.

## Results

[FDB-v3 Pass@1 vs. reply quality](https://github.com/speakrail/speakrail/blob/main/docs/charts/fdb3_pass_vs_reply_wide.png)

[FDB-v3 Pass@1 vs. time to task done](https://github.com/speakrail/speakrail/blob/main/docs/charts/fdb3_pass_vs_done_wide.png)

### Conversational dynamics

Pause handling, turn-taking, interruptions and backchannels, scored with the Artificial Analysis formula on Full-Duplex-Bench v1.0 and v1.5: 94.0, the top open-weights score.

[Conversational dynamics](https://github.com/speakrail/speakrail/blob/main/docs/charts/aa_conversational_dynamics.png)

Latency on real recorded conversations: the reply reaches your ear about 0.7-0.8 s after your last word (median).

## Prerequisites

- RTX 4090 (or any other 24 GB NVIDIA card, it should work. If it doesn't - feel free to open an issue. I only have a 4090, so I couldn't test it on anything else.) The prebuilt images cover the RTX 3090 and 4090 generations. An RTX 5090 needs a build from source (see below).
- Linux with a recent NVIDIA driver
- Docker with Compose v2 and the [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html)
- About 75 GB of disk (models ~22 GB, container images ~50 GB)

## Installation

The one-liner above is all you need: every setting has a default. To change settings (port, voice, home city, web search), copy the example file first and edit it:

```
cp .env.example .env
docker compose up -d
```

The first start takes a while, about 20 minutes on a fast connection: about 22 GB of models are downloaded and verified, the services warm up, and short reply openers are pre-rendered in your voice. Later starts take a few minutes. `docker compose logs -f` shows the progress.

The containers come from Docker Hub: [speakrail/voxtral-stt-server](https://hub.docker.com/r/speakrail/voxtral-stt-server), [speakrail/breeze-tts-server](https://hub.docker.com/r/speakrail/breeze-tts-server), [speakrail/app](https://hub.docker.com/r/speakrail/app), [speakrail/model-init](https://hub.docker.com/r/speakrail/model-init), plus vLLM's official image. To build everything from source instead:

```
docker compose up -d --build
```

The speech recognizer is compiled for the RTX 3090 and 4090 generations by default; for an RTX 5090, set `ASR_CUDA_ARCHS=86;89;120` in `.env` before building.

Browsers allow the microphone only on `localhost` or over HTTPS. If Speakrail runs on another machine, use an SSH tunnel:

```
ssh -L 8080:localhost:8080 your-gpu-box
```

`http://localhost:8080/debug/` shows what the system sees and decides: turn-head probabilities, every decision, per-turn latencies.

## Configuration

Everything is in `.env` (see [.env.example](https://github.com/speakrail/speakrail/blob/main/.env.example)):

| Setting | Default |  |
| --- | --- | --- |
| `UI_HOST` / `UI_PORT` | `127.0.0.1` / `8080` | where the web UI listens |
| `LLM_GPU_MEMORY_UTILIZATION` | `0.48` | lower it if the card also drives your display |
| `VOICE` | `female_a.wav` | a reference clip in `app/voices/` with its transcript next to it |
| `HOME_CITY` / `HOME_TIMEZONE` | London | "home" for the weather and time tools |
| `SILENCE_MS` | `1000` | silence that ends your turn when the head is unsure |
| `NOTES` | `1` | listening notes (better answers to long requests) |
| `TOOL_HOLD_MS` | `300` | tools run only after you've been quiet this long |
| `SEARCH` | `off` | `searxng` (local, start with `docker compose --profile search up -d`), `serper` or `brave` (API key) |
| `FIREWORKS_API_KEY` | empty | enables asking a larger model for hard questions (paid) |
| `SESSIONS_DIR` | off | save every session (log, context, audio) for debugging |

## Limitations

- English only.
- One conversation at a time.
- Echo cancellation comes from the client: browsers do it; a bare microphone and speaker on a device without it will hear the assistant as you.
- Following written-format instructions (word counts, markdown) is weaker than base Gemma: the model is tuned for short spoken answers.
- The default voice (Breeze TTS 2) is for research and non-commercial use only; see the license below.

## License

Speakrail's code is licensed under the [Apache License 2.0](https://github.com/speakrail/speakrail/blob/main/LICENSE).

The models it downloads keep their own licenses (details in [NOTICE](https://github.com/speakrail/speakrail/blob/main/NOTICE)):

- **Gemma 4 12B** (Google DeepMind): Apache 2.0
- **Voxtral Mini 4B Realtime** (Mistral AI): Apache 2.0
- **Breeze TTS 2** (BreezeBlue): code Apache 2.0; model weights and the audio they generate are for research and non-commercial use only. If you want to use speakrail commercially, you will have to swap that for another TTS. Any streaming TTS should work (in theory).

Speech recognition runs on [our fork of audio.cpp](https://github.com/speakrail/audio.cpp) (Apache 2.0, by [ShugoAI LLC](https://github.com/0xShug0/audio.cpp)), which adds the turn-taking head and peek decoding to its Voxtral realtime model. The voice runs on [our fork of Breeze TTS](https://github.com/speakrail/breeze-tts), which adds int8 serving next to an LLM on one card.

## Troubleshooting

**The UI doesn't open, or `docker compose ps` shows the app as `Created`:** the UI port is probably taken by another program. Pick a free port in `.env` (for example `UI_PORT=8085`), then recreate the app container:

```
docker compose up -d --force-recreate app
```

A container that failed to start on a busy port keeps a broken network setup, so a plain restart is not enough: it has to be recreated.
danil_rootint1813
๐ŸŸง echo.github โญFull-duplex voice assistant on a single RTX 4090: Voxtral Mini 4B Realtime with an 80ms turn-taking head, Gemma 4 12B with a Speakrail LoRA danil_rootint (Speakrail)โ€”โ€”

Interpretation history

Decision trace