2026-10-11 17:09 UTC

Fusion-runtime maintainer SamarthUrs18 claims the released single-process speech-to-text, LLM, and text-to-speech stack delivers roughly 991-millisecond end-of-speech response latency with interruption handling on an RTX 3090, potentially simplifying responsive self-hosted voice-agent deployment.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumvoice-agents agent-runtime local-inferenceSamarthUrs18

What is this?

The supplied case describes Fusion-runtime as a self-hosted voice-agent runtime released by maintainer SamarthUrs18, combining speech recognition, an LLM, and speech synthesis in one process with browser and CLI clients. The maintainer reportedly claims roughly 991 milliseconds from end of speech to response on an RTX 3090, with interruption handling. None of the supplied web results directly covers Fusion-runtime or its maintainer, so the release details and performance remain unverified here. The surrounding snippets identify streaming, end-of-turn detection, service handoffs, and interruption handling as important voice-runtime concerns, but their differing latency figures do not establish a comparable benchmark for this project.

Why it matters to Scott

Fusion-runtime independently packages the STTβ†’LLMβ†’TTS loop Scott explored in his Twilio realtime voice AI laboratory, making it a concrete self-hosted candidate to evaluate against Practice Trainer’s Ultravox-based voice sessions and on his gamepc speech-serving substrate. The release and roughly 991-ms latency remain unverified maintainer claims, so this warrants an integration and interruption-handling benchmark rather than a change to his architectural claims; the supplied radar hits track adjacent local speech runtimes, not this development.
dev:project.twiliodev:project.sales-trainerdev:project.gamepcradar:concept.voice-agentsradar:concept.local-inferenceradar:crispasr-local-audio-runtimeradar:nemo-speech-cpp-local-stack
queries asked of Scott's wikis
  • voice-agent projects speech recognition speech synthesis integration
  • single-process agent runtimes versus distributed orchestration
  • local GPU inference deployment cost operational simplicity
  • conversational latency end-of-turn detection streaming benchmarks
  • agent interruption handling cancellation realtime interaction

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 480h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-21 16:27 (minted)⭐ origin echo-reconstructedPublishes an installable self-hosted streaming voice runtime with browser and CLI clients, reporting about 490 ms processing plus a configur
SamarthUrs18 on github (echo) Β· attributed from hn.story.49789320 Β· published time unknown
β€”
09-21 16:19first on hacker news Β· published Β· lag ?Show HN: Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process
samarthurs18
β€”
09-21 16:19amplified on hacker news πŸ‘‘hn.story.49789320
samarthurs18
peak 2 Β· 1 comments Β· 101% of case engagement
09-21 16:20our radar first saw it Β· lag ?discovery anchor: hn.story.49789320β€”
pace: p32 vs 1032 stories at the 336h mark (now 480h old) β€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process
Retrieved article excerpt

Open article Β· Retrieved 2026-09-21T16:23:40.126475+00:00

fusion-runtime

[GitHub stars](https://github.com/SamarthUrs18/fusion-runtime)
[PyPI](https://pypi.org/project/fusion-runtime/)
[Docs](https://fusion-runtime.dev/docs)
[Python 3.11 to 3.13](https://camo.githubusercontent.com/23b639181075ba5ffe31a53c6e4b5668ebf2dc7ff043ffb7491be7182ba8d68a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f707974686f6e2d332e3131253230253743253230332e3132253230253743253230332e31332d626c7565)

**A self-hosted voice agent runtime.** Speech-to-text, the LLM and text-to-speech run together on
one machine and stream into each other, so a reply starts playing while it's still being generated.

**On an RTX 3090 with a 7B model: about 490 ms of processing once a turn ends**, or 991 ms
stopwatched from your last syllable β€” the difference is a silence wait you can configure. 127
tokens/sec, interruptions honoured mid-sentence.

## Quickstart

Requires Python 3.11–3.13.

```
pip install fusion-runtime
```

An agent is one file. This is the whole thing:

```
# agent.py
from fusion_runtime import Agent, LLM, STT, TTS, Turns

agent = Agent(
    name="shopkart-orders",
    prompt="You are the order line for ShopKart. Keep answers to one short sentence.",
    stt=STT("whisper-tiny.en"),          # or "whisper-small" for better accuracy
    llm=LLM("qwen2.5-0.5b-q4", max_tokens=256),
    tts=TTS("kokoro-v1.0", voice="af_heart"),
    turns=Turns(wait_ms=500, interrupt_after_ms=300),
)
```

```
frun models pull agent.py     # exactly the models it names, nothing else
frun up agent.py              # add --reload to restart on every edit
```

Then talk to it from a second terminal:

```
pip install "fusion-runtime[talk]"
frun talk
```

That is the whole loop β€” one file, two commands, a conversation. Talk over the agent to
interrupt it.

`frun up` with no file runs a default agent if you just want to hear it work, and `frun doctor`
checks libraries, GPU, models and audio and says how to fix what it finds.

### In a browser instead

The runtime serves a browser client at **<http://localhost:8000>** β€” the same one you would embed
in your own page.

With no keys configured, open it and click Talk. With keys configured (`FUSION_ACCEPTED_KEYS`),
a page can't hold a secret, so it needs a short-lived session token:

```
frun token        # prints a URL with a token in it β€” open that
```

Tokens are single-use and expire in about a minute. The page is handed its next one over the
socket it already has, so a conversation keeps going without asking again. If you open the bare
URL on a server with keys, the connection closes and the page says the token wasn't accepted.

### Naming models

A model is a catalog id (`frun models list`), a file path, `hf:owner/repo` for anything on
Hugging Face, or a URL for an OpenAI-compatible endpoint. Settings the config knows are applied;
anything else is passed through to that runtime.

Secrets never go in the agent file β€” it names the *variable* holding a key
(`api_key_env="GROQ_API_KEY"`), so `agent.py` is safe to commit.

## On your own site

```
<script src="https://your-server/fusion-runtime.js"></script>
<button id="talk"></button>
<script>FusionRuntime.attach({ button: "#talk" });</script>
```

The runtime serves the browser client it uses itself, so the page you demo with is the one your
site embeds. With no `url` it connects back to wherever the script came from.

Browsers only allow a microphone on `https://`, so a deployment needs TLS and `wss://`. A page
never holds an API key: your backend mints it a short-lived token.

## Authentication

```
frun key new
FUSION_ACCEPTED_KEYS=web:frun_kR7m...
```

Without keys the server answers on `localhost` only, and `frun up --host 0.0.0.0` refuses to
start. `frun talk`, a backend or curl send the key in an `Authorization` header; a browser page
gets a short-lived, single-use token from `POST /v1/sessions` instead, because a page can hold
neither a secret nor a header.

Concurrency caps, message and audio limits, idle timeouts, origin allowlists and proxy trust all
have working defaults β€” see the docs.

## Performance

Measured, not estimated. The production profile as it ships β€” RTX 3090, Qwen 7B q4 + Whisper
small + Kokoro on the one card β€” through the browser client, 21 September 2026:

|  | Median | Range |
| --- | --- | --- |
| **Processing** β€” turn ends, audio comes back | **~490 ms** | 288–657 |
| **Stopwatch from your last syllable** | **991 ms** | 858–1061 |
| ↳ of which: silence wait before the turn is judged over | ~500 ms | `turns.wait_ms` |
| Speech-to-text | 119 ms | 58–329 |
| LLM first token | 27 ms | 20–70 |
| First token β†’ first audio (a sentence gets written, then spoken) | 430 ms | 320–509 |
| Text-to-speech real-time factor | 0.09 | speech is synthesized ~11Γ— faster than real time |
| LLM tokens/sec | 127 | 106–130 |

**Two numbers, because there are two honest answers.** A stopwatch started at your last syllable
reads 991 ms. About 500 ms of that is the runtime waiting through silence to decide you've
finished β€” which elapses while you're still finishing, so people don't experience it as waiting.
What a caller feels is closer to the 490 ms of processing. Quote whichever you like, but say
which one: a voice stack claiming a number under 500 ms is almost always measuring from "we
decided the caller stopped", not "the caller stopped".

**The stages don't sum, and that's not sleight of hand.** Transcription of what you already said
runs during the silence wait. And "first token β†’ first audio" is mostly the language model
writing a sentence β€” text-to-speech can't start on half a clause β€” so it is not a measure of how
fast Kokoro is. Kokoro's own speed is the real-time factor: 0.09, or about 126 ms of compute for
1.4 seconds of speech.

Every figure is the runtime's own per-turn telemetry (`frun talk --verbose`, or the browser
console), so you can reproduce them rather than trusting ours. Barge-in fired on every attempt.

## Several callers at once

Measured on the same 3090, real WebSocket sessions, three turns each:

| Callers | Response, median | Turns/sec |
| --- | --- | --- |
| 1 | ~460 ms | 0.21 |
| 4 | ~740 ms | 0.55 |
| 8 | ~4600 ms | 0.69 |
| 12 | ~7500 ms | 0.74 |

**Four simultaneous callers land in the same range as one**, within run-to-run variance. Past
that it saturates: throughput plateaus around 0.7 turns/sec, so an extra caller past the knee
buys queue time rather than capacity. Eight is not a conversation.

The bottleneck is one specific thing. At twelve callers the language model's first token takes
4790 ms of a 5312 ms response, while speech-to-text stays at 76 ms and text-to-speech at 469 ms.
A single in-process llama.cpp context decodes one reply at a time; the speech stages do not care
how many callers there are.

So to go past four, move the language model out and leave speech where it is:

```
llm = LLM("http://localhost:8080/v1", model_name="qwen2.5-7b-instruct")
```

vLLM and `llama-server -np N` both speak the API the `openai_http` runtime uses. Whether that
moves the knee, and how far, is not yet measured.

## The `frun` CLI

|  |  |
| --- | --- |
| `frun up [agent.py]` | Starts the server. `--host`, `--port`, `--reload`, `--config` |
| `frun talk` | Talks to it from a terminal, with a latency summary per turn |
| `frun models list` / `pull` | What's available, and downloading it |
| `frun key new` / `keys list` / `token` | Keys and browser tokens |
| `frun doctor` | Checks the machine and says how to fix what's wrong |
| `frun version` | The installed version. `--version` and `-V` work too |

`fusion-runtime` works as an alias for `frun`.

## Using it as a library

`frun up agent.py` covers running an agent. The pipeline can also run inside your own process β€”
for a queue worker, a test, or a batch job over recorded calls β€” with no server involved. See
[`examples/sdk_example.py`](https://github.com/SamarthUrs18/fusion-runtime/blob/main/examples/sdk_example.py), which is runnable, and
[the docs](https://fusion-runtime.dev/docs#python).

## Documentation and contact

Everything else β€” configuration, turn detection, languages, the server API, telemetry, limits,
GPU setup and deployment β€” is at **[fusion-runtime.dev/docs](https://fusion-runtime.dev/docs)**.

|  |  |
| --- | --- |
| Site and docs | **[fusion-runtime.dev](https://fusion-runtime.dev)** |
| Questions, or anything else | **[email protected]** |
| Security problems | **[email protected]** β€” not a public issue, please ([why](https://github.com/SamarthUrs18/fusion-runtime/blob/main/CONTRIBUTING.md#security)) |

## Development

```
uv sync --extra dev --extra talk     # or: pip install -e ".[dev,talk]"
pytest
```

CI runs the suite on Python 3.11, 3.12 and 3.13. [CONTRIBUTING.md](https://github.com/SamarthUrs18/fusion-runtime/blob/main/CONTRIBUTING.md) has the
layout, the design rules a review will hold you to, and how to add a runtime.

## License

[Apache-2.0](https://github.com/SamarthUrs18/fusion-runtime/blob/main/LICENSE). Embed it in a commercial product, rebrand it, ship it closed β€” keep the
copyright notice and the `NOTICE` file in what you distribute, and don't use the project's name
to imply it endorses you.

The models it downloads by default are permissive too (Whisper MIT, Silero VAD MIT, Qwen2.5
Apache-2.0, Kokoro Apache-2.0), so the whole default path is clear for commercial use. A model
you point it at yourself carries its own licence β€” check that one before you ship it.
samarthurs1821
🟧 echo.github ⭐Publishes an installable self-hosted streaming voice runtime with browser and CLI clients, reporting about 490 ms processing plus a configurSamarthUrs18β€”β€”

Interpretation history

Decision trace