NineNineSix claims its Apache-2.0 Gepard 1.0 model reaches 68.7 milliseconds median time-to-first-audio and 5.23% WER on one RTX 4090, which would make it a leading low-latency open TTS option on commodity GPU hardware.
state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference open-models text-to-speechNineNineSixCoval
What is this?
Gepard 1.0 is a 555M-parameter, generative autoregressive streaming text-to-speech model released by NineNineSix under Apache 2.0 and designed for real-time dialogue on standard LLM-serving infrastructure. NineNineSix reports 68.7 ms median time-to-first-audio, 85.8 ms p90, and 5.23% WER from its own run of Coval’s public benchmark on one RTX 4090. The supplied snippets do not independently verify that Coval result, and NineNineSix’s Hugging Face table reports a different WER of 3.6% under an apparently separate evaluation, so the benchmark conditions and comparative “leading” claim remain insufficiently established here.
Why it matters to Scott
Gepard converges with Scott’s active work on low-latency, locally served TTS and is a concrete Apache-2.0 candidate for benchmarking against Kokoro, Parler-TTS and F5-TTS on his GPU stack. It could change model selection for his voice systems, but the claimed latency and WER remain vendor-reported and need controlled comparison before carrying more weight.
dev:project.audiodev:project.gamepcdev:technology.kokoro-ttsdev:concept.hardware-aware-local-inferencedev:concept.latency-aware-parallel-tts-chunkingradar:concept.local-inferenceradar:concept.open-modelsradar:concept.text-to-speechradar:concept.local-ttsradar:nari-sub-50ms-ttsradar:indextts-25-local-tts-validation
queries asked of Scott's wikis
- local voice inference latency economics
- open-weight speech models and model sovereignty
- streaming TTS for real-time agent interfaces
- serving speech models with standard LLM infrastructure
- voice-agent latency and evaluation benchmarks
- Apache-2.0 models in local AI stacks
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-28T18:39:17Z
No independent reproduction, implementation uptake, or benchmark clarification arrived within the active window. Gepard remains a potentially useful benchmarking candidate, but its leading-performance claim is still vendor-reported rather than a developing signal.
2026-08-26T17:42:24Z
No independent reproduction or comparable benchmark has arrived; the case remains a useful but vendor-reported candidate rather than evidence of a leading open TTS option. The unchanged discussion cools the case without altering its relevance to Scott’s local voice stack.
2026-08-26T17:34:14Z
grounded: converges/medium — Gepard converges with Scott’s active work on low-latency, locally served TTS and is a concrete Apache-2.0 candidate for benchmarking against Kokoro, Parler-TTS
2026-08-26T17:31:27Z
case created — The open model and specific single-GPU benchmark results form a bounded performance claim with a first-party artifact.
Decision trace
- 08-29 04:39expireNo independent reproduction, implementation uptake, or benchmark clarification arrived within the active window. Gepard remains a potentially useful benchmarking candidate, but its leading-performance
- 08-29 04:39alert_silentThe only delta is minor engagement drift without comments or substantive evidence; it changes neither access nor confidence and can remain in the archive until an independent benchmark or implementati
- 08-29 04:39alert_routeThe only delta is minor engagement drift without comments or substantive evidence; it changes neither access nor confidence and can remain in the archive until an independent benchmark or implementati
- 08-27 03:42repriceNo independent reproduction or comparable benchmark has arrived; the case remains a useful but vendor-reported candidate rather than evidence of a leading open TTS option. The unchanged discussion coo
- 08-27 03:42alert_silentThe only new delta is an inconsequential engagement reobservation; it adds no validation, access change, or implementation evidence and can wait for normal briefing cadence.
- 08-27 03:42alert_routeThe only new delta is an inconsequential engagement reobservation; it adds no validation, access change, or implementation evidence and can wait for normal briefing cadence.
- 08-27 03:40alert_silentNineNineSix has published a reproducible-looking, self-run benchmark for an Apache-2.0 TTS candidate relevant to Scott’s local voice stack, but its headline latency is not directly comparable with Cov
- 08-27 03:40surface_candidateNineNineSix has published a reproducible-looking, self-run benchmark for an Apache-2.0 TTS candidate relevant to Scott’s local voice stack, but its headline latency is not directly comparable with Cov
- 08-27 03:40alert_routeNineNineSix has published a reproducible-looking, self-run benchmark for an Apache-2.0 TTS candidate relevant to Scott’s local voice stack, but its headline latency is not directly comparable with Cov
- 08-27 03:34groundGepard converges with Scott’s active work on low-latency, locally served TTS and is a concrete Apache-2.0 candidate for benchmarking against Kokoro, Parler-TTS and F5-TTS on his GPU stack. It could ch
- 08-27 03:31createThe open model and specific single-GPU benchmark results form a bounded performance claim with a first-party artifact.