2026-10-11 17:11 UTC

NineNineSix claims its Apache-2.0 Gepard 1.0 model reaches 68.7 milliseconds median time-to-first-audio and 5.23% WER on one RTX 4090, which would make it a leading low-latency open TTS option on commodity GPU hardware.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference open-models text-to-speechNineNineSixCoval

What is this?

Gepard 1.0 is a 555M-parameter, generative autoregressive streaming text-to-speech model released by NineNineSix under Apache 2.0 and designed for real-time dialogue on standard LLM-serving infrastructure. NineNineSix reports 68.7 ms median time-to-first-audio, 85.8 ms p90, and 5.23% WER from its own run of Coval’s public benchmark on one RTX 4090. The supplied snippets do not independently verify that Coval result, and NineNineSix’s Hugging Face table reports a different WER of 3.6% under an apparently separate evaluation, so the benchmark conditions and comparative “leading” claim remain insufficiently established here.

Why it matters to Scott

Gepard converges with Scott’s active work on low-latency, locally served TTS and is a concrete Apache-2.0 candidate for benchmarking against Kokoro, Parler-TTS and F5-TTS on his GPU stack. It could change model selection for his voice systems, but the claimed latency and WER remain vendor-reported and need controlled comparison before carrying more weight.
dev:project.audiodev:project.gamepcdev:technology.kokoro-ttsdev:concept.hardware-aware-local-inferencedev:concept.latency-aware-parallel-tts-chunkingradar:concept.local-inferenceradar:concept.open-modelsradar:concept.text-to-speechradar:concept.local-ttsradar:nari-sub-50ms-ttsradar:indextts-25-local-tts-validation
queries asked of Scott's wikis
  • local voice inference latency economics
  • open-weight speech models and model sovereignty
  • streaming TTS for real-time agent interfaces
  • serving speech models with standard LLM infrastructure
  • voice-agent latency and evaluation benchmarks
  • Apache-2.0 models in local AI stacks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditRan our Apache 2.0 Gepard TTS through Coval's public benchmark. 68.7 ms to first audio on one RTX 4090.
LocalLLaMA
ylankgz60
🟧 echo.blog ⭐NineNineSix reports that Gepard 1.0 achieved 68.7 ms p50 and 85.8 ms p90 time-to-first-audio with 5.23% WER in its self-run use of Coval's pNineNineSix——

Interpretation history

Decision trace