Qwen3-TTS is a multilingual text-to-speech model family from the Qwen research team, with 0.6B and 1.7B streaming variants designed for real-time synthesis. Nari Labs claims its optimized serving implementation of the 1.7B CustomVoice model reaches 10 requests per second and sub-50 ms p95 latency, but the supplied evidence does not establish the exact latency metric, hardware, speech-quality trade-offs, or cost. Qwen’s technical report instead lists roughly 101 ms first-packet latency for the 1.7B model at concurrency 1 and 333 ms at concurrency 6, so independent, like-for-like testing is needed to assess Nari Labs’ claim.
Nari Labs’ optimization claim converges with Scott’s active local speech-engine experiments, latency-aware TTS chunking, and evaluation-driven approach to real-time voice systems. A reproducible comparison covering hardware, concurrency, first-audio latency, speech quality, and unit cost could affect model and serving choices for his dental-receptionist work, but the supplied claim alone is not yet actionable.
dev:project.audiodev:concept.latency-aware-parallel-tts-chunkingip:concept.evaluation-driven-developmentwork:concept.dynaquest-ai-dental-receptionistradar:indextts-25-local-tts-validationradar:audio-cpp-0-4-local-speech-validationradar:vllm-h100-config-latency-gainsradar:concept.local-tts
queries asked of Scott's wikis
- voice-agent end-to-end latency budgets
- streaming TTS inference optimization
- self-hosted voice model economics
- latency versus speech-quality trade-offs
- voice-agent benchmarking methodology
- GPU serving concurrency and tail latency
2026-08-25T15:47:41Z
No independent benchmark or implementation result emerged after repeated checks, and the surrounding discussion has faded into repetitive amplification. The claim remains technically testable but no longer warrants an active episode absent fresh reproducible evidence.
2026-08-23T15:40:09Z
The additional discussion is repetitive amplification rather than independent validation; no new benchmark, implementation result, or quality-and-cost evidence changes the vendor-measured status of the sub-50 ms claim.
2026-08-22T11:29:32Z
The refreshed discussion remains repetitive amplification and supplies no independent benchmark or material technical clarification. Nari Labs’ sub-50 ms TTFA claim remains testable but vendor-measured, with practical quality, concurrency, and cost still unresolved.
2026-08-22T01:30:52Z
The refreshed comments remain repetitive practitioner anecdotes and add no independent, reproducible validation of latency, quality, concurrency, hardware, or cost. The case still hinges on a like-for-like benchmark of Nari Labs’ released implementation.
2026-08-22T00:24:32Z
The refreshed discussion remains repetitive practitioner context rather than independent validation or rebuttal. The case still hinges on a reproducible, like-for-like benchmark covering TTFA, hardware, concurrency, quality, and cost.
2026-08-21T23:25:30Z
Refreshed practitioner discussion reinforces that TTS TTFA is only one component of voice-agent latency, but adds no reproducible benchmark or quality-and-cost validation. The sub-50 ms result remains a testable first-party claim rather than independent evidence.
2026-08-21T19:31:21Z
Practitioner comments add skepticism about attainable TTFA and emphasize on-device economics, but provide no reproducible, like-for-like benchmark. The central sub-50 ms quality-and-cost claim remains vendor-measured and uncorroborated.
2026-08-21T17:53:02Z
The refreshed discussion adds only minor amplification and no independent benchmark, implementation result, or clarification of hardware, quality, concurrency, or cost. The released implementation remains testable, but the sub-50 ms claim is still vendor-measured.
2026-08-21T17:39:32Z
grounded: converges/medium — Nari Labs’ optimization claim converges with Scott’s active local speech-engine experiments, latency-aware TTS chunking, and evaluation-driven approach to real-
2026-08-21T17:37:25Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49389952 -> echo.blog.ad3e7a95a8 by Nari Labs Team
2026-08-21T17:36:23Z
case created — The first-party engineering report makes a concrete latency and economics claim relevant to real-time voice-agent infrastructure.