2026-10-11 16:38 UTC

Nari Labs claims its Qwen3-ASR and Qwen3-TTS hosted endpoints achieve 44 ms median final-segment latency and 63 ms median first-audio latency with competitive error rates and pricing, potentially lowering production voice-agent latency and cost as they move to paid general availability.

state: watchingheat: lowuncertainty: highknownscott: mediumvoice-agents inference-serving inference-economicsNari LabsCoval

What is this?

The case concerns Nari Labs’ claimed hosted Qwen3 speech-recognition and text-to-speech endpoints, reporting 44 ms median final-segment latency for ASR and 63 ms median first-audio latency for TTS as they move toward paid general availability. The supplied web snippets establish the underlying Qwen3-TTS as an open multilingual speech-generation model family from Alibaba’s Qwen team, with reported first-packet latency around 97–101 ms at concurrency one. None of the returned snippets corroborates Nari’s endpoints, pricing, launch status, or the attributed Coval September 14 benchmark snapshot; the underlying model’s measurements do not establish or directly refute Nari’s hosted-service claims.

Why it matters to Scott

The radar already tracks Nari’s Qwen3-TTS serving claims in radar:nari-sub-50ms-tts; the added ASR, Coval snapshot and paid-GA claims remain uncorroborated in the supplied grounding, making this a follow-up rather than independent validation. Verified endpoint performance could inform provider comparisons in Scott’s audio laboratory and Twilio STT→LLM→TTS loop, particularly his first-audio chunking strategy, but these segment-level medians do not establish end-to-end conversational latency gains.
dev:project.audiodev:project.twiliodev:concept.latency-aware-parallel-tts-chunkingradar:nari-sub-50ms-ttsradar:concept.inference-latencyradar:concept.text-to-speechradar:concept.speech-to-text
queries asked of Scott's wikis
  • voice-agent projects speech recognition synthesis providers
  • conversational turn-taking end-to-end latency budgets
  • streaming inference serving optimization concurrency
  • open-model self-hosting versus hosted API economics
  • speech benchmarks latency definitions quality cost tradeoffs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 648h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 18:29 (minted)⭐ origin echo-reconstructedNari reports Coval's September 14 snapshot places its public Qwen3 speech endpoints on quality, latency, and cost Pareto frontiers and annou
Nari Labs Team on blog (echo) · attributed from hn.story.49699267 · published time unknown
—
09-14 16:07first on hacker news · published · lag ?Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
toebee
—
09-14 16:07amplified on hacker news 👑hn.story.49699267
toebee
peak 91 · 31 comments · 100% of case engagement
09-14 18:20our radar first saw it · lag ?discovery anchor: hn.story.49699267—
pace: p69 vs 1032 stories at the 336h mark (now 648h old) — ahead of claude-mem-windows-credential-polling (1.0x), behind gpt6-luna-terminal-bench-regression (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
Retrieved article excerpt

Open article · Retrieved 2026-09-14T18:22:36.262441+00:00

## TL;DR

**Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.**

[Coval](https://www.coval.ai/) is a leading provider of voice AI evaluation and benchmarks. They help speech AI agents perform better in production and publish one of the most widely cited benchmarks in the industry.

The [Text-to-Speech (TTS) benchmark](https://benchmarks.coval.ai/tts) evaluates latency from text input to first audible chunk of audio (time-to-first-audio or TTFA) and Word Error Rate (WER). The [Speech-to-Text (STT) benchmark](https://benchmarks.coval.ai/stt) evaluates latency from user’s finalize request to the final text output (time-to-final-segment or TTFS) and Word Error Rate (WER).

TTFA and TTFS are critical for voice agents, where latency can make a voice AI agent feel unresponsive. Low WER is an obvious key factor for model performance as well.

As of mid September 2026, Nari Labs tops both the Speech-to-Text and Text-to-Speech benchmarks. **STT: #1 Latency, #2 WER**. **TTS: #2 Latency, #1 WER**. Note that [Coval’s benchmarks](https://benchmarks.coval.ai/) can fluctuate every 30 minutes[\*](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/#coval-snapshot-note). We only include publicly available endpoints in our rankings and charts.

## Speech-to-Text

**Our [Qwen3-ASR Fast](https://narilabs.com/product/stt/) model is ranked #1 in Time-to-Final-Segment (TTFS), at p50 of 44 ms and WER of 3.6%, placing #2 behind AssemblyAI’s Universal 3.5 Pro at 3.5%.**

The pricing makes it even better. At **$0.12 / hour**, our Fast endpoint ties for the **2nd-lowest price** among models with known public rates in [Coval’s pricing directory](https://benchmarks.coval.ai/pricing). Universal 3.5 Pro costs 3.75× more, and Deepgram Nova 3 costs 2.4× more. Our Standard endpoint would be the cheapest at **$0.06 / hour**.

[Coval STT latency and accuracy: Nari Qwen3-ASR Fast on the Pareto frontier at 44 ms median TTFS and 3.6% WER.Nari Labs](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/assets/stt-latency-accuracy.png)
[Coval STT word error rates: Nari Qwen3-ASR Fast ranks second at 3.6%, behind AssemblyAI Universal 3.5 Pro at 3.5%.Nari Labs](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/assets/stt-word-error-rate.png)

## Text-to-Speech

**Our [Qwen3-TTS Fast](https://narilabs.com/product/qwen3-tts/) model is ranked #2 in Time-to-First-Audio (TTFA), at p50 of 63 ms and WER of 3.8%, coming in at #1.**

The only model with a lower median TTFA than ours is vui from Fluxions, at 49 ms. It is a 300M parameter model, compared to the 1.7B Qwen3-TTS that we serve.

At **$10 per 1M characters**, our Fast endpoint is tied for the **#1 cheapest** model on [Coval’s pricing directory](https://benchmarks.coval.ai/pricing). ElevenLabs Eleven v3 Conversational costs 5x more, and Cartesia Sonic 3.6 costs 6.5x more. Our Standard endpoint would be the cheapest at **$5 per 1M characters**.

[Coval TTS latency and accuracy: Nari Qwen3-TTS Fast on the Pareto frontier at 63 ms median TTFA and 3.8% WER.Nari Labs](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/assets/tts-latency-accuracy.png)
[Coval TTS word error rates: Nari Qwen3-TTS Fast leads at 3.8%.Nari Labs](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/assets/tts-word-error-rate.png)

Interestingly, the official Qwen3 TTS Flash Realtime endpoint sits at 8.8% WER and 692 ms median TTFA. We both serve the same model.

We also surpass Baseten’s dedicated Qwen3-TTS endpoint, which records 6.0% WER and 101 ms median TTFA.

## Get Started

Try both our [Speech-to-Text](https://narilabs.com/product/stt/) and [Text-to-Speech](https://narilabs.com/product/qwen3-tts/) models for free for a limited period of time. We are moving our Public Beta APIs to a paid GA within this week and will provide **$20 in credits** for everyone who has created an account when the switch happens.

[Try STT and TTS](https://app.narilabs.com/)

Need help meeting the latency and capacity requirements of your voice application? [Talk to our engineers](https://cal.com/toby-kim/30min?utm_source=nari-labs&utm_medium=blog&utm_campaign=coval-voice-ai-benchmarks)

[\*](https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/#coval-snapshot-ref) Benchmark values in this post are based on Coval’s **1-day view as of September 14, 2026, at 15:00 UTC**. WER is pooled across datasets. Rankings exclude dedicated inference endpoints. Prices compare Nari’s published rates with known public rates in Coval’s pricing directory.
toebee9131
🟧 echo.blog ⭐Nari reports Coval's September 14 snapshot places its public Qwen3 speech endpoints on quality, latency, and cost Pareto frontiers and annouNari Labs Team——

Interpretation history

Decision trace