2026-10-11 17:21 UTC

Tavus claims its released Sparrow-2 audio-understanding and turn-taking model provides a new approach to conversational-flow understanding, potentially improving how real-time voice agents decide when to speak or wait beyond conventional noise cancellation.

state: expiredheat: lowuncertainty: highconvergesscott: mediumconversational-agents speech-models turn-takingTavusBrian

What is this?

Tavus has announced Sparrow-2, an audio-native streaming model for deciding when real-time conversational agents should listen, wait, speak, or continue speaking; its research listing names Brian Johnson. Tavus says the model jointly interprets turn-taking, interruptions, backchannels, speaker identity, and environmental audio, using an encoder and causal transformer at a native 10 ms frame rate. The supplied snippets establish the announcement and Tavus's whole-scene-understanding approach, but do not independently demonstrate performance gains over noise cancellation or establish standalone model availability.

Why it matters to Scott

Tavus’s dedicated listen/wait/speak controller converges with the interaction-management side of Scott’s Cognitive Pipelining and The Fast Slow Split, offering a candidate to evaluate for turn-taking in his Twilio voice lab and Practice Trainer—not evidence that Tavus has adopted his separate reasoning/authority lane. Neither performance gains nor standalone availability are established, so this warrants investigation rather than a stack change; the radar’s GPT Live architecture page tracks a related approach, not this Sparrow-2 announcement.
ip:concept.cognitive-pipeliningip:source.the-fast-slow-splitdev:project.twiliodev:project.sales-trainerradar:openai-gpt-live-voice-architectureradar:concept.voice-agents
queries asked of Scott's wikis
  • real-time voice agents turn-taking floor ownership
  • streaming audio inference latency interaction timing
  • multimodal agents prosody speaker identity acoustic context
  • conversational UX interruptions backchannels
  • voice pipeline endpoint detection noise suppression

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Sparrow-2 – Noise cancellation isn't designed for conversational AIcode_brian114
🟧 echo.blog ⭐Tavus developer Brian introduces Sparrow-2 as the team's latest audio-understanding/turn-taking model and a new approach to conversational-fTavus——

Interpretation history

Decision trace