2026-10-11 18:01 UTC

voice-agents

band: hotmomentum: stable score: 1.0
temperature history

Episodes (19)

Follow-up reporting will determine whether Kinney Drugs keeps its AI phone assistant withdrawn or materially redesigns it after hundreds of customer complaints about the deployed service.
expiredconvergesscott: medium
Independent use will determine whether Riffn provides a reliable hands-free mobile voice interface for coding agents and local models beyond ordinary voice-note capture.
expiredknownscott: low
Independent implementations will determine whether OpenAI’s published GPT Live architecture provides a practical low-latency pattern for continuous, responsive voice-agent interaction.
expiredconvergesscott: high
Independent deployments will determine whether Speko’s benchmark-driven routing across speech-to-text, language, and text-to-speech models materially improves voice-agent quality, latency, or cost over fixed vendor stacks.
expiredconvergesscott: medium
Independent deployments will determine whether OpenCyvis provides a practical self-hosted phone-agent stack that works reliably with user-selected local or hosted LLMs.
expiredconvergesscott: medium
Independent benchmarks will determine whether Nari Labs' Qwen3-TTS serving optimizations deliver sub-50-millisecond response latency at practically acceptable speech quality and cost.
expiredconvergesscott: medium
Daily claims its open-weight PhoneLLM Alpha 1 provides a foundation model specialized for low-latency voice-agent workflows, potentially reducing dependence on general-purpose hosted models for conversational audio systems.
expiredconvergesscott: medium
Pipecat AI claims PhoneLLM Alpha matches GPT-5.6 Terra on typical voice-agent tasks at one-third the latency and one-eighteenth the cost, potentially improving the economics of real-time voice agents.
expiredknownscott: medium
1Dial claims its phone-and-text agent has completed more than 100 delegated real-world coordination tasks and can reliably handle calls, appointments, quotes, and follow-ups well enough to reduce users’ manual service work.
expiredknownscott: low
Cortexist claims its open-source Little Gemma CUDA engine runs Gemma 4 E2B voice conversations on Jetson Orin NX faster than llama.cpp without degradation on long voice prompts, potentially enabling sustained local voice agents on edge hardware.
seedknownscott: low
Egma’s builders claim their released platform supports repository-based simulated voice conversations, mocked tool responses, and production grading for LiveKit and Retell agents, enabling repeatable pre-deployment regression testing alongside production monitoring.
watchingconvergesscott: medium
Nari Labs claims its Qwen3-ASR and Qwen3-TTS hosted endpoints achieve 44 ms median final-segment latency and 63 ms median first-audio latency with competitive error rates and pricing, potentially lowering production voice-agent latency and cost as they move to paid general availability.
watchingknownscott: medium
Google claims its released Gemini 3.8 Live models combine uninterrupted voice dialogue with background tool execution and, in Extended Thinking, simultaneous multi-step reasoning, enabling more complex voice-agent workflows without conversational pauses.
corroboratedconvergesscott: medium
Chert claims its released experimental CLI/SDK bridges independently hosted LiveKit agents into FaceTime audio/video calls through Chrome and manual iPhone admission, enabling consumer-call integration without Chert-hosted agent execution.
watchingconvergesscott: medium
Fusion-runtime maintainer SamarthUrs18 claims the released single-process speech-to-text, LLM, and text-to-speech stack delivers roughly 991-millisecond end-of-speech response latency with interruption handling on an RTX 3090, potentially simplifying responsive self-hosted voice-agent deployment.
seedconvergesscott: medium
VOYGR's ex-Google founders claim their YC W26 PlaceCall API lets agents call businesses by voice and complete real-world tasks like verifying hours, making agentic telephony a production action surface for real-world agent work.
corroboratedconvergesscott: high
ElevenLabs claims its released Eleven v4 and v4 Turbo — a new expressive TTS architecture with ~100ms-median-latency Turbo aimed at agents, 10-second instant voice cloning, IPA pronunciation control, and 90+ languages — sets a new commercial speech standard; adoption in agent and media workflows settles whether v4 becomes the leading speech model.
watchingconvergesscott: medium
Speakrail's creator claims the released open-source full-duplex voice stack — Voxtral Realtime with an 80ms turn-taking head, a microturn-tuned Gemma 4 12B, and Breeze TTS 2 on one RTX 4090 — rivals GPT-Live (94.0 Full-Duplex-Bench conversational dynamics, ~0.7–0.8s median reply latency) and becomes a widely adopted self-hosted alternative to hosted realtime voice APIs; independent replication of its results and real adoption confirm it.
seedconvergesscott: high
Neuphonic claims its open-source NeuDecide — a 43MB model that maps audio directly to tool calls without transcription — enables practical voice-enabled agent workflows on edge devices via WASM browser deployment.
corroboratedconvergesscott: high

Trajectory notes