Neuphonic claims its open-source NeuDecide β a 43MB model that maps audio directly to tool calls without transcription β enables practical voice-enabled agent workflows on edge devices via WASM browser deployment.
state: corroboratedheat: mediumuncertainty: lowconvergesscott: highaudio-to-tool edge-inference agent-harnesses voice-agentsNeuphonicTeamNeuphonic
What is this?
Neuphonic is a voice AI company releasing open-source, on-device speech models (NeuTTS Air, NeuCodec) built on small LLM backbones (Qwen 0.5B) for CPU-native inference. The web results confirm their TTS/voice-cloning models and LiveKit integration, but do not surface a model called 'NeuDecide' that maps audio directly to tool calls β the search returns only NeuTTS Air (text-to-speech) and NeuCodec. The case's central claim (a 43MB audio-to-tool model with WASM browser demo) is not corroborated by the supplied snippets; it may be a newer/ ΰ€
ΰ€²ΰ€ release not yet indexed, or the name may be conflated with NeuTTS Air.
Why it matters to Scott
A 43MB open-weight audio-to-tool model running in-browser via WASM is a concrete arrival at positions Scott's canon has argued for years: the fast lane of the Fast-Slow Split served by tiny local models (ip:framework.fast-slow-split), model sovereignty via open weights that run on edge hardware without vendor permission (ip:framework.sovereign-software-assurance, ip:concept.model-perishability), hardware-aware local inference as explicit runtime policy (dev:concept.hardware-aware-local-inference), and the shallow-action pool executed without round-tripping through ASR (ip:framework.voice-ais-fork). If the release holds up, it gives Scott a dated-receipts opportunity: a consequential other party shipping the architecture he specified.
ip:framework.fast-slow-splitip:framework.voice-ais-forkip:framework.sovereign-software-assuranceip:concept.model-perishabilitydev:concept.hardware-aware-local-inferenceip:framework.micro-agents-architecturedev:concept.padded-cell-agent-architecturedev:concept.deterministic-agent-control-planeip:framework.agent-native-computingdev:project.audioradar:aiope-android-agent-runtimeradar:1dial-real-world-task-agentradar:adaptive-kv-cache-streamingradar:aa-agentperf-local-benchmark
queries asked of Scott's wikis
- audio-to-tool calling without ASR transcription
- edge voice agent architectures WASM browser deployment
- on-device function calling from raw audio
- model sovereignty open-weight voice agents
- local inference economics for voice agents
- agent harness patterns for streaming audio input
Measured heat
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p26momentum: steady2 platformsage 77h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p45 vs 1243 stories at the 72h mark (now 77h old) β ahead of apowerb-open-agent-runtime (1.2x), behind agentic-flooding-public-services (0.9x)
Evidence (2) β β canonical anchor
Interpretation history
2026-10-11T12:23:34Z
Moved to corroborated: first-party release announcement plus a third-party demo (voice racer game) built on NeuDecide provide two independent evidence lines. Concrete benchmarks (72.4% SLURP tool accuracy, 43MB, sub-100ms on mobile/edge) and working WASM demo confirm the core claims. Directly converges with Scott's fast-slow split, sovereign software, hardware-aware inference, and voice-AI-fork frameworks.
2026-10-11T11:36:50Z
evidence attached: hn.story.50041695 β Voice racer game is a concrete demo built with the NeuDecide model the case tracks.
2026-10-08T13:03:19Z
grounded: converges/high β A 43MB open-weight audio-to-tool model running in-browser via WASM is a concrete arrival at positions Scott's canon has argued for years: the fast lane of the F
2026-10-08T12:45:28Z
case created β First-party release of a novel tiny audio-to-tool-calling model with browser-run WASM demo, directly relevant to voice agent workflows and edge inference.
Decision trace
- 10-11 23:37attention_routeThe editor compared this story and chose to keep watching.
- 10-11 23:23attention_candidatematerial_reprice
- 10-11 23:23repriceMoved to corroborated: first-party release announcement plus a third-party demo (voice racer game) built on NeuDecide provide two independent evidence lines. Concrete benchmarks (72.4% SLURP tool accu
- 10-11 22:42attention_routeThe editor compared this story and chose to keep watching.
- 10-11 22:36attention_candidateattach
- 10-11 22:36attachVoice racer game is a concrete demo built with the NeuDecide model the case tracks.
- 10-11 22:36propose_attachVoice racer game is a concrete demo built with the NeuDecide model the case tracks.
- 10-09 18:41feedback_briefingScott vote via UI
- 10-09 18:08attention_communicatedNeuDecide maps raw audio directly to tool calls with arguments (72.4% tool accuracy on SLURP, 3Γ better than Parakeet+FunctionGemma). Model totals 43MB, runs on single CPU thread: 46ms on M3 MacBook,
- 10-09 18:08attention_routeLead story for 6 PM briefing: concrete arrival at positions Scott's canon has argued for years β tiny open-weight model enabling voice-enabled agent workflows on edge devices. High relevance to S
- 10-09 10:25attention_routeConcrete arrival at positions Scott's canon has argued for years: tiny open-weight model enabling voice-enabled agent workflows on edge devices. High relevance to Scott's agent-harness patte
- 10-09 10:06attention_routeConcrete arrival at positions Scott's canon has argued for years: tiny open-weight model enabling voice-enabled agent workflows on edge devices. High relevance to Scott's agent-harness patte
- 10-09 01:03attention_routeConcrete arrival at positions Scott's canon has argued for years: tiny open-weight model enabling voice-enabled agent workflows on edge devices. High relevance to Scott's agent-harness patte
- 10-09 00:56attention_candidatecreate
- 10-09 00:03groundA 43MB open-weight audio-to-tool model running in-browser via WASM is a concrete arrival at positions Scott's canon has argued for years: the fast lane of the Fast-Slow Split served by tiny local
- 10-08 23:45createFirst-party release of a novel tiny audio-to-tool-calling model with browser-run WASM demo, directly relevant to voice agent workflows and edge inference.