2026-10-11 17:14 UTC

inclusionAI claims its released Realtime-Venus-Omni 9B checkpoint continuously watches and listens, decides when to respond, and generates synchronized text and speech, potentially enabling self-hosted audiovisual agents beyond turn-based interaction.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumrealtime-multimodal-agents open-audio-visual-models streaming-inferenceinclusionAI

What is this?

The case describes inclusionAI’s claimed release of Realtime-Venus-Omni, a 9B audiovisual checkpoint that continuously processes audio and video, chooses when to respond, and produces synchronized text and speech. InclusionAI’s publication page lists research attributed to “Inclusion AI, Ant Group,” including unified multimodal perception and generation, but the supplied web snippets do not cover this specific release. The case’s evidence titles point to a Hugging Face repository and a quoted model card mentioning adaptation from MiniCPM-o 4.5; without the card’s contents, the release details, capabilities, licensing, and practical self-hosting requirements remain unverified.

Why it matters to Scott

The claimed continuous listening and selective response converge with Scott’s Ambient conversation copilot, which defaults to silence, and could offer a backend to evaluate for Practice Trainer and his local GPU laboratory rather than another turn-based speech pipeline. This remains a candidate evaluation, not a demonstrated replacement: the supplied evidence does not verify capabilities, licensing or hardware requirements, and the radar’s related voice-interaction episodes do not establish prior coverage of this release.
dev:project.listendev:concept.ambient-conversation-augmentationdev:project.sales-trainerdev:project.gamepcradar:openai-gpt-live-voice-architectureradar:tavus-sparrow-2-conversational-flowradar:concept.multimodal-agentsradar:ling-flash-vl-release
queries asked of Scott's wikis
  • continuous perception agents versus turn-based interaction
  • voice agent turn detection interruption response timing
  • self-hosted multimodal models inference hardware economics
  • streaming audio video agent runtime integration
  • synchronized speech text generation latency evaluation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 553h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 16:23 (minted)⭐ origin echo-reconstructedThe quoted model card says the repository hosts two Realtime-Venus checkpoints, including a 9B Omni model adapted from MiniCPM-o 4.5 that co
inclusionAI on blog (echo) · attributed from reddit.post.1wjtav9 · published time unknown
—
09-18 15:27first on r/LocalLLaMA · published · lag ?inclusionAI/Realtime-Venus · Hugging Face
jacek2023
—
09-18 15:27amplified on r/LocalLLaMA 👑reddit.post.1wjtav9
jacek2023
peak 125 · 36 comments · 100% of case engagement
09-18 16:20our radar first saw it · lag ?discovery anchor: reddit.post.1wjtav9—
pace: p72 vs 1032 stories at the 336h mark (now 553h old) — ahead of intern-s2-397b-release (1.0x), behind anthropic-antspace-deployment-discovery (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditinclusionAI/Realtime-Venus · Hugging Face
LocalLLaMA
jacek202312436
🟧 echo.blog ⭐The quoted model card says the repository hosts two Realtime-Venus checkpoints, including a 9B Omni model adapted from MiniCPM-o 4.5 that coinclusionAI——

Interpretation history

Decision trace