2026-10-11 17:12 UTC

TontaubeV1’s developers claim their released 2.9B open-weight model enables expressive long-form speech, low-latency local inference, and zero-shot voice cloning in English and German, potentially expanding practical self-hosted TTS workflows.

state: expiredheat: lowuncertainty: highknownscott: mediumopen-models local-inference audio-modelsTontaubeV1

What is this?

The case describes TontaubeV1 as a released 2.9B open-weight, character-level TTS model for expressive long-form generation, local inference, and zero-shot voice cloning, particularly in English and German. However, the supplied web results do not substantiate those details: they concern Mistral AI’s distinct 4B Voxtral TTS model and appear to be the source of the claims about nine-language support and roughly 70–90 ms latency. TontaubeV1’s developers, release provenance, actual language coverage, latency, and self-hosting requirements therefore remain unverified from these snippets.

Why it matters to Scott

The radar already tracks this exact development on `radar:tontaubev1-local-longform-tts`. It directly intersects Scott’s local speech-engine lab, self-hosted GPU stack, and long-form TTS work, making it a plausible evaluation candidate, but the supplied evidence does not verify its performance, language coverage, latency, or deployment requirements.
dev:project.audiodev:project.gamepcdev:project.podcastdev:concept.latency-aware-parallel-tts-chunkingradar:tontaubev1-local-longform-ttsradar:concept.local-ttsradar:concept.text-to-speechradar:concept.voice-cloning
queries asked of Scott's wikis
  • self-hosted TTS and local voice-agent stacks
  • open-weight audio models and model sovereignty
  • long-form speech generation workflows
  • zero-shot voice cloning risks and consent
  • latency requirements for conversational agents
  • character-level generation for multilingual speech

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditWe released TontaubeV1, a character-level TTS model for long-form generation [P]
MachineLearning
EAVDR41
🟧 echo.other ⭐The initial Hugging Face release describes TontaubeV1 as a multilingual TTS model for expressive voice cloning, long-form generation, and loFritz Cremer / TontaubeAI——

Interpretation history

Decision trace