2026-10-11 16:38 UTC

VoiceStudio’s maintainers claim their released beta integrates voice cloning, dubbing, transcription, and audiobook production with local engines and APIs, potentially replacing hosted audio workflows without accounts or metered inference.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference speech-synthesis speech-recognitiondebpalashVoiceStudio

What is this?

VoiceStudio, previously OmniVoice-Studio, is an open-source audio application maintained under the GitHub account debpalash. Its repository advertises voice cloning, voice design, dubbing, dictation, transcription, and audiobook creation, combining 16 text-to-speech and 11 speech-recognition engines for macOS, Windows, Linux, and Docker. The repository says local workflows need no account, API key, subscription, or usage meter, while users supply hardware and model weights and remote features are opt-in. The snippets establish the project's advertised capabilities, but do not independently verify performance, parity with hosted services, or the timing of the claimed beta release; language coverage varies by engine.

Why it matters to Scott

VoiceStudio’s advertised integrated local-audio workflow converges with Scott’s audio speech-engine laboratory, gamepc voice-cloning setup, and multi-backend podcast generator, offering a concrete candidate to evaluate for reducing custom integration work. The supplied radar hits track adjacent local-speech runtimes, not VoiceStudio itself; hardware fit, output quality, and replacement of hosted workflows remain unverified.
dev:project.audiodev:project.gamepcdev:project.podcastradar:concept.local-audio-inferenceradar:concept.voice-cloningradar:audio-cpp-0-4-local-speech-validation
queries asked of Scott's wikis
  • local inference economics versus metered cloud APIs
  • private offline speech recognition and voice synthesis workflows
  • audiobook production dubbing and content publishing automation
  • swappable model engines unified APIs and tool orchestration
  • self-hosted AI hardware requirements and model licensing

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 639h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 02:24 (minted)⭐ origin echo-reconstructedThe repository presents an active-beta, fully local audio application with 16 TTS engines, 11 ASR engines, desktop and Docker packages, and
debpalash on github (echo) · attributed from hn.story.49706572 · published time unknown
—
09-15 01:26first on hacker news · published · lag ?VoiceStudio – local open-source ElevenLabs alternative
iamsyr
—
09-15 01:26amplified on hacker news 👑hn.story.49706572
iamsyr
peak 6 · 0 comments · 101% of case engagement
09-15 02:20our radar first saw it · lag ?discovery anchor: hn.story.49706572—
pace: p42 vs 1032 stories at the 336h mark (now 639h old) — ahead of agent-memory-add-search-evaluation (1.2x), behind agenticos-self-hosted-governance (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnVoiceStudio – local open-source ElevenLabs alternative
Retrieved article excerpt

Open article · Retrieved 2026-09-15T02:21:52.639549+00:00

### NOTE: Electron Rewrite Ongoing: Please dont't create desktop app related issues and pr

[VoiceStudio logo](https://github.com/debpalash/VoiceStudio/blob/main/docs/logo.png)

# VoiceStudio

[VoiceStudio ranking on Trendshift](https://trendshift.io/repositories/28176?utm_source=repository-badge&utm_medium=badge&utm_campaign=badge-repository-28176)

Previously OmniVoice-Studio

### Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

[Install](https://github.com/debpalash/VoiceStudio#install) ·
[Features](https://github.com/debpalash/VoiceStudio#features) ·
[Compare](https://github.com/debpalash/VoiceStudio#comparison) ·
[Requirements](https://github.com/debpalash/VoiceStudio#requirements) ·
[Hardware](https://github.com/debpalash/VoiceStudio#hardware-recommendations) ·
[Engines](https://github.com/debpalash/VoiceStudio#engines) ·
[Architecture](https://github.com/debpalash/VoiceStudio#architecture) ·
[API](https://github.com/debpalash/VoiceStudio#api) ·
[Docs](https://github.com/debpalash/VoiceStudio#documentation) ·
[FAQ](https://github.com/debpalash/VoiceStudio#faq) ·
[**简体中文**](https://github.com/debpalash/VoiceStudio/blob/main/README_CN.md)

[CI status](https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml)
[GitHub stars](https://github.com/debpalash/VoiceStudio/stargazers)
[Total downloads](https://github.com/debpalash/VoiceStudio/releases)
[Latest release](https://github.com/debpalash/VoiceStudio/releases/latest)
[AGPL-3.0 license](https://github.com/debpalash/VoiceStudio/blob/main/LICENSE)
[Discord community](https://discord.gg/bzQavDfVV9)

[Download VoiceStudio](https://github.com/debpalash/VoiceStudio/releases/latest)

[Switching TTS engines from the VoiceStudio status bar](https://github.com/debpalash/VoiceStudio/blob/main/docs/media/0.5.0/quick-switch.gif)

Warning

**Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work. `main` contains the newest fixes and may change between releases. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues).

## At a glance

|  | VoiceStudio |
| --- | --- |
| **Workflows** | Voice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation |
| **Language catalogue** | 646 TTS languages; actual coverage and quality depend on the selected engine |
| **Engines** | 16 TTS · 11 ASR · switch in Model Catalogue or with `Ctrl`/`Cmd`+`E` |
| **Platforms** | macOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86\_64 with glibc 2.39+ |
| **Compute** | CUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers |
| **Interfaces** | Desktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server |
| **Storage** | Voices, projects, settings, and outputs stay on the machine by default |
| **License** | AGPL-3.0 application; downloaded models keep their upstream terms |

The Voice workspace starts with three tabs: **From audio** for cloning, **By design** for creating a voice, and **Convert** for speech-to-speech conversion. Each tab displays its own workflow, with Synthesize Audio or Convert pinned below the scrolling form. The top-bar **Engines** panel combines engine selection, loaded models, and unload/flush controls; `Ctrl`/`Cmd`+`E` opens it. The searchable language picker shares Dubbing’s flags and language list layout, selects one output language, and retains Auto and the full cloning catalogue. Language options flow into multiple columns when space allows. Expand **Workspaces** in the sidebar to reveal navigation labels; Escape collapses it.

Dubbing starts with file upload or URL import and nearby language choices. Its **Projects** panel lists previous dubs so they can be reopened by clicking anywhere on a card; action buttons operate independently. Advanced import options include captions and optional YouTube sign-in. Dubbing places playback controls over the video with background blur and combines the waveform and timed transcript in one compact editing surface. Drag the zoomed waveform left or right to pan; click to seek. Translation language and ISO-code controls stay synchronized; Auto clears any previous language code and dialect. Transcript items group editable text, timing and status, and voice controls into three readable rows that wrap with the panel width. Output Options stays compact with the active settings shown in its summary; expand it to change output, timing, or voice matching. Transcript, glossary, and paste controls share a toolbar above the segment editor. Project details, workflow steps, and Generate/Verify/Export actions use an unfilled header.

The Audiobook Script editor fills the available workspace beneath its markup toolbar; Voices and Book settings stay in their own tabs.

Output settings use aligned rows; review status appears before the collapsible transcript and glossary. Glossary terms have labelled entry fields and an explicit edit action. Launchpad arranges recent files and saved voices side by side when space allows, with responsive card grids and visible Open actions.

The casting board shows icon-based voice cards and searchable selectors for each speaker. Drag a card onto a speaker or choose a voice from that speaker’s menu.

## Install

Download a package from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest), then follow the platform guide.

| Platform | Package | Guide |
| --- | --- | --- |
| macOS 13.3+ | Apple Silicon DMG | [Install on macOS](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/macos.md) |
| Windows 10/11 | x64 MSI; choose the current-user build when listed to install without admin access | [Install on Windows](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/windows.md#install-pre-built-msi) |
| Linux | AppImage, x86\_64 with glibc 2.39+ | [Install on Linux](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/linux.md) |
| Docker | Linux/AMD64 images; CUDA, ROCm, CPU, and worker-only GPU profiles | [Run with Docker](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/docker.md) |

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/macos.md) instead.

### Quick Docker run

The published images are **`linux/amd64` only**. On Apple Silicon, use the
[native macOS app](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/macos.md) for GPU acceleration. ARM64 hosts
should read the [architecture requirements](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/docker.md#architecture)
before pulling an image.

```
docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable
```

### First voice

1. Launch VoiceStudio and open **Voice Cloning**.
2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
3. Enter text, choose a language, then select **Generate**.

Tip

**Try without installing:** Run VoiceStudio in the cloud via the [Google Colab notebook](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb). Explore audio quality comparisons in [benchmarks](https://github.com/debpalash/VoiceStudio/blob/main/docs/benchmarks.md) and prompt design tips in [expressive speech](https://github.com/debpalash/VoiceStudio/blob/main/docs/expressive-speech.md).

### Audio samples

Listen to sample outputs produced locally with VoiceStudio:

| Workflow | Prompt / Reference Audio | Generated Audio |
| --- | --- | --- |
| **Voice Cloning** | [demo\_voice.wav](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo_voice.wav) | [demo\_clone\_output.wav](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo_clone_output.wav) |
| **Voice Design** (US News Anchor) | *"Clear, authoritative American broadcast tone"* | [demo\_voice\_design\_us\_news\_anchor.wav](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) |
| **Voice Design** (UK Audiobook) | *"Warm, expressive British storytelling voice"* | [demo\_voice\_design\_audiobook\_uk\_narrator.wav](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) |
| **Video Dubbing** (Multilingual) | [source.src.wav](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo/dubbing/source.src.wav) | [Spanish](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [French](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [Japanese](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [Chinese](https://github.com/debpalash/VoiceStudio/blob/main/backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) |

### Run from source

Install the [development prerequisites](https://github.com/debpalash/VoiceStudio/blob/main/.github/CONTRIBUTING.md#development-setup) (Node 20+/Bun and Python 3.11+), then:

```
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop
```

The desktop launcher configures Python dependencies on first run via `uv` automatically. Use `bun run dev` for the browser UI. See [Contributing](https://github.com/debpalash/VoiceStudio/blob/main/.github/CONTRIBUTING.md) for services, tests, and platform packages.

### If setup fails

- Run **Settings → About → Run self-check** or `uv run python backend/main.py --diagnose --deep`.
- Check [install troubleshooting](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/troubleshooting.md).
- Save a scrubbed diagnostic bundle from the app when opening an issue.
- For slow generation, compare [measured benchmarks](https://github.com/debpalash/VoiceStudio/blob/main/docs/benchmarks.md) and [performance settings](https://github.com/debpalash/VoiceStudio/blob/main/docs/performance.md).

## Features

| Area | Included |
| --- | --- |
| **Voice Cloning** | Zero-shot synthesis from a short reference clip ([guide](https://github.com/debpalash/VoiceStudio/blob/main/docs/engines/README.md)) |
| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions ([expressive speech](https://github.com/debpalash/VoiceStudio/blob/main/docs/expressive-speech.md)) |
| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video; compact translation settings include track selection, and completed dubs flag timing issues for review ([export guide](https://github.com/debpalash/VoiceStudio/blob/main/docs/dubbing/export.md)) |
| **Stories and audiobooks** | Multi-voice scripts · EPUB/PDF import · chapter rendering · `.m4b` export |
| **[Dictation Widget](https://github.com/debpalash/VoiceStudio/blob/main/docs/features/dictation.md)** | System-wide shortcut, live transcription, optional local-LLM cleanup |
| **Vocal Isolation** | Demucs speech/background separation |
| **Speaker Diarization** | Pyannote and WhisperX speaker assignment ([guide](https://github.com/debpalash/VoiceStudio/blob/main/docs/features/diarization.md)) |
| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos |
| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models ([catalogue](ht
iamsyr60
🟧 echo.github ⭐The repository presents an active-beta, fully local audio application with 16 TTS engines, 11 ASR engines, desktop and Docker packages, and debpalash——

Interpretation history

Decision trace