VoiceStudio.sh

Make voices. Tell stories. Keep the files.

languages
646
TTS + ASR engines
16 + 11
desktop app + API
1 studio

Overview

Clone, design, dub, dictate, and build audiobooks in one open-source desktop studio. Local-first by default — no subscription, no usage meter; online services stay opt-in.

Capabilities

Own the Voice Pipeline

Clone, design, generate, dub, and dictate from one desktop studio, with the models and output under your control.

Move from Clip to Production

A multi-track dubbing workflow, audiobook editor, project library, voice gallery, and engine controls turn isolated model demos into repeatable work.

Run Where the Work Lives

Use MLX on Apple silicon, CUDA on NVIDIA, ROCm on Linux, Docker on a server — or the OpenAI-compatible API from another tool.

Versus the cloud

The Scoreboard

Same jobs, different terms. This is the table that made me build it.

ElevenLabsVoiceStudio
PricingSubscription + usage meterFree, open source — AGPL-3.0
Voice cloning3s clip3s clip, zero-shot
Voice designGender, ageGender, age, accent, pitch, style, dialect
AudiobooksFull editor — EPUB/PDF in, .m4b out
LanguagesPlan-dependent646
Video dubbingCloud-onlyFully local
Where audio goesTheir serversYour machine — online is opt-in
API keysAccount requiredNone for local work
GPUN/A — it's a browser tabCUDA · Apple Silicon · ROCm · CPU
Desktop appmacOS · Windows · Linux
TTS engines116
ASR engines111
MCP serverClaude, Cursor, any MCP client
SourceClosedFork it, extend it, ship it

Will it run

The System Floor

A GPU is optional. The whole pipeline runs on CPU — slower, but yours.

MinimumRecommended
OSWindows 10 · macOS 13.3+ (Apple Silicon) · Ubuntu 24.04+Any modern 64-bit OS
RAM8 GB16 GB+
VRAM4 GB — TTS auto-offloads to CPU8 GB+ (RTX 3060+)
Disk10 GB free20 GB+ SSD
GPUOptional — CPU worksNVIDIA CUDA · Apple MPS · AMD ROCm (Linux)

Engine room

Sixteen Voices In, Eleven Ears Out

One picker, Cmd/Ctrl+E from anywhere. Defaults are chosen; the rest auto-detect or install on first use.

CountDefaultStandouts
TTS16VoiceStudio — 600+ languages, clone + instructCosyVoice 3 · VoxCPM2 · IndexTTS 2.5 · MLX-Audio on Apple Silicon
ASR11WhisperX — ~100 languages, word-level timingParakeet TDT at ~10× realtime on CPU · FunASR with inline diarization

Drop-in OpenAI/ElevenLabs-compatible API: point your existing SDK at localhost — voice takes your own cloned-profile IDs, model pins an engine per request.