Local voice AI · 646 languages · no API keys

Create Voices
with VoiceStudio

Clone a voice from a three-second sample, design a new one from a description, dub video into 646 languages, and turn long recordings into searchable transcripts. Everything renders on your own machine, so your voice data stays yours.

Try the studio console
Free local tier646 languagesCloning, dubbing, transcriptsRuns on your own hardwareNo API keys required
Studio console

One console for cloning, design, dubbing, and transcripts

This is the workflow the desktop studio walks you through: point it at a source, describe the result, pick a local engine, and render. The console on this page runs the same stages as a demonstration, without sending a single request.

Source
reference-clip.wav · 3.2s · 48 kHz
Script

Ship the trailer narration in my own voice, then keep the same timbre for the three follow-up lines.

Engine · xtts-v2local cuda:0OutputWAV 48 kHz
Run logno network requests

Press “Run local render” to walk through a voice clone job.

Speaker 1 · clonedIdle
11,000GitHub stars
294,413Downloads and pulls
646Language catalog
v0.5.0Latest release

Observed Aug 21, 2026 · upstream project snapshot, recorded on this site for reference.

Key capabilities

A full studio, not a single text box

VoiceStudio covers the whole local voice workflow: reference-driven cloning, descriptive voice design, timing-locked dubbing, speaker-aware transcription, and an API for automation.

Everything runs locally

Engines run on your own hardware, so voice samples, scripts, and finished renders never leave your machine and never hit a usage counter.

646 languages

Clone, dub, and transcribe across a catalog of 646 languages that varies by engine, from widely spoken languages to long-tail ones.

No API keys, no cloud bill

Install an engine once and keep using it. There is no per-character billing, no rate limit, and no account needed for local renders.

Voice cloning from a short sample

A few seconds of clean speech are enough to embed a speaker identity you can reuse across narration, dialogue, and dubbing.

Voice design with controls

Steer gender, age, accent, pitch, and emotion instead of hunting for the right reference recording.

Timing-locked dubbing

Speaker separation, translation, and re-voicing keep each line aligned to the original cut, including multi-speaker scenes.

Speaker-aware transcription

Diarisation labels who said what, and captions export as SRT or VTT in every language the source provides.

Vocal isolation and cleanup

Separate speech from music and room tone, then match loudness and de-ess before you export.

OpenAI-compatible local API

Point your own tools at the local API and drive synthesis, transcription, and dubbing from scripts or your own app.

macOS, Linux, WSL, Windows

The same studio and the same projects on every desktop platform you work on, with GPU or CPU execution.

Engine switching

Swap between bundled engines per job, so a quick transcription and a long-form narration can use different models.

Private by design

Nothing is uploaded for a local render. Only hosted plans send work to our infrastructure, and only when you ask for it.

How it works

Install, choose an engine, render

No training run, no dataset upload, and no waiting for a queue. The studio is useful within minutes of installing it.

01

Install the studio

Download VoiceStudio for macOS, Linux, WSL, or Windows and let it detect your GPU. The free tier is a full local install, not a trial.

02

Pick an engine

Choose the engine that fits the job: a cloning engine for narration, a transcription engine for long recordings, or both.

03

Clone, design, or dub

Use a reference clip, describe a voice, or hand over a video. The studio keeps a separate voice profile per speaker.

04

Render and export

Export WAV, MP3, MP4, SRT, or VTT. Projects stay on disk, so you can revisit a job and re-render a single line.

curl -fsSL https://voicestudio.lol/install | shor use the desktop installer for macOS, Windows, and Linux
Engines and platforms

Pick the engine per job, keep the machine you already have

Cloning, design, dubbing, and transcription can each use a different local engine, and heavy jobs fall back to hosted rendering on paid plans.

JobEngineWhere it runsOutput
Voice cloningClone engine · localYour machine (GPU or CPU)WAV, MP3
Voice designDesign engine · localYour machine (GPU or CPU)WAV, MP3
Video dubbingTranscribe + clone · localYour machine (GPU or CPU)MP4, MKV
TranscriptionLarge-v3 transcription · localYour machine (GPU or CPU)TXT, SRT, VTT
Isolation and diarisationSeparation + diarisation · localYour machine (GPU or CPU)WAV, labels
Batch and automationOpenAI-compatible local APIYour machine or your serverJSON, files
Use cases

Creators, developers, and teams

Start local and stay local, or move heavy renders to hosted plans when a deadline matters more than your GPU.

Creators

Narrate your own scripts in your own voice, dub a back catalogue, or cast an audiobook without booking a studio.

  • Voiceovers and trailers
  • Multi-language dubbing
  • Chaptered audiobooks
  • Transcripts for publishing

Developers

Drive synthesis and transcription from code through an OpenAI-compatible local API, with no per-request cost and no data leaving the box.

  • Local API for pipelines
  • Batch jobs over folders
  • Deterministic voice profiles
  • Runs offline in CI or on-prem

Teams

Give a team one studio instead of a stack of subscriptions, with commercial rights and hosted access when local hardware is not enough.

  • Shared projects and presets
  • Commercial usage rights
  • Hosted rendering options
  • Priority support
Pricing

Free to run locally, paid when you need the cloud

The local studio is free forever for personal projects. Pro and Studio add hosted rendering, credits, commercial rights, and support.

Free

$0forever

The full local studio for personal projects, with no account and no usage limit.

Run the studio demo
  • Local engines on your own hardware
  • Voice cloning, design, dubbing, transcripts
  • 646-language catalog
  • No API keys, no cloud bill
  • Personal use

Studio

$39/month

For teams and studios that need concurrency, rights, and support.

See plan details
  • Everything in Pro
  • 4,000 studio credits every month
  • Team seats and shared presets
  • Extended commercial and client rights
  • Higher concurrency and priority queue
  • Priority support

Annual billing saves 17%. Plans are managed in your account area and can be cancelled at any time. Hosted renders consume studio credits; local renders never do.

Questions

VoiceStudio FAQ

Short answers on privacy, licensing, languages, and what the plans include.

Does VoiceStudio run offline?+

Yes. The local engines run on your own machine, and cloning, dubbing, transcription, and design all work without an internet connection. Only hosted plans send work to our infrastructure, and only when you ask for it.

How much audio do I need to clone a voice?+

About three seconds of clean speech is usually enough for a usable voice profile. Longer, quieter reference clips improve stability for long-form narration or emotional delivery.

Can I use the results commercially?+

The Free tier covers personal projects. Pro and Studio plans include commercial usage rights, and Studio extends them to client and team work. You are always responsible for having the rights to the voice you clone.

Which languages are supported?+

The catalog covers 646 languages across the bundled engines, so the exact list depends on the engine you pick for a job. Dubbing and transcription can use different engines per project.

What do the studio credits pay for?+

Credits cover hosted rendering: cloud GPU time, hosted projects, and concurrency. Local renders on the Free tier never consume credits.

Do I need an account to start?+

No. Download the studio and render locally without an account. An account is only needed for hosted rendering, plan management, and billing, and sign-in is Google only.

Turn your script into finished audio

Sign in with Google, pick a plan, and render your first track in the studio console.

Compare plans
Local-first by default Free tier included 646 languages Cancel anytime
VoiceStudio - Local Voice Cloning, Dubbing, and Text to Speech