jamiepine/voiceboxPublic

The open-source AI voice studio. Clone, dictate, create.

AI summary: An open-source, local-first voice assistant application focused on privacy and customization.

Stars
49.6K
+272 today
Forks
6.1K
Watchers
232
Open issues
471
Open PRs
117
Contributors
~75
Commits
638
Branches
67

TypeScriptMITCreated Jan 25, 2026Last push 10d agoLatest release v0.5.0+2.1K stars this week+2.3K this month

Star history

since Jan 25, 2026
020K40KJan 2026Mar 2026Jun 2026Aug 2026
49.6K stars as of Aug 7, 2026, tracked back to Jan 25, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 70 commits2026-01-26: 38 commits2026-01-27: 14 commits2026-01-28: 24 commits2026-01-29: 34 commits2026-01-30: 27 commits2026-01-31: 9 commits2026-02-01: 1 commit2026-02-02: 5 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 1 commit2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 1 commit2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 5 commits2026-02-18: 2 commits2026-02-19: 0 commits2026-02-20: 5 commits2026-02-21: 2 commits2026-02-22: 1 commit2026-02-23: 4 commits2026-02-24: 1 commit2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 3 commits2026-02-28: 3 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 1 commit2026-03-05: 0 commits2026-03-06: 3 commits2026-03-07: 2 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 6 commits2026-03-13: 38 commits2026-03-14: 16 commits2026-03-15: 39 commits2026-03-16: 50 commits2026-03-17: 23 commits2026-03-18: 6 commits2026-03-19: 8 commits2026-03-20: 2 commits2026-03-21: 1 commit2026-03-22: 2 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 2 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 23 commits2026-04-17: 1 commit2026-04-18: 19 commits2026-04-19: 4 commits2026-04-20: 8 commits2026-04-21: 5 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 2 commits2026-04-26: 1 commit2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 2 commits2026-06-27: 3 commits2026-06-28: 5 commits2026-06-29: 1 commit2026-06-30: 4 commits2026-07-01: 0 commits2026-07-02: 1 commit2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 1 commit2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 8 commits2026-07-21: 11 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 14 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
562 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    49,644 stars

  • Permissive license

    MIT

  • Repeat trending

    19 trending appearances

What voicebox does

Voicebox is a desktop application that provides a fully local, customizable voice assistant. It leverages local, lightweight models for wake-word detection, speech-to-text, and natural language processing, ensuring that user audio data never leaves their machine. The project emphasizes modularity, allowing users to swap out specific models or integrate custom skills and APIs. It tailors the assistant's capabilities to specific user needs without relying on cloud services.

It is for privacy advocates, tinkerers, and developers wanting a fully controllable, local voice assistant. Users who want to integrate voice commands into their specific desktop workflows will find it highly adaptable.

  • Local on-device execution: All processing happens on-device, maximizing privacy and offline capability.
  • Modular model architecture: You can easily swap wake-word, STT, or LLM engines based on hardware capabilities.
  • Custom extensible skills: It includes a plugin system for adding new commands and API integrations.
  • Cross-platform compatibility: It is built to run on Windows, macOS, and Linux desktop environments.
  • Low latency responses: It is optimized for fast response times by avoiding cloud round-trips.

Where teams use it

Privacy-focused Assistance

Using a voice assistant without transmitting audio to large tech companies.

Smart Home Control

Integrating with local Home Assistant setups for offline device management.

Developer Workflow Automation

Creating voice commands for executing scripts or managing local development environments.

Accessibility

Providing a customizable, always-on voice interface for navigating the OS.

Getting started: npm run dev

README

main branch

Voicebox

Voicebox

The open-source AI voice studio.
Clone any voice. Generate speech. Dictate into any app. Talk to agents in voices you own.
The full voice I/O stack, running locally on your machine.

Downloads Release Stars License Ask DeepWiki

jamiepine%2Fvoicebox | Trendshift

voicebox.shDocsDownloadFeaturesAPITroubleshooting


Voicebox App Screenshot

Click the image above to watch the demo video on voicebox.sh


Voicebox Screenshot 2

Voicebox Screenshot 3


What is Voicebox?

Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app. Clone voices from a few seconds of audio, generate speech in 23 languages across 7 TTS engines, dictate into any text field with a global hotkey, and give any MCP-aware AI agent a voice of your choosing.

The two cloud incumbents sit on opposite halves of the voice I/O loop — ElevenLabs on output, WisprFlow on input. Voicebox does both, bridges them with a bundled local LLM for refinement and per-profile personas, and runs the whole thing on your machine.

  • Complete privacy — models, voice data, and captures never leave your machine
  • 7 TTS engines — Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, and Kokoro
  • Voice cloning and preset voices — zero-shot cloning from a reference sample, or 50+ curated preset voices via Kokoro and Qwen CustomVoice
  • 23 languages — from English to Arabic, Japanese, Hindi, Swahili, and more
  • Post-processing effects — pitch shift, reverb, delay, chorus, compression, and filters
  • Expressive speech — paralinguistic tags like [laugh], [sigh], [gasp] via Chatterbox Turbo; natural-language delivery control via Qwen CustomVoice
  • Unlimited length — auto-chunking with crossfade for scripts, articles, and chapters
  • Stories editor — multi-track timeline for conversations, podcasts, and narratives
  • Voice input — global dictation hotkey with push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, in-app mic on every text field, Whisper-based STT
  • Agent voice output — one tool call (voicebox.speak) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned
  • Voice personalities — attach a free-form persona to any voice profile, then Compose, Rewrite, or Respond via a bundled local LLM — agents can invoke the same modes over MCP
  • API-first — REST API plus a built-in MCP server for integrating voice I/O into your own apps and agents
  • Native performance — built with Tauri (Rust), not Electron
  • Runs everywhere — macOS (MLX/Metal), Windows (CUDA), Linux, AMD ROCm, Intel Arc, Docker

Download

Platform Download
macOS (Apple Silicon) Download DMG
macOS (Intel) Download DMG
Windows Download MSI
Docker docker compose up

View all binaries →

Linux — Pre-built binaries are not yet available. See voicebox.sh/linux-install for build-from-source instructions.

Having trouble? See the Troubleshooting Guide for common install, generation, model-download, and GPU issues.


Features

Multi-Engine Voice Cloning

Seven TTS engines with different strengths, switchable per-generation:

Engine Languages Strengths
Qwen3-TTS (0.6B / 1.7B) 10 High-quality multilingual cloning, delivery instructions ("speak slowly", "whisper")
Qwen CustomVoice 10 9 curated preset voices with natural-language delivery control — no reference audio required
LuxTTS English Lightweight (~1GB VRAM), 48kHz output, 150x realtime on CPU
Chatterbox Multilingual 23 Broadest language coverage — Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, Turkish and more
Chatterbox Turbo English Fast 350M model with paralinguistic emotion/sound tags
TADA (1B / 3B) 10 HumeAI speech-language model — 700s+ coherent audio, text-acoustic dual alignment
Kokoro 8 50 curated preset voices, tiny 82M model, fast CPU inference

Emotions & Paralinguistic Tags

Only Chatterbox Turbo interprets paralinguistic tags like [laugh] and [sigh]. Qwen3-TTS, LuxTTS, Chatterbox Multilingual, and HumeAI TADA read them literally as text.

With Chatterbox Turbo selected, type / in the text input to open the tag inserter and add expressive tags inline with speech:

[laugh] [chuckle] [gasp] [cough] [sigh] [groan] [sniff] [shush] [clear throat]

Post-Processing Effects

8 audio effects powered by Spotify's pedalboard library. Apply after generation, preview in real time, build reusable presets.

Effect Description
Pitch Shift Up or down by up to 12 semitones
Reverb Configurable room size, damping, wet/dry mix
Delay Echo with adjustable time, feedback, and mix
Chorus / Flanger Modulated delay for metallic or lush textures
Compressor Dynamic range compression
Gain Volume adjustment (-40 to +40 dB)
High-Pass Filter Remove low frequencies
Low-Pass Filter Remove high frequencies

Ships with 4 built-in presets (Robotic, Radio, Echo Chamber, Deep Voice) and supports custom presets. Effects can be assigned per-profile as defaults.

Unlimited Generation Length

Text is automatically split at sentence boundaries and each chunk is generated independently, then crossfaded together. Works with all engines.

  • Configurable auto-chunking limit (100–5,000 chars)
  • Crossfade slider (0–200ms) for smooth transitions
  • Max text length: 50,000 characters
  • Smart splitting respects abbreviations, CJK punctuation, and [tags]

Generation Versions

Every generation supports multiple versions with provenance tracking:

  • Original — clean TTS output, always preserved
  • Effects versions — apply different effects chains from any source version
  • Takes — regenerate with a new seed for variation
  • Source tracking — each version records its lineage
  • Favorites — star generations for quick access

Async Generation Queue

Generation is non-blocking. Submit and immediately start typing the next one.

  • Serial execution queue prevents GPU contention
  • Real-time SSE status streaming
  • Failed generations can be retried
  • Stale generations from crashes auto-recover on startup

Voice Profile Management

  • Create profiles from audio files or record directly in-app
  • Import/export profiles to share or back up
  • Multi-sample support for higher quality cloning
  • Per-profile default effects chains
  • Organize with descriptions and language tags

Stories Editor

Multi-voice timeline editor for conversations, podcasts, and narratives.

  • Multi-track composition with drag-and-drop
  • Inline audio trimming and splitting
  • Auto-playback with synchronized playhead
  • Version pinning per track clip

Global Dictation & Voice Input

The other half of the voice I/O loop. Hold a hotkey anywhere on your system, speak, release — on macOS the transcript pastes straight into the focused text field. Or hit the mic on any Voicebox text input and dictate directly into the app.

  • Configurable chord bindings — hold-to-speak and tap-to-toggle chords, each rebindable in the in-app chord picker. Holding push-to-talk and tapping Space mid-hold upgrades into a toggle session without a gap in audio
  • Target-aware paste (macOS) — accessibility-verified injection into the focused text field, with atomic clipboard save/restore so your clipboard isn't clobbered
  • First-run permissions UX — in-app gates walk you through the macOS Accessibility and Input Monitoring grants with deep-links to System Settings
  • In-app mic button on every Voicebox text field — generation form, profile descriptions, story titles, anywhere you'd type
  • LLM refinement — optional cleanup of ums, stutters, and false starts before paste
  • On-screen pill — floating overlay surfacing recording, transcribing, refining, and speaking states. Same pill agents use when they speak to you, so there's one mental model for both directions of the loop

Speech-to-Text

Voicebox runs OpenAI Whisper for transcription — the same model that backs dictation, the Captures tab, and the /transcribe API. Running on MLX (Apple Silicon) or PyTorch (CUDA / ROCm / DirectML / CPU) depending on your platform.

Size Notes
Base / Small / Medium / Large Standard Whisper quality ladder
Turbo ~8x faster than Whisper Large, minimal quality loss

More engines (Parakeet v3, Qwen3-ASR) are planned — see Roadmap.

Captures

Every dictation, in-app recording, and uploaded audio file lands in the Captures tab — original audio paired with transcript, always preserved.

  • Replay, re-transcribe, refine — rerun STT with any Whisper size, or re-run the raw transcript through the local LLM with different flags (filler cleanup, self-correction removal, technical-term preservation)
  • Edit inline — tweak the transcript and save on blur
  • Play as voice profile — turn any capture into speech with a cloned voice, one click
  • Promote to voice sample — use a capture's audio + transcript as a reference sample on any voice profile
  • Local capture storage — original audio and transcript stay in your Voicebox data directory, with a folder shortcut in Settings

Agent Voice Output

Every agent gets a voice. One tool call and any MCP-aware agent can speak to you in a voice you've cloned — task completions, questions, notifications. The same pill that surfaces during dictation surfaces during agent speech, so you always see what's coming out of your machine.

// In any MCP-aware agent:
await voicebox.speak({
  text: "Deploy complete.",
  profile: "Morgan",
});

Also exposed as POST /speak for anything that doesn't speak MCP — ACP, A2A, shell scripts, custom harnesses.

  • Bidirectional pillrecording, transcribing, refining, and speaking are all states of the same OS-level overlay, so dictation and agent speech share one surface
  • Per-agent voice binding — in Settings → MCP, pin Claude Code to Morgan and Cursor to Scarlett so you can tell which agent is talking without looking. Each client's last_seen_at timestamp confirms the install actually took
  • Always visible — no silent background TTS; every agent-initiated speak surfaces the pill with the voice profile name for the full duration
  • HTTP + stdio transports — install as a URL in Claude Code / Cursor / Windsurf / VS Code MCP, or point stdio-only clients at the bundled voicebox-mcp binary

Voice Personalities

Attach a free-form personality to any voice profile — who this voice is, how they speak, what they care about. Two actions appear on the generate box when a personality is set, powered by a bundled Qwen3 LLM running entirely locally.

  • Compose — a shuffle button that drops a fresh in-character line into the textarea; edit and speak, or click again for a different take
  • Speak in character — a toggle that routes your input text through the personality LLM to be rewritten in their voice before TTS

Agents can reach the same rewrite path over MCP by passing personality: true to voicebox.speak, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint.

Local LLM options: Qwen3 0.6B / 1.7B / 4B, sharing the TTS runtime (MLX on Apple Silicon, PyTorch elsewhere).

Use cases: agent dev loops (dictate a question, hear the answer in a cloned voice), interactive characters for games and narrative tools, speech assistance for people who can't speak in their original voice.

Model Management

  • Per-model unload to free GPU memory without deleting downloads
  • Custom models directory via VOICEBOX_MODELS_DIR
  • Model folder migration with progress tracking
  • Download cancel/clear UI

GPU Support

Platform Backend Notes
macOS (Apple Silicon) MLX (Metal) 4-5x faster via Neural Engine
Windows (NVIDIA) PyTorch (CUDA) Auto-downloads CUDA binary from within the app
Linux (NVIDIA) PyTorch (CUDA) Use a local/remote Python backend with CUDA PyTorch
Linux (AMD) PyTorch (ROCm) Auto-configures HSA_OVERRIDE_GFX_VERSION
Windows (any GPU) DirectML Universal Windows GPU support
Intel Arc IPEX/XPU Intel discrete GPU acceleration
Any CPU Works everywhere, just slower

API

Voicebox exposes a REST API for integrating voice I/O into your own apps and agents.

# Generate speech
curl -X POST http://127.0.0.1:17493/generate \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello world", "profile_id": "abc123", "language": "en"}'

# Agent voice output — any app or script can speak in a cloned voice
curl -X POST http://127.0.0.1:17493/speak \
  -H "Content-Type: application/json" \
  -H "X-Voicebox-Client-Id: my-script" \
  -d '{"text": "Deploy complete.", "profile": "Morgan"}'

# Transcribe an audio file
curl -X POST http://127.0.0.1:17493/transcribe \
  -F "audio=@recording.wav" \
  -F "model=whisper-turbo"

# List voice profiles
curl http://127.0.0.1:17493/profiles

POST /speak accepts profile as a name (case-insensitive) or id, and resolves via the same precedence as the MCP tool: explicit arg → per-client binding → capture_settings.default_playback_voice_id.

MCP server

Voicebox ships a built-in Model Context Protocol server so any MCP-aware agent (Claude Code, Cursor, Windsurf, Cline, VS Code MCP extensions) can speak, transcribe, and browse captures and profiles.

Claude Code one-liner:

claude mcp add voicebox \
  --transport http \
  --url http://127.0.0.1:17493/mcp \
  --header "X-Voicebox-Client-Id: claude-code"

Any HTTP MCP client (Cursor, Windsurf, VS Code, etc.):

{
  "mcpServers": {
    "voicebox": {
      "url": "http://127.0.0.1:17493/mcp",
      "headers": { "X-Voicebox-Client-Id": "cursor" }
    }
  }
}

Stdio fallback for clients that don't speak HTTP MCP — point at the bundled voicebox-mcp binary inside the app:

{
  "mcpServers": {
    "voicebox": {
      "command": "/Applications/Voicebox.app/Contents/MacOS/voicebox-mcp",
      "env": { "VOICEBOX_CLIENT_ID": "claude-desktop" }
    }
  }
}

Four tools ship: voicebox.speak, voicebox.transcribe, voicebox.list_captures, voicebox.list_profiles. Per-client voice bindings are managed in Voicebox → Settings → MCP. See the full MCP guide for tool signatures, resolution precedence, the speaking-pill contract, and security notes.

// In any MCP-aware agent:
await voicebox.speak({
  text: "Tests passing. Ready to merge.",
  profile: "Morgan",      // optional — falls back to the per-client binding
  personality: true,      // optional — rewrites text through the profile's personality LLM first
});

Use cases: agent dev loops (voice in, voice out), game dialogue, podcast production, accessibility tools, voice assistants, content automation.

Full API documentation available at http://127.0.0.1:17493/docs.


Tech Stack

Layer Technology
Desktop App Tauri (Rust)
Frontend React, TypeScript, Tailwind CSS
State Zustand, React Query
Backend FastAPI (Python)
TTS Engines Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox, Chatterbox Turbo, TADA, Kokoro
STT Whisper / Whisper Turbo (PyTorch or MLX)
Local LLM Qwen3 (0.6B / 1.7B / 4B), shared runtime with TTS / STT
MCP Server FastMCP mounted at /mcp (Streamable HTTP) + bundled stdio shim binary
Native Shim Rust (inside Tauri) for global hotkey, paste injection, focus introspection
Effects Pedalboard (Spotify)
Inference MLX (Apple Silicon) / PyTorch (CUDA/ROCm/XPU/CPU)
Database SQLite
Audio WaveSurfer.js, librosa

Roadmap

Feature Description
Windows / Linux auto-paste Dictation paste parity — SendInput on Windows, uinput / AT-SPI on Linux
STT engine expansion Parakeet v3 and Qwen3-ASR joining Whisper — 50+ languages, better non-English quality
Pipeline routing Configurable source → transform → sink chains with webhook + MCP sinks and a preset editor
Streaming transcription WebSocket /transcribe/stream for partial transcripts as you speak
End-to-end speech LLMs Moshi, GLM-4-Voice, Qwen2.5 Omni — real voice-to-voice, no text between
Voice Design Create new voices from text descriptions
Long-form capture Dual-stream recorder (mic + system audio) with summary LLM transform
Platform sinks Apple Notes, Obsidian, and other opt-in integrations
Plugin architecture Extend with custom models, transforms, and sinks
Mobile companion Control Voicebox from your phone

For the full engineering status, open-issue triage, and prioritized work queue, see docs/PROJECT_STATUS.md — a living document that tracks what's shipped, what's in-flight, candidate TTS engines under evaluation, and why we've accepted or backlogged specific integrations.


Development

See CONTRIBUTING.md for detailed setup and contribution guidelines.

Quick Start

git clone https://github.com/jamiepine/voicebox.git
cd voicebox

just setup   # creates Python venv, installs all deps
just dev     # starts backend + desktop app

Install just: brew install just or cargo install just. Run just --list to see all commands.

Prerequisites: Bun, Rust, Python 3.11+, Tauri Prerequisites, and Xcode on macOS.

The repo ships a pre-wired .mcp.json at the root — running Claude Code inside this checkout picks up the Voicebox MCP tools automatically once the dev app is running.

Building Locally

just build          # Build CPU server binary + Tauri app
just build-local    # (Windows) Build CPU + CUDA server binaries + Tauri app

Adding New Voice Models

The multi-engine architecture makes adding new TTS engines straightforward. A step-by-step guide covers the full process: dependency research, backend protocol implementation, frontend wiring, and PyInstaller bundling.

The guide is optimized for AI coding agents. An agent skill can pick up a model name and handle the entire integration autonomously — you just test the build locally.

Project Structure

voicebox/
├── app/              # Shared React frontend
├── tauri/            # Desktop app (Tauri + Rust)
├── web/              # Web deployment
├── backend/          # Python FastAPI server
├── landing/          # Marketing website
└── scripts/          # Build & release scripts

Contributing

Contributions welcome! See CONTRIBUTING.md for guidelines.

  1. Fork the repo
  2. Create a feature branch
  3. Make your changes
  4. Submit a PR

Security

Found a security vulnerability? Please report it responsibly. See SECURITY.md for details.


License

MIT License — see LICENSE for details.


voicebox.sh

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

25 total
  1. v0.5.0v0.5.0Apr 25, 20261.5M downloads

    ## The Capture release. Voicebox stops being just a voice-cloning studio and becomes a full AI voice studio. Hold a key anywhere on your machine, speak, release — the transcript lands in the focused text field. Flip the primitive around and any MCP-aware agent — Claude Code, Cursor, Spacebot — speaks back through an on-screen pill in one of your cloned voices. A local LLM sits between the two, so transcripts come out clean and voice profiles can carry a personality that reshapes what the agent says before it gets spoken. <img width="1354" height="990" alt="Screenshot 2026-04-22 at 2 37 10 PM" src="https://github.com/user-attachments/assets/853c730f-64f3-4573-bcd8-1000bdb8e11f" /> ### Dictation — speak anywhere, paste anywhere - **Global hotkey capture.** Hold a customizable chord anywhere on your machine (defaults: right-Cmd + right-Option on macOS, right-Ctrl + right-Shift on Windows), speak, release. A floating on-screen pill walks through recording → transcribing → refining → done with a live elapsed timer. The transcript lands as clean text. - **Push-to-talk and toggle modes, each with its own chord.** The default toggle chord adds Space to the push-to-talk chord.

  2. v0.4.5v0.4.5Apr 22, 202661.7K downloads

    Second hotfix for the "offline mode is enabled" crash on model load. 0.4.4 reverted the inference-path offline guards but kept the same trap on the load path, so users who updated to 0.4.4 kept hitting the exact error the release was supposed to fix ([#526](https://github.com/jamiepine/voicebox/issues/526)). This release removes the load-path guards and patches the transformers tokenizer load to be robust to HuggingFace metadata failures at the source, so the class of bug can't recur. ### Reliability - **Load no longer fails with "offline mode is enabled"** ([#530](https://github.com/jamiepine/voicebox/pull/530), fixes [#526](https://github.com/jamiepine/voicebox/issues/526)). transformers 4.57.x added an unconditional `huggingface_hub.model_info()` call inside `AutoTokenizer.from_pretrained` (via `_patch_mistral_regex`) that runs for every non-local repo load, regardless of cache state or whether the target model is actually a Mistral variant. The load-time `HF_HUB_OFFLINE` guard from 0.4.2 turned that into a hard crash for cached online users the moment 0.4.4 removed the inference-path guard that had been masking the problem. Fix wraps `_patch_mistral_regex` so any exceptio

  3. v0.4.4v0.4.4Apr 21, 202615.7K downloads

    Hotfix for a regression in 0.4.3 where generation and transcription could fail outright with "offline mode is enabled" even when the user was online. ### Reliability - **Inference no longer fails with "offline mode is enabled" while online** ([#524](https://github.com/jamiepine/voicebox/pull/524), reverts the inference-path guards from [#503](https://github.com/jamiepine/voicebox/pull/503)). 0.4.3 wrapped every inference body (`generate`, `transcribe`, `create_voice_clone_prompt`) with a process-wide `HF_HUB_OFFLINE` flip to stop lazy HuggingFace lookups from hanging when the network drops mid-inference ([#462](https://github.com/jamiepine/voicebox/issues/462)). That flag also blocks legitimate metadata calls (e.g. `HfApi().model_info` for revision resolution) so online users started seeing generation fail outright. Inference now runs with the process's default HF state. Load-time offline guards — which weren't the source of the regression — stay in place. **Known caveat**: users generating without an internet connection may see brief pauses during inference while HuggingFace metadata lookups time out (typically ~30s, after which the library recovers). A proper offline-mod

  4. voicebox v0.4.3v0.4.3Apr 21, 20264.5K downloads

    A patch focused on two user-impacting reliability fixes: macOS DMG notarization (unblocks `brew install voicebox` on macOS 15 Sequoia and fixes spurious "app isn't signed" Gatekeeper dialogs on older Intel Macs) and Kokoro Japanese voice initialization on fresh installs. ### macOS - **DMGs are now notarized and stapled** ([#523](https://github.com/jamiepine/voicebox/pull/523)). Tauri's bundler notarizes the `.app` inside the DMG but ships the DMG wrapper itself unnotarized. Gatekeeper rejects that on macOS 15 Sequoia (confirmed by Homebrew Cask CI failing on both arm and intel Sequoia runners) and causes the "the app is not signed" dialog on older Intel Macs when Apple's notarization servers are slow or unreachable ([#509](https://github.com/jamiepine/voicebox/issues/509)). The release workflow now submits each DMG to `notarytool`, staples the ticket, verifies with `spctl`, and overwrites the draft-release asset `tauri-action` uploaded. Adds ~5-10 min per macOS job. ### Backend - **Kokoro Japanese voices no longer crash on fresh installs** ([#521](https://github.com/jamiepine/voicebox/pull/521), fixes [#514](https://github.com/jamiepine/voicebox/issues/514)). `misaki[ja

  5. voicebox v0.4.2v0.4.2Apr 21, 20265.3K downloads

    This release localizes the entire app. English, Simplified Chinese (zh-CN), Traditional Chinese (zh-TW), and Japanese (ja) are wired up end-to-end across every tab, modal, dialog, and toast — 559 translation keys per locale, parity verified. Plus a batch of reliability fixes: offline-mode now actually stays offline, Chatterbox accepts reference samples it used to reject, MLX Qwen 0.6B points at the right repo, and macOS system audio survives backgrounding. ### Internationalization ([#508](https://github.com/jamiepine/voicebox/pull/508)) - **i18next foundation** with an in-app language switcher that re-renders the tree on change — lazy-loaded components were holding stale strings without an explicit key-bump on the React root. - **Four locales** at full coverage: English, Simplified Chinese, Traditional Chinese, Japanese. No partial/English-fallback surfaces. - **Every user-visible surface translated**: Stories (list, content editor, dialogs, toasts), Effects (list, detail, chain editor, built-in preset names), Voices (table, search, inspector, Create/Edit modal, audio sample panels), Audio Channels (list, dialogs, device picker), history + story dropdown menus, ProfileCard /

Code frequency

additions and deletions
+118K-118KWeek of 2026-01-25: +118,031 linesWeek of 2026-01-25: -39,006 linesWeek of 2026-02-01: +8,925 linesWeek of 2026-02-01: -19,755 linesWeek of 2026-02-08: +1 linesWeek of 2026-02-08: -1 linesWeek of 2026-02-15: +664 linesWeek of 2026-02-15: -163 linesWeek of 2026-02-22: +681 linesWeek of 2026-02-22: -34 linesWeek of 2026-03-01: +626 linesWeek of 2026-03-01: -117 linesWeek of 2026-03-08: +16,613 linesWeek of 2026-03-08: -3,300 linesWeek of 2026-03-15: +29,608 linesWeek of 2026-03-15: -23,139 linesWeek of 2026-03-22: +3 linesWeek of 2026-03-22: -4 linesWeek of 2026-03-29: +82 linesWeek of 2026-03-29: -50 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +6,643 linesWeek of 2026-04-12: -3,935 linesWeek of 2026-04-19: +26,143 linesWeek of 2026-04-19: -2,419 linesWeek of 2026-04-26: +4 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +2,088 linesWeek of 2026-06-21: -591 linesWeek of 2026-06-28: +6,763 linesWeek of 2026-06-28: -440 linesWeek of 2026-07-05: +505 linesWeek of 2026-07-05: -5 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +4,423 linesWeek of 2026-07-19: -41 linesWeek of 2026-07-26: +589 linesWeek of 2026-07-26: -132 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesJan 25, 2026Aug 2, 2026
+222.4K lines added, -93.1K removed over the last year.

Commits per week

last 52 weeks
2160Week of 2025-08-09: 0 commitsWeek of 2025-08-16: 0 commitsWeek of 2025-08-23: 0 commitsWeek of 2025-08-30: 0 commitsWeek of 2025-09-06: 0 commitsWeek of 2025-09-13: 0 commitsWeek of 2025-09-20: 0 commitsWeek of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 216 commitsWeek of 2026-02-01: 7 commitsWeek of 2026-02-08: 1 commitsWeek of 2026-02-15: 14 commitsWeek of 2026-02-22: 12 commitsWeek of 2026-03-01: 6 commitsWeek of 2026-03-08: 60 commitsWeek of 2026-03-15: 129 commitsWeek of 2026-03-22: 2 commitsWeek of 2026-03-29: 2 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 43 commitsWeek of 2026-04-19: 19 commitsWeek of 2026-04-26: 1 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 5 commitsWeek of 2026-06-28: 11 commitsWeek of 2026-07-05: 1 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 19 commitsWeek of 2026-07-26: 14 commitsWeek of 2026-08-02: 0 commitsAug 9, 2025Aug 2, 2026
562 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 2 commitsSun 1:00 — 1 commitsSun 2:00 — 2 commitsSun 3:00 — 3 commitsSun 4:00 — 11 commitsSun 5:00 — 1 commitsSun 6:00 — 3 commitsSun 7:00 — 2 commitsSun 8:00 — 5 commitsSun 9:00 — 7 commitsSun 10:00 — 8 commitsSun 11:00 — 6 commitsSun 12:00 — 9 commitsSun 13:00 — 6 commitsSun 14:00 — 4 commitsSun 15:00 — 2 commitsSun 16:00 — 4 commitsSun 17:00 — 8 commitsSun 18:00 — 3 commitsSun 19:00 — 13 commitsSun 20:00 — 4 commitsSun 21:00 — 13 commitsSun 22:00 — 5 commitsSun 23:00 — 2 commitsMon 0:00 — 9 commitsMon 1:00 — 11 commitsMon 2:00 — 9 commitsMon 3:00 — 11 commitsMon 4:00 — 8 commitsMon 5:00 — 4 commitsMon 6:00 — 2 commitsMon 7:00 — 1 commitsMon 8:00 — 3 commitsMon 9:00 — 1 commitsMon 10:00 — 1 commitsMon 11:00 — 3 commitsMon 12:00 — 7 commitsMon 13:00 — 1 commitsMon 14:00 — 7 commitsMon 15:00 — 0 commitsMon 16:00 — 7 commitsMon 17:00 — 13 commitsMon 18:00 — 1 commitsMon 19:00 — 3 commitsMon 20:00 — 6 commitsMon 21:00 — 3 commitsMon 22:00 — 10 commitsMon 23:00 — 9 commitsTue 0:00 — 5 commitsTue 1:00 — 5 commitsTue 2:00 — 3 commitsTue 3:00 — 4 commitsTue 4:00 — 5 commitsTue 5:00 — 0 commitsTue 6:00 — 7 commitsTue 7:00 — 4 commitsTue 8:00 — 0 commitsTue 9:00 — 6 commitsTue 10:00 — 1 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 6 commitsTue 14:00 — 2 commitsTue 15:00 — 0 commitsTue 16:00 — 4 commitsTue 17:00 — 6 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 0 commitsTue 21:00 — 1 commitsTue 22:00 — 3 commitsTue 23:00 — 0 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 1 commitsWed 8:00 — 0 commitsWed 9:00 — 1 commitsWed 10:00 — 2 commitsWed 11:00 — 1 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 3 commitsWed 15:00 — 3 commitsWed 16:00 — 1 commitsWed 17:00 — 1 commitsWed 18:00 — 0 commitsWed 19:00 — 3 commitsWed 20:00 — 6 commitsWed 21:00 — 0 commitsWed 22:00 — 6 commitsWed 23:00 — 3 commitsThu 0:00 — 2 commitsThu 1:00 — 2 commitsThu 2:00 — 9 commitsThu 3:00 — 5 commitsThu 4:00 — 4 commitsThu 5:00 — 1 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 0 commitsThu 9:00 — 3 commitsThu 10:00 — 4 commitsThu 11:00 — 1 commitsThu 12:00 — 1 commitsThu 13:00 — 0 commitsThu 14:00 — 1 commitsThu 15:00 — 7 commitsThu 16:00 — 6 commitsThu 17:00 — 2 commitsThu 18:00 — 4 commitsThu 19:00 — 9 commitsThu 20:00 — 2 commitsThu 21:00 — 3 commitsThu 22:00 — 0 commitsThu 23:00 — 5 commitsFri 0:00 — 11 commitsFri 1:00 — 4 commitsFri 2:00 — 5 commitsFri 3:00 — 4 commitsFri 4:00 — 7 commitsFri 5:00 — 4 commitsFri 6:00 — 5 commitsFri 7:00 — 2 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 5 commitsFri 11:00 — 2 commitsFri 12:00 — 0 commitsFri 13:00 — 2 commitsFri 14:00 — 4 commitsFri 15:00 — 3 commitsFri 16:00 — 5 commitsFri 17:00 — 4 commitsFri 18:00 — 2 commitsFri 19:00 — 2 commitsFri 20:00 — 3 commitsFri 21:00 — 4 commitsFri 22:00 — 0 commitsFri 23:00 — 3 commitsSat 0:00 — 4 commitsSat 1:00 — 3 commitsSat 2:00 — 6 commitsSat 3:00 — 4 commitsSat 4:00 — 0 commitsSat 5:00 — 2 commitsSat 6:00 — 0 commitsSat 7:00 — 5 commitsSat 8:00 — 2 commitsSat 9:00 — 2 commitsSat 10:00 — 2 commitsSat 11:00 — 2 commitsSat 12:00 — 4 commitsSat 13:00 — 3 commitsSat 14:00 — 0 commitsSat 15:00 — 3 commitsSat 16:00 — 2 commitsSat 17:00 — 3 commitsSat 18:00 — 2 commitsSat 19:00 — 2 commitsSat 20:00 — 1 commitsSat 21:00 — 2 commitsSat 22:00 — 1 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits527 (83%)
Community commits111 (17%)

638 commits in total over the last year.

DateListRankStars gained
Aug 5, 2026daily#13+412
Aug 4, 2026daily#13+412
Jul 21, 2026daily#9+5
Jul 14, 2026daily#13+5
Jul 2, 2026daily#24+20
Jun 30, 2026daily#21+16
Jun 24, 2026daily#16+18
Jun 21, 2026daily#21+23
Apr 18, 2026daily#20+197
Apr 17, 2026daily#18+111
Apr 15, 2026daily#19+121
Apr 14, 2026daily#23+105
Feb 22, 2026daily#9+306
Feb 21, 2026daily#3+547
Feb 20, 2026daily#2+548
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    453.6K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    385.5K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    268.6K stars · Shell