Text To Speech
Voicebox is a free, open-source, local-first AI voice studio that lets users clone voices, generate speech, and dictate system-wide entirely on their own machine. It runs seven TTS engines (including Qwen3-TTS and Kokoro) plus Whisper-based transcription, and ships a REST/WebSocket API and MCP server so AI agents like Claude Code or Cursor can speak and listen. It is unrelated to Meta's unreleased research model of the same name, and serves developers, content creators, and accessibility users.
Other tools in our directory you may want to compare.
Also mentioned
Full overview from our catalog (read-only reference).
Editorial notes to help compare fit before opening the vendor site.
Context from the listing review and editorial research.
Launched February 4, 2026; its GitHub repository (jamiepine/voicebox) surpassed 47,000 stars within roughly six months. Note: unrelated to Meta's "Voicebox" speech research model, which was never released publicly — a frequent source of naming confusion.
In-depth description and capability notes.
Voice cloning from a few seconds of reference audio Seven interchangeable TTS engines including Qwen3-TTS and Kokoro System-wide dictation into any app via a global hotkey Built-in MCP server exposing voicebox.speak and voicebox.transcribe to AI agents REST and WebSocket API with no API keys or rate limits Cross-platform local GPU inference (Metal, CUDA, ROCm, Intel Arc, DirectML) Whisper-based transcription supporting 99 languages Multi-track timeline editor for audio production
Developers give their own apps or AI coding agents a spoken voice via the local API or MCP server, while creators and accessibility users rely on it for narration, dictation, and speech-to-text without any cloud subscription.