Text To Speech
    Freemium
    Free Trial
    Sponsored

    Voicebox

    Text To Speech

    Voicebox is a free, open-source, local-first AI voice studio that lets users clone voices, generate speech, and dictate system-wide entirely on their own machine. It runs seven TTS engines (including Qwen3-TTS and Kokoro) plus Whisper-based transcription, and ships a REST/WebSocket API and MCP server so AI agents like Claude Code or Cursor can speak and listen. It is unrelated to Meta's unreleased research model of the same name, and serves developers, content creators, and accessibility users.

    Alternatives & similar tools

    Other tools in our directory you may want to compare.

    Also mentioned

    • ElevenLabs (leading cloud voice cloning and TTS platform)
    • WisprFlow (cloud dictation for agents and power users)
    • Coqui and OpenVoice (other open-source voice cloning projects)

    Features & details

    Full overview from our catalog (read-only reference).

    Category (short)
    Text To Speech
    Category
    Text To Speech
    Pricing (CSV)
    Free Trial
    Directory pricing
    Freemium
    Sponsored note
    no
    Target audience
    Developers, privacy-conscious creators, and power users who want a free, fully local alternative to ElevenLabs or WisprFlow with complete API control. Not ideal for: non-technical users wanting a polished hosted cloud product, or teams needing enterprise SLAs and dedicated support.
    Best for
    Developers, privacy-conscious creators, and power users who want a free, fully local alternative to ElevenLabs or WisprFlow with complete API control. Not ideal for: non-technical users wanting a polished hosted cloud product, or teams needing enterprise SLAs and dedicated support.
    Pricing notes
    Verified July 2026 at voicebox.sh and docs.voicebox.sh. Voicebox is entirely free and open source under the MIT license, with no subscription fees, per-character charges, or API rate limits, since all processing runs locally on the user's own hardware.

    Pros & cons

    Editorial notes to help compare fit before opening the vendor site.

    Pros

    • Completely free and open source with no usage limits
    • Runs fully offline, preserving privacy and eliminating per-use costs
    • Supports 7 TTS engines and 23 output languages
    • Native MCP integration lets AI agents speak in cloned voices

    Cons

    • Requires local GPU/compute for best performance
    • Desktop-only, with no hosted or cloud-hosted option
    • Less production-hardened than established commercial cloud services
    • Support is community-based (Discord/GitHub) with no enterprise SLA

    Review notes

    Context from the listing review and editorial research.

    Launched February 4, 2026; its GitHub repository (jamiepine/voicebox) surpassed 47,000 stars within roughly six months. Note: unrelated to Meta's "Voicebox" speech research model, which was never released publicly — a frequent source of naming confusion.

    Extended features

    In-depth description and capability notes.

    Voice cloning from a few seconds of reference audio Seven interchangeable TTS engines including Qwen3-TTS and Kokoro System-wide dictation into any app via a global hotkey Built-in MCP server exposing voicebox.speak and voicebox.transcribe to AI agents REST and WebSocket API with no API keys or rate limits Cross-platform local GPU inference (Metal, CUDA, ROCm, Intel Arc, DirectML) Whisper-based transcription supporting 99 languages Multi-track timeline editor for audio production

    Use cases

    Developers give their own apps or AI coding agents a spoken voice via the local API or MCP server, while creators and accessibility users rely on it for narration, dictation, and speech-to-text without any cloud subscription.