Text To Speech
    Freemium
    Free Trial
    Sponsored

    Fish.audio

    Text To Speech

    Fish Audio is an AI voice platform for expressive text-to-speech, instant voice cloning, and speech-to-text. Its S2.1 Pro model supports emotion tags, long-form narration, and sub-500ms streaming latency. The platform hosts 2M+ community voices and serves creators, developers, and enterprises through a web studio and developer API.

    Alternatives & similar tools

    Other tools in our directory you may want to compare.

    Also mentioned

    • ElevenLabs (premium voice quality and cloning ecosystem)
    • Play.ht (broad TTS library with commercial licensing)
    • Resemble AI (voice synthesis plus deepfake detection)

    Features & details

    Full overview from our catalog (read-only reference).

    Category (short)
    Text To Speech
    Category
    Text To Speech
    Pricing (CSV)
    Free Trial
    Directory pricing
    Freemium
    Sponsored note
    no
    Target audience
    Agencies & studios
    Best for
    Budget-conscious creators and developers needing competitive API pricing with voice cloning. Not ideal for: teams requiring only a simple consumer TTS app without credit management, or buyers expecting broadcast human voice talent quality on every output.
    Pricing notes
    Verified July 2026 at fish.audio/plan and docs.fish.audio. Web plans: Free $0/mo, 8,000 credits (~7 min/mo), 500 chars/generation, personal use per FAQ. Plus: $15/mo monthly or $5.50/mo annual ($66/yr), 250,000 credits (~200 min), 1 PVC slot, commercial use. Pro: $100/mo monthly or $37.50/mo annual ($450/yr), 2M credits (~1,620 min), 3 team seats, 5 PVC slots. Max: $999/mo monthly or $749/mo annual ($8,988/yr), 25M credits (~6,250 min), 10 seats, 15 PVC slots. API (pay-as-you-go): S2.1 Pro/S2 Pro/S1 TTS $15/M UTF-8 bytes (~180K English words or ~12 hrs speech); ASR transcribe-1 $0.36/audio hour; voice-design-1 $0.01/request. Enterprise: custom from ~$999/mo.

    Pros & cons

    Editorial notes to help compare fit before opening the vendor site.

    Pros

    • Very competitive API pricing ($15/M UTF-8 bytes) Large community voice library (2M+) Strong emotion tag control for expressive narration Generous free tier for testing (no credit card)

    Cons

    • Credit math and monthly reset can frustrate heavy users Pro monthly list price jumped to $100/mo on official plan page Commercial rights require paid tier
    • free tier personal-only per FAQ Learning curve for emotion tags and API concurrency tiers

    Review notes

    Context from the listing review and editorial research.

    Raised $52M seed; 8M+ builders cited on site. API pricing significantly undercuts ElevenLabs per-character rates. Free tier FAQ restricts commercial use despite plan card wording—verify terms before monetizing. Credits reset monthly without rollover.

    Extended features

    In-depth description and capability notes.

    S2.1 Pro TTS with inline emotion tags (angry, whispering, laughing, pauses, etc.) Voice cloning from ~15 seconds; Professional Voice Clone (PVC) on paid plans 2M+ user-uploaded voices in Discovery library; Voice Design API Speech-to-text (transcribe-1) and real-time WebSocket streaming Pay-as-you-go API with Python SDK; open-source models (Apache 2.0) Story Studio for audiobook and long-form narration workflows 30+ languages; enterprise tier with on-prem and zero data retention options

    Use cases

    YouTubers and podcasters generate narration without studio time. Developers embed low-latency TTS and voice agents via API. Game and animation teams clone character voices with emotional control tags.