AI Inference & Training Cloud Infrastructure
    Freemium
    Free Trial
    Sponsored

    Together AI

    AI Inference & Training Cloud Infrastructure

    Together AI is a cloud platform for running, fine-tuning, and training open-source and proprietary AI models at scale. It offers serverless inference across chat, vision, image, audio, video, transcription, embedding, and reranking models, along with dedicated GPU endpoints, on-demand/reserved GPU clusters (H100, H200, B200, GB200), and managed fine-tuning. It serves AI developers, startups, and enterprises such as Cursor, Zoom, and Salesforce building on open models.

    Alternatives & similar tools

    Related picks from editorial notes.

    Also mentioned

    • Fireworks AI (competing open-model inference platform)
    • Replicate (model hosting and inference marketplace)
    • AWS Bedrock and Google Vertex AI (hyperscaler-native model hosting)

    Features & details

    Full overview from our catalog (read-only reference).

    Category (short)
    AI Inference & Training Cloud Infrastructure
    Category
    AI Inference & Training Cloud Infrastructure
    Pricing (CSV)
    Free Trial
    Directory pricing
    Freemium
    Sponsored note
    no
    Target audience
    Developers and infrastructure/ML teams wanting flexible, usage-based access to open-source models and raw GPU compute without building their own inference stack. Not ideal for: teams wanting a single flat subscription price, or those needing only one specific closed-source frontier model rather than a multi-model marketplace.
    Best for
    Developers and infrastructure/ML teams wanting flexible, usage-based access to open-source models and raw GPU compute without building their own inference stack. Not ideal for: teams wanting a single flat subscription price, or those needing only one specific closed-source frontier model rather than a multi-model marketplace.
    Pricing notes
    Verified July 2026 at together.ai/pricing. Serverless inference is priced per model, e.g., Llama 3.3 70B at $1.04 per 1M tokens and Qwen3.5 9B at $0.17 per 1M tokens, with cached-input discounts on several models; image generation ranges from $0.0019 to $0.134 per image, and video generation from $0.14 to $3.20 per video depending on model. GPU Clusters: on-demand pricing runs from $3.99/hour (H100) to $8.19/hour (B200), with reserved rates as low as $3.19/hour (H100, 181+ day terms). Fine-tuning starts at $0.48 per 1M tokens (LoRA, models up to 16B) and rises to $100 per 1M tokens for large specialized models like GLM-5.1. The largest GPU tiers and dedicated inference require contacting sales.

    Pros & cons

    Editorial notes to help compare fit before opening the vendor site.

    Pros

    • Very broad model catalog (text, image, video, audio) with transparent per-unit pricing
    • Competitive GPU cluster rates with volume discounts for longer reservations
    • Supports the full workflow from inference to fine-tuning to dedicated training clusters
    • Used in production by well-known AI companies (Cursor, ElevenLabs, Salesforce)

    Cons

    • Pricing complexity across many models and units makes upfront cost estimation harder
    • Largest enterprise GPU tiers are quote-only
    • Usage-based costs can scale unpredictably at high volume without monitoring
    • Requires technical/API integration rather than being a turnkey end-user product

    Review notes

    Context from the listing review and editorial research.

    Together AI recently announced a Series C funding round. The platform lists major customers and partners including Cursor, Salesforce, Zoom, ElevenLabs, and DeepMind, and has a partnership with Y Combinator for a dedicated GPU cluster.

    Extended features

    In-depth description and capability notes.

    Serverless pay-per-token inference for 30+ open models (Llama, Qwen, DeepSeek, Kimi, GLM, gpt-oss, and more) Image, video, and audio generation model hosting (FLUX, Veo, Kling, Seedance, Sora 2 access) Provisioned Throughput reserved-capacity pricing for predictable high-volume workloads Dedicated single-tenant GPU inference endpoints On-demand and reserved GPU Clusters (H100, H200, B200, GB200) for training Supervised fine-tuning and Direct Preference Optimization pipelines Code Sandbox and Code Interpreter for agent execution Managed high-bandwidth shared filesystem storage

    Use cases

    AI startups serve chat or agent products on open-weight models without managing GPU infrastructure themselves, ML teams fine-tune open models on proprietary data, and companies rent dedicated GPU clusters for large training runs.