Together AI is a cloud platform for running, fine-tuning, and training open-source and proprietary AI models at scale. It offers serverless inference across chat, vision, image, audio, video, transcription, embedding, and reranking models, along with dedicated GPU endpoints, on-demand/reserved GPU clusters (H100, H200, B200, GB200), and managed fine-tuning. It serves AI developers, startups, and enterprises such as Cursor, Zoom, and Salesforce building on open models.
Related guides
Browse curated shortlists where this tool appears.
Fireworks AI (competing open-model inference platform)
Replicate (model hosting and inference marketplace)
AWS Bedrock and Google Vertex AI (hyperscaler-native model hosting)
Features & details
Full overview from our catalog (read-only reference).
Category (short)
3D
Category
3D
Pricing (CSV)
Paid Service
Directory pricing
Paid
Sponsored note
no
Target audience
Developers and infrastructure/ML teams wanting flexible, usage-based access to open-source models and raw GPU compute without building their own inference stack. Not ideal for: teams wanting a single flat subscription price, or those needing only one specific closed-source frontier model rather than a multi-model marketplace.
Best for
Developers and infrastructure/ML teams wanting flexible, usage-based access to open-source models and raw GPU compute without building their own inference stack. Not ideal for: teams wanting a single flat subscription price, or those needing only one specific closed-source frontier model rather than a multi-model marketplace.
Pricing notes
Verified July 2026 at together.ai/pricing. Serverless inference is priced per model, e.g., Llama 3.3 70B at $1.04 per 1M tokens and Qwen3.5 9B at $0.17 per 1M tokens, with cached-input discounts on several models; image generation ranges from $0.0019 to $0.134 per image, and video generation from $0.14 to $3.20 per video depending on model. GPU Clusters: on-demand pricing runs from $3.99/hour (H100) to $8.19/hour (B200), with reserved rates as low as $3.19/hour (H100, 181+ day terms). Fine-tuning starts at $0.48 per 1M tokens (LoRA, models up to 16B) and rises to $100 per 1M tokens for large specialized models like GLM-5.1. The largest GPU tiers and dedicated inference require contacting sales.
Pros & cons
Editorial notes to help compare fit before opening the vendor site.
Pros
Very broad model catalog (text, image, video, audio) with transparent per-unit pricing
Competitive GPU cluster rates with volume discounts for longer reservations
Supports the full workflow from inference to fine-tuning to dedicated training clusters
Used in production by well-known AI companies (Cursor, ElevenLabs, Salesforce)
Cons
Pricing complexity across many models and units makes upfront cost estimation harder
Largest enterprise GPU tiers are quote-only
Usage-based costs can scale unpredictably at high volume without monitoring
Requires technical/API integration rather than being a turnkey end-user product
Review notes
Context from the listing review and editorial research.
Together AI recently announced a Series C funding round. The platform lists major customers and partners including Cursor, Salesforce, Zoom, ElevenLabs, and DeepMind, and has a partnership with Y Combinator for a dedicated GPU cluster.
Extended features
In-depth description and capability notes.
Serverless pay-per-token inference for 30+ open models (Llama, Qwen, DeepSeek, Kimi, GLM, gpt-oss, and more)
Image, video, and audio generation model hosting (FLUX, Veo, Kling, Seedance, Sora 2 access)
Provisioned Throughput reserved-capacity pricing for predictable high-volume workloads
Dedicated single-tenant GPU inference endpoints
On-demand and reserved GPU Clusters (H100, H200, B200, GB200) for training
Supervised fine-tuning and Direct Preference Optimization pipelines
Code Sandbox and Code Interpreter for agent execution
Managed high-bandwidth shared filesystem storage
Use cases
AI startups serve chat or agent products on open-weight models without managing GPU infrastructure themselves, ML teams fine-tune open models on proprietary data, and companies rent dedicated GPU clusters for large training runs.