Icon for together.ai

together.ai

The fastest cloud platform for building and running generative AI.

Together.ai Inference provides fast, scalable, and cost-efficient serverless API endpoints for deploying and fine-tuning leading open-source models like Llama-2 and Mistral. It emphasizes speed and efficiency, claiming up to 3x faster performance and 6x lower costs than competitors, alongside automatic scaling to meet growing API request volumes. The platform supports over 100 models.

Pricing: Per token usage

Hosting Cloud + Self-hosted
Pricing Usage Based, Pay-per-token
HQ 🇺🇸 United States
Founded 2022
License PROPRIETARY
Compliance SOC 2 · HIPAA · GDPR · SSO
Screenshot of together.ai webpage

Together AI runs open-weight models as a hosted service behind an OpenAI-compatible API, alongside fine-tuning, dedicated endpoints, batch jobs, an evaluations API, a code-execution sandbox and custom model hosting. For teams that want the same models on reserved hardware, it also rents GPU clusters directly.

Serverless pricing is per model and per million tokens. Read on 16 August 2026 from Together's own pricing page and cross-checked against its serverless model docs: Llama 3.3 70B Instruct Turbo at $1.04 in and $1.04 out, GLM-5.2 at $1.40 and $4.40 with a 512K context window, and DeepSeek-V4-Flash at $0.14 and $0.28 with a 1M context window. Dedicated inference is listed at $5.49 per H100 GPU-hour and $8.99 for B200; GPU clusters start at $3.99 per H100-hour on demand and $3.19 reserved for 181 days or more. The model catalogue changes quickly, so treat any of these figures as needing a re-check rather than as settled.

The company was founded in 2022, is led by Vipul Ved Prakash, holds ISO 27001:2022 certification, and announced an $800 million Series C on 1 July 2026 with investors including NVIDIA, Aramco Ventures, Vista Equity and General Catalyst.

Also listed in

Work on together.ai? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →