≫ Home / LLM API pricing / Llama 3.3 70B Instruct

Llama 3.3 70B Instruct API pricing

Of the 6 providers checked, Novita AI lists the lowest blended rate for Llama 3.3 70B Instruct: $0.135 input and $0.40 output per 1M tokens, on its pricing page as of 25 September 2026.

Meta's text-only 70B model from December 2024, with a 128K context. It is older than the rest of this list but still widely hosted, which makes it a useful yardstick for what providers charge. Meta's Llama 3.3 Community License comes with its own use policy.

Creator
Meta
Released
6 December 2024
Weights
Open
License
Llama 3.3 Community License
Size
70B total
Context
128K tokens
Input
text
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 6 providers, the most expensive charges 5.2ร— the cheapest on a blended basis.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 Novita AI $0.135 โ€“ $0.40 25 Sep 2026
2 DeepInfra $0.23 โ€“ $0.40 25 Sep 2026 The Turbo variant is $0.10 / $0.32
3 IONOS AI Model Hub โ‚ฌ0.65 โ‰ˆ $0.7389 โ€“ โ‚ฌ0.65 โ‰ˆ $0.7389 25 Sep 2026 Price on ionos.de; the US site lists $0.71 / $0.71
4 Cloudflare Workers AI $0.293 โ€“ $2.253 25 Sep 2026 @cf/meta/llama-3.3-70b-instruct-fp8-fast
5 Scaleway โ‚ฌ0.90 โ‰ˆ $1.023 โ€“ โ‚ฌ0.90 โ‰ˆ $1.023 25 Sep 2026
6 Together AI $1.04 โ€“ $1.04 25 Sep 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →