Server room with racks of GPU hardware for AI inference

Cheapest AI Inference Providers (September 2026)

/ Updated / Arvid Andersson

The cheapest inference provider depends on the model you run, and the numbers move month to month. This page pins prices to one model, gpt-oss-120B, a current open-source model that most serverless providers host, so the rates are directly comparable. Every figure below is taken from the provider's own pricing page on the date noted, not estimated.

Price comparison: gpt-oss-120B

Per 1M tokens, sorted cheapest first by blended cost (3:1 input to output). Click a provider for its full profile.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 $0.03 – $0.17 27 Sep 2026
2 $0.037 – $0.17 25 Sep 2026
3 €0.04 β‰ˆ $0.0455 – €0.20 β‰ˆ $0.2273 1 Oct 2026
4 $0.05 – $0.25 25 Sep 2026
5 $0.05 – $0.45 25 Sep 2026
6 $0.10 – $0.50 25 Sep 2026
7 $0.15 – $0.60 25 Sep 2026
8 $0.15 – $0.60 25 Sep 2026
9 €0.15 β‰ˆ $0.1705 – €0.60 β‰ˆ $0.682 25 Sep 2026
10 €0.15 β‰ˆ $0.1705 – €0.65 β‰ˆ $0.7389 25 Sep 2026 Price on ionos.de; the US site lists $0.17 / $0.71
11 €0.16 β‰ˆ $0.1819 – €0.63 β‰ˆ $0.7161 1 Oct 2026
12 €0.16 β‰ˆ $0.1819 – €0.83 β‰ˆ $0.9435 1 Oct 2026 Listed in SimpleCredits (100 SC = EUR 1); homepage advertises 20% off during the open beta
13 $0.35 – $0.75 25 Sep 2026
14 $0.35 – $0.75 27 Sep 2026 @cf/openai/gpt-oss-120b
15 €1.00 β‰ˆ $1.1367 – €4.20 β‰ˆ $4.7741 1 Oct 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.

For every provider in the category, see the filterable inference APIs table. The same prices with the model's own facts are on gpt-oss-120b API pricing. For DeepSeek, GLM, Kimi, Qwen and Llama models, see LLM API pricing by model.

How to read this

On gpt-oss-120B, CoreWeave lists the lowest blended rate in the table above, with DeepInfra and Melious AI next. Groq and Together.ai sit at the same published rate on this model. They differ on other axes: Groq on first-token latency from its LPU hardware, Together on fine-tuning and production features on the same platform.

Output tokens cost more than input tokens at every provider here, so for generation-heavy workloads the output column matters most. For prompt-heavy workloads (long context, short answers) the input column carries more weight. The table keeps the two separate so the blend reflects your own traffic mix rather than a fixed assumption.

Cheapest is not always best

Price is a starting filter, not the whole decision. A provider that wins on per-token cost may not host the specific model you need, may have lower rate limits, or may be slower for interactive use where time-to-first-token matters more than raw throughput. Data residency is another constraint: if a workload is GDPR-sensitive, EU-hosted options (Nebius, Scaleway, Mistral La Plateforme) matter more than a few cents per million tokens. The European providers page lists those with hosting regions.

For a fuller picture across speed, model catalog, and API compatibility, see AI Inference API Providers Compared.

Why these numbers move

Per-token prices shift as providers compete on hardware utilization and model demand. A clear example: gpt-oss-120B on DeepInfra was listed around $0.08 per 1M blended in March 2026 and had dropped to roughly $0.05 blended by June 2026. That is the norm, not the exception. Treat any pricing table, including this one, as a snapshot, and check the provider's own page for the model you actually plan to run.

Frequently asked questions

What is the cheapest AI inference provider?

On gpt-oss-120b, the open model the most providers here publish a price for, the cheapest per-token rates on the providers' own pricing pages are CoreWeave at $0.03 input and $0.17 output; DeepInfra at $0.037 input and $0.17 output; Melious AI at €0.04 input and €0.20 output (per 1M tokens, checked 25 September 2026). 15 providers are compared in the table on this page. Per-token pricing shifts month to month and varies a lot by model, so check the provider's own pricing page before committing to high-volume workloads.

Why do inference prices change so often?

Providers compete on price as hardware utilization and model demand shift, so per-token rates move month to month. As a concrete example, gpt-oss-120B on DeepInfra was listed around $0.08 per 1M blended in March 2026 and dropped to roughly $0.05 blended by June 2026. Always check the provider's own pricing page for the specific model before relying on a number.

Is the cheapest inference provider the best option?

Not necessarily. Per-token price is one factor; throughput, time-to-first-token, model catalog, rate limits, and data residency all matter. A provider that is cheapest on one model may not host the model you need, or may be slower for interactive workloads. Use price as a starting filter, then weigh speed and model availability for your specific use case.

How is the blended price calculated?

Blended price assumes a fixed ratio of input to output tokens (commonly 3:1) to combine the two rates into a single number for comparison. Because output tokens are usually priced higher than input tokens, the output rate dominates real-world cost on generation-heavy workloads. The table here shows input and output separately so you can compute the blend for your own traffic mix.

Is your product missing?

Add it here →