≫ Home / LLM API pricing / gpt-oss-20b

gpt-oss-20b API pricing

Of the 6 providers checked, CoreWeave lists the lowest blended rate for gpt-oss-20b: $0.03 input and $0.13 output per 1M tokens, on its pricing page as of 27 September 2026.

The smaller of OpenAI's two open-weight models, meant for lower latency and for running locally. It has the same Apache 2.0 license as gpt-oss-120b.

Creator
OpenAI
Released
5 August 2025
Weights
Open
License
Apache 2.0
Size
21B total, 4B active
Context
128K tokens
Input
text
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 6 providers, the most expensive charges 4.1ร— the cheapest on a blended basis.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 CoreWeave $0.03 โ€“ $0.13 27 Sep 2026
2 DeepInfra $0.03 โ€“ $0.14 25 Sep 2026
3 Novita AI $0.04 โ€“ $0.15 25 Sep 2026
4 SiliconFlow $0.04 โ€“ $0.18 25 Sep 2026
5 Groq $0.075 โ€“ $0.30 25 Sep 2026
6 Cloudflare Workers AI $0.20 โ€“ $0.30 27 Sep 2026 @cf/openai/gpt-oss-20b

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →