Icon for DigitalOcean Inference Engine

DigitalOcean Inference Engine

DigitalOcean's managed inference: serverless LLM and image APIs, dedicated GPU endpoints and batch jobs

DigitalOcean Inference Engine is DigitalOcean's managed inference service. It offers serverless per-token APIs, dedicated GPU endpoints billed by the hour, batch jobs and a router that picks a model per request. Serverless covers OpenAI and Anthropic models plus open-weight models such as gpt-oss. Image models served through fal draw from the same prepaid balance.

Open models start at $0.05 per 1M input tokens (gpt-oss-20b). FLUX Schnell images cost $0.003 per megapixel. Dedicated endpoints start at $2.59 an hour on an AMD MI300X.

Pricing: Per token usage

Hosting Cloud
Pricing $0.05/1M input tokens (gpt-oss-20b)
HQ ๐Ÿ‡บ๐Ÿ‡ธ United States
Screenshot of DigitalOcean Inference Engine webpage

DigitalOcean Inference Engine prices by model

Per 1M tokens, read off DigitalOcean Inference Engine's own pricing page on the date shown.

Model Input / 1M Cached / 1M Output / 1M Checked Notes
gpt-oss-120b $0.10 โ€“ $0.70 11 Oct 2026
gpt-oss-20b $0.05 โ€“ $0.45 11 Oct 2026

Work on DigitalOcean Inference Engine? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →