Aurora Inference
OpenAI-compatible API for open-weight models, with in-region deployment on reserved capacity
Aurora Inference serves open-weight models, including DeepSeek V4, GLM 5.2 and Kimi K3, through a single OpenAI-compatible endpoint. Metered per-token access has no commitment or minimum, and reserved capacity is available for steady workloads.
On reserved capacity, prompts, outputs and weights stay in the chosen region, with deployments across North America, Europe, the Nordics, the Middle East and APAC. Cached input is priced separately, from $0.006/1M tokens on DeepSeek V4 Flash. It is part of Aurora Infra, which also sells GPU clusters and storage.
Pricing: Per token usage
Aurora Inference Alternatives
Explore 104 products in the Inference APIs category. View all Aurora Inference alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
deepinfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
Mistral
Use models in a few clicks with our platform. Download our open models for deep access.
Work on Aurora Inference? Feature it at the top of Inference APIs.
Is your product missing?