SovInfra
OpenAI-compatible inference API for open-weight models, run on EU GPUs with no request retention.
SovInfra serves Gemma 4 31B and Qwen 3.8 27B on H200 GPUs in the Czech Republic through an OpenAI-compatible API. Call a model by name, or set the model to "sovapi" to route between the two. Whisper Large v3 transcription, BGE-M3 embeddings and Kokoro text-to-speech run on the same key.
Requests are not logged or used for training. The sovereignty page names each subprocessor; a DPA is available. The operator's own ISO 27001 is still in preparation. The first billion tokens are free.
Pricing: Per token usage
SovInfra prices by model
Per 1M tokens, read off SovInfra's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| Qwen3.8 27B | $0.12 | $0.04 | $0.38 | 2 Oct 2026 |
SovInfra Alternatives
Explore 116 products in the Inference APIs category. View all SovInfra alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
DeepInfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
Mistral
Use models in a few clicks with our platform. Download our open models for deep access.
Work on SovInfra? Feature it at the top of Inference APIs.
Is your product missing?