Icon for SovInfra

SovInfra

Free Trial

OpenAI-compatible inference API for open-weight models, run on EU GPUs with no request retention.

SovInfra serves Gemma 4 31B and Qwen 3.8 27B on H200 GPUs in the Czech Republic through an OpenAI-compatible API. Call a model by name, or set the model to "sovapi" to route between the two. Whisper Large v3 transcription, BGE-M3 embeddings and Kokoro text-to-speech run on the same key.

Requests are not logged or used for training. The sovereignty page names each subprocessor; a DPA is available. The operator's own ISO 27001 is still in preparation. The first billion tokens are free.

Pricing: Per token usage

Hosting Cloud
Pricing Usage Based, $0.12 / 1M input tokens
HQ 🇫🇷 France
Compliance GDPR
Screenshot of SovInfra webpage

SovInfra prices by model

Per 1M tokens, read off SovInfra's own pricing page on the date shown.

Model Input / 1M Cached / 1M Output / 1M Checked Notes
Qwen3.8 27B $0.12 $0.04 $0.38 2 Oct 2026

Work on SovInfra? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →