KV Cache Store
Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill
KV Cache Store lets you precompute the KV cache for a long prompt once, then load it before inference instead of paying the prefill cost on every request. The open-source kvcdn CLI builds, verifies, quantizes and benchmarks cache artifacts locally.
An optional hosted registry stores artifacts behind stable URLs with SHA-256 digests and model, dtype and tokenizer metadata, so a cache can be shared across machines without ambiguity about what produced it.
Useful when many requests share a large fixed context, such as a long system prompt or a document being queried repeatedly. Currently in public beta.
Pricing: Monthly subscriptions
KV Cache Store Alternatives
Explore 82 products in the Inference APIs category. View all KV Cache Store alternatives.
Scaleway
European serverless AI inference APIs, 100% hosted in Europe
Infercom
European sovereign AI inference with OpenAI-compatible APIs hosted in EU datacenters
Nebius
Full-stack AI cloud with GPU infrastructure for training and inference
IONOS AI Model Hub
OpenAI-compatible API for open-weight LLMs and image models, hosted in IONOS EU data centers
Work on KV Cache Store? Feature it at the top of Inference APIs.
Is your product missing?