Icon for KV Cache Store

KV Cache Store

Open Source Free Trial

Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill

KV Cache Store lets you precompute the KV cache for a long prompt once, then load it before inference instead of paying the prefill cost on every request. The open-source kvcdn CLI builds, verifies, quantizes and benchmarks cache artifacts locally.

An optional hosted registry stores artifacts behind stable URLs with SHA-256 digests and model, dtype and tokenizer metadata, so a cache can be shared across machines without ambiguity about what produced it.

Useful when many requests share a large fixed context, such as a long system prompt or a document being queried repeatedly. Currently in public beta.

Pricing: Monthly subscriptions

Screenshot of KV Cache Store webpage

Work on KV Cache Store? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →