Long prompts, multi-turn chat and agent loops all share one hidden cost: the key-value (KV) cache. Every token an LLM processes leaves behind attention keys and values that sit in scarce GPU memory. When that memory fills up, your inference server either evicts the cache and recomputes the prefill, or it stops accepting new requests. KV cache offloading fixes this by moving those tensors to cheaper tiers such as CPU RAM, local NVMe or a shared cache cluster, then pulling them back when a request needs them.
In 2026 three options dominate production conversations for vLLM-based stacks: LMCache, NVIDIA Dynamo’s KV Block Manager (KVBM), and FlexKV. This guide explains how KV cache offloading works, compares the three, and helps you pick the right one for your workload.
What Is KV Cache Offloading and Why Does It Matter?

During the prefill phase, a transformer computes keys and values for every input token. The decode phase then reuses them for each new token. That cache grows linearly with context length, batch size and model depth, so a long-context model serving dozens of concurrent users can exhaust HBM long before compute becomes the bottleneck.
Engines like vLLM already manage GPU memory in pages and reuse shared prefixes. Offloading extends that idea across a memory hierarchy:
- GPU HBM (tier 1): fastest, smallest, most expensive.
- CPU pinned RAM (tier 2): often several times larger, reachable over PCIe or NVLink-C2C.
- Local SSD / NVMe (tier 3): large and cheap, useful for long-lived prefixes.
- Remote / shared stores: Redis, Mooncake or similar, so many engines can reuse one cache.
The payoff is simple: loading a cached prefix from CPU or disk is usually much cheaper than recomputing a long prefill. That lowers time-to-first-token (TTFT) for repeated context — system prompts, RAG documents, chat history, agent scratchpads — and lets you pack more concurrent sessions onto each GPU.
Workloads that benefit most
- Multi-turn chatbots where each turn resends the full conversation.
- RAG pipelines that repeatedly query the same large documents.
- Coding agents and tool-calling loops with long, stable system prompts.
- Prefill-decode (PD) disaggregated deployments that must ship KV blocks between GPUs.
If your prompts are short and rarely repeated, offloading adds little. Measure your prefix-hit rate first.
LMCache: The Open-Source KV Cache Layer
LMCache is an open-source project that extracts KV caches from engines such as vLLM and SGLang, stores them outside GPU memory, and shares them across queries and engines. According to the LMCache paper on arXiv, it supports two main modes: cache offloading for prefix reuse, and cross-engine transfer for PD disaggregation. The authors report up to 15x throughput gains on workloads like multi-round QA and document analysis — a best-case figure you should validate on your own traffic.
Key strengths
- Multi-level storage backends: CPU RAM, local disk, Redis, GPUDirect Storage and Mooncake/InfiniStore.
- Engine-agnostic design that plugs into vLLM through its KV connector interface.
- Research-backed extras such as CacheGen (KV compression) and CacheBlend (reusing non-prefix chunks, handy for RAG).
- Integrates with the vLLM Production Stack Helm charts on Kubernetes.
Best for: teams that want a vendor-neutral, open-source cache layer and need sharing across many engine replicas.
NVIDIA Dynamo KVBM: Built for Disaggregated Serving
Dynamo is NVIDIA’s open-source distributed inference framework, and KVBM is its built-in KV block manager. The Dynamo KV cache offloading docs describe a three-layer design: the LLM runtime, a logical block-management layer, and NIXL as the transport. Blocks move between a GPU device pool, a pinned CPU host pool and disk through asynchronous, per-path transfer queues.
Key strengths
- Native integration with Dynamo’s KV-aware router, which sends requests to the worker that already holds the matching prefix.
- First-class support for prefill-decode disaggregation across nodes.
- NIXL transport tuned for NVIDIA networking and storage paths.
- Works with vLLM and other Dynamo-supported backends.
Best for: large NVIDIA GPU fleets already adopting Dynamo, where KV-aware routing and disaggregation matter as much as raw cache capacity.
FlexKV: Distributed Multi-Tier Reuse
FlexKV is listed in the Dynamo documentation as another offloading backend for vLLM. It focuses on multi-level caching (GPU, CPU, SSD), distributed KV cache reuse across nodes, and high-throughput I/O through io_uring and GPUDirect Storage. It is a strong fit when SSD capacity is your main lever and you want reuse to span a cluster rather than a single host.
Best for: SSD-heavy deployments with very long or very numerous reusable prefixes.
LMCache vs KVBM vs FlexKV: Side-by-Side
| Feature | LMCache | Dynamo KVBM | FlexKV |
|---|---|---|---|
| Tiers | CPU, disk, Redis, GDS, Mooncake | GPU, CPU pinned, disk | GPU, CPU, SSD |
| Cross-node reuse | Yes (remote backends) | Via Dynamo routing + NIXL | Yes |
| PD disaggregation | Supported | Core design goal | Not the main focus |
| Engines | vLLM, SGLang | vLLM and other Dynamo backends | vLLM |
| Ecosystem lock-in | Low | Tied to Dynamo | Low–moderate |
| Standout feature | CacheBlend for RAG | KV-aware routing | io_uring + GDS I/O |
vLLM also ships native CPU offloading, which is a sensible baseline before adding any external system.
How to Choose (and Roll Out Safely)
- Start with metrics. Track GPU KV utilisation, preemption counts, prefix-cache hit rate and TTFT p95 before changing anything.
- Try native offloading first. If CPU RAM alone fixes preemption, you may not need more.
- Pick LMCache for open-source flexibility, cross-engine sharing or RAG-heavy traffic.
- Pick KVBM if you are standardising on Dynamo and need disaggregated serving with smart routing.
- Pick FlexKV when large SSD tiers and cluster-wide reuse are the priority.
- Benchmark with real prompts. Synthetic tests with no shared prefixes will hide the benefit — or exaggerate it.
For Java and Spring teams calling these servers, nothing changes on the client: offloading lives entirely in the serving layer, behind the same OpenAI-compatible API. Pair it with our guides on vLLM vs SGLang vs TensorRT-LLM, speculative decoding and LLM semantic caching to cut latency end to end.

Frequently Asked Questions
Is KV cache offloading the same as prefix caching?
No. Prefix caching reuses KV blocks that are still in GPU memory. KV cache offloading moves those blocks to CPU, disk or remote storage so they survive eviction and can be reused later or by other servers.
Does offloading slow down token generation?
Decode speed is largely unaffected because active blocks are loaded back to the GPU first. The main trade-off is transfer time, which is usually far smaller than recomputing a long prefill.
Can I use LMCache with NVIDIA Dynamo?
Yes. Dynamo’s documentation lists LMCache as one of its supported offloading backends for vLLM, alongside KVBM and FlexKV.
How much CPU RAM do I need?
It depends on model size, context length and concurrency. A practical start is sizing the CPU tier at a few times your GPU KV capacity, then tuning from hit-rate metrics.
Conclusion
KV cache offloading is one of the cheapest ways to raise throughput and cut TTFT for long-context and multi-turn LLM workloads. LMCache offers open, flexible, multi-backend caching; Dynamo KVBM shines in disaggregated NVIDIA fleets; and FlexKV targets SSD-scale reuse. Measure your prefix reuse, start with native offloading, then layer on the tool that matches your architecture.
Running vLLM in production? Bookmark NewsifyAll and subscribe for weekly, hands-on guides to faster and cheaper LLM inference.

