Paging the KV Cache to NVMe: The Per-Token Tax
When context exceeds VRAM, paging the KV cache to NVMe adds a measurable per-token latency tax. Here is how to think about it.
When your context window outgrows VRAM, something has to give. The KV cache, which holds the keys and values for every token in the sequence, grows linearly with context length. On a 24GB card running a 7B model in FP16, you can fit roughly 16k tokens before the cache alone consumes half your memory. Push to 32k or 128k and the math stops working. Teams typically respond by paging the KV cache to NVMe. It works, but it is not free. Every token generated from a paged cache pays a latency tax.
This article breaks down that tax: where it comes from, how to estimate it, and what you can do to reduce it.
The anatomy of the tax
The per-token latency tax from NVMe paging has three components:
- PCIe transfer time – moving the KV blocks between host memory and the GPU.
- NVMe read latency – the time to fetch blocks from the SSD.
- Software overhead – the paging logic inside the inference engine.
Each component scales differently with block size and access pattern. Understanding them is the first step to mitigation.
PCIe transfer time
A PCIe 4.0 x16 slot delivers about 25 GB/s in each direction. A KV block for a 7B model in FP16 with 32 layers, 32 heads, and head dimension 128 is roughly 32 * 2 * 32 * 128 * 2 bytes = 512 KB per token. At 25 GB/s, transferring one token's KV block takes about 20 microseconds. For a 4k-token context, that is 80 milliseconds just to move the data if you had to transfer it all at once. In practice, you only transfer the blocks needed for the current token, but the cache is accessed per layer, per head, so the effective transfer size is smaller and more frequent.
NVMe read latency
Enterprise NVMe drives have read latencies in the 80–120 microsecond range for 4 KB random reads. Larger reads amortize better. A 512 KB read might take 150–200 microseconds. If your paging strategy fetches one token's KV block per layer, you are looking at 32 separate reads, each with its own latency. That is 32 * 100 µs = 3.2 milliseconds per token just in NVMe latency. For a 50 token/s generation rate, that is a 16% overhead. At 10 token/s, it is 3.2%.
Software overhead
The inference engine must decide which blocks to evict, which to prefetch, and how to schedule the transfers. vLLM's paged attention already manages KV blocks in GPU memory. Extending that to NVMe requires a block manager that tracks residency across tiers. The overhead of that bookkeeping is typically in the tens of microseconds per token, but it can spike if the eviction policy thrashes.
Estimating the tax for your workload
The total per-token tax is the sum of the three components. A simple model:
def per_token_tax(context_len, block_size_kb, pcie_bw_gbps, nvme_latency_us, layers):
# PCIe transfer: assume we transfer one block per layer per token
pcie_time_us = (block_size_kb * 1024 * 8) / (pcie_bw_gbps * 1e9) * 1e6
# NVMe read: one read per layer
nvme_time_us = nvme_latency_us * layers
# Software overhead: assume 20 us per token
sw_time_us = 20
return pcie_time_us + nvme_time_us + sw_time_us
# Example: 7B model, 32 layers, 512 KB blocks, PCIe 4.0, NVMe 100 us
print(per_token_tax(4096, 512, 25, 100, 32)) # ~3.3 msThis model ignores batching and prefetching, which can hide some of the latency. But it gives a lower bound. If your generation loop is 20 ms per token (50 token/s), adding 3.3 ms is a 16% slowdown. If you are already at 100 ms per token, the tax is 3%.
When paging makes sense
Paging to NVMe is not always the wrong choice. It makes sense when:
- Context length is variable and bursty. Most requests are short, but a few need 128k. Paging lets you serve the long tail without provisioning for it.
- You are memory-bound, not latency-bound. If your SLA allows 200 ms per token, the tax is acceptable.
- You can batch. Batching amortizes the per-token overhead across multiple sequences. If you process 8 sequences at once, the NVMe reads can be coalesced.
It makes less sense when:
- You need sub-50 ms per token. The tax will dominate.
- Your context is uniformly long. If every request needs 128k, you are better off with a larger GPU or a smaller model.
- Your NVMe is shared. Contention with other workloads can spike latency unpredictably.
Mitigation strategies
If you must page, here is how to reduce the tax.
1. Prefetch aggressively
The inference engine knows the next token's KV blocks are likely to be needed soon. Prefetching them into GPU memory while the current token is being computed can hide NVMe latency. vLLM's block manager can be extended to issue prefetch requests for blocks that are likely to be accessed based on the attention pattern.
2. Use larger blocks
Small blocks mean more NVMe reads. If you page in 2 MB blocks instead of 512 KB, you halve the number of reads. The tradeoff is that you may fetch data you do not need. But for sequential access patterns, larger blocks are almost always better.
3. Pin hot blocks in VRAM
Not all KV blocks are equally likely to be accessed. The most recent tokens are accessed every step. The first few tokens (the system prompt) are accessed every step. Middle tokens are accessed less frequently. A tiered cache that keeps hot blocks in VRAM, warm blocks in host DRAM, and cold blocks on NVMe can reduce the tax significantly.
# Pseudocode for a tiered KV cache
class TieredKVCache:
def __init__(self, vram_capacity, dram_capacity):
self.vram = LRUCache(vram_capacity)
self.dram = LRUCache(dram_capacity)
self.nvme = NVMeStore()
def get(self, block_id):
if block_id in self.vram:
return self.vram[block_id]
if block_id in self.dram:
block = self.dram[block_id]
self.vram.put(block_id, block)
return block
block = self.nvme.read(block_id)
self.dram.put(block_id, block)
self.vram.put(block_id, block)
return block4. Compress the KV cache
Quantizing the KV cache to INT8 or INT4 reduces its size by 2–4x. That means fewer blocks to page, and less PCIe traffic. The tradeoff is a small accuracy drop, but for many workloads it is acceptable. Recent work on KV cache quantization shows that INT8 is nearly lossless for most models.
5. Use a faster interconnect
PCIe 5.0 doubles bandwidth to 50 GB/s. If your GPU and CPU support it, the PCIe transfer time halves. NVMe drives with PCIe 5.0 also have lower latency. This is a hardware upgrade, but it directly reduces the tax.
The real-world tradeoff
In practice, teams that page the KV cache to NVMe typically see a per-token latency increase in the 10–30% range, depending on context length, block size, and batching. The exact number is workload-specific. The key is to measure it in your own setup. Do not assume the tax is negligible, and do not assume it is prohibitive.
A common pattern is to start with paging, measure the latency, and then decide whether to invest in mitigation or in more VRAM. If your workload is latency-sensitive and your context is consistently long, more VRAM is the simpler answer. If your context is variable and your latency budget is loose, paging is a cost-effective way to handle the long tail.
What to watch
The KV cache is the new bottleneck in LLM inference. As context windows grow to 1M tokens and beyond, paging to NVMe will become more common, not less. The tooling is still maturing. vLLM's paged attention is a start, but it does not yet handle multi-tier paging natively. Expect to see more engines add NVMe-backed KV caches in the coming year.
For now, if you are paging, measure the tax. It is the only way to know if it is worth it.