Shared KV-Cache Poisoning: Isolating Paged Attention Between Tenants

Why a single vLLM engine serving multiple agents needs more than a request queue.

by

Paged attention is the reason a single inference engine can serve more than one agent without pre-allocating a contiguous KV buffer per request. Instead of one long tensor per sequence, the cache is split into fixed-size blocks, and each sequence carries a block table mapping logical positions to physical slots. The engine's allocator hands out blocks, the attention kernel gathers them, and when a sequence finishes the blocks go back to the pool. It is virtual memory for attention.

The convenience hides a sharp edge. The block table is shared state. The prefix cache is shared state. The eviction policy is shared state. Any path where one tenant's request can influence another tenant's block table, prefix hash, or eviction order is a path where context can be read, overwritten, or silently corrupted. In a multi-agent deployment, that is not a theoretical concern. Agents routinely share an engine, and their prompts are not equally trusted.

What paged attention actually shares

A vLLM-style engine keeps three structures that matter for isolation:

  1. The physical block pool. A flat array of KV blocks, indexed by block number.
  2. The block table per sequence. Logical token position to physical block.
  3. The prefix cache. A hash map from token-prefix hash to a reference-counted block.

The prefix cache is the interesting one. It exists so that two requests with the same system prompt do not recompute the same KV. That is a pure win when both requests belong to the same tenant. When they belong to different tenants, the cache becomes a side channel. A request that shares a prefix with another tenant's request can, depending on the engine's matching rules, attach to blocks that were computed under a different trust boundary.

# Simplified block table, one per sequence.
# The allocator is global; the table is per-sequence.
class BlockTable:
    def __init__(self, block_size: int):
        self.block_size = block_size
        self.blocks: list[int] = []

    def append(self, physical_block: int) -> None:
        self.blocks.append(physical_block)

    def slot(self, logical_pos: int) -> tuple[int, int]:
        block_idx = logical_pos // self.block_size
        offset = logical_pos % self.block_size
        return self.blocks[block_idx], offset

The table itself is cheap. The bug is never in the table arithmetic. It is in who is allowed to reuse a block and under what hash.

Three concrete poisoning shapes

1. Prefix-cache collision across tenants

The prefix cache is keyed by a hash of the token prefix, usually with the model, adapter, and some sampling parameters folded in. If the key does not include a tenant identifier, two tenants with identical prefixes share blocks. Identical prefixes are not rare. System prompts, tool schemas, and few-shot exemplars are often copy-pasted across teams. The result is that tenant B's request can read KV computed for tenant A's request. Even if the tokens are identical, the attention state was produced under A's adapter, A's quantization settings, or A's LoRA weights.

2. Block reuse without a generation counter

A block is freed when a sequence finishes. If the allocator hands that block to a new sequence before the old sequence's writes are fully fenced, the new sequence can observe stale KV. On a single GPU this is usually masked by kernel ordering. On a multi-GPU tensor-parallel setup, the fence is across ranks, and the window is real. A common pattern in bug reports is a rare, non-reproducible token corruption that only appears under high concurrency.

3. Eviction as a cross-tenant signal

Eviction is not just a performance policy. The order in which blocks are evicted leaks information about other tenants' activity. A tenant that can observe its own eviction rate can infer load from other tenants. In a shared engine, that is a side channel. It is not a memory-safety bug, but it is a confidentiality bug, and it is the one most teams notice last.

Isolation strategies that actually hold

Partition the block pool

The bluntest fix is to give each tenant its own block pool. The allocator never crosses tenants, so no block table can reference another tenant's physical memory. The cost is fragmentation. If tenant A is idle and tenant B is saturated, B cannot use A's blocks. On a single GPU this is often unacceptable, because the whole point of paging is to oversubscribe.

A middle path is a pool per trust class. Group tenants that are allowed to share (same team, same adapter, same data classification) into one pool, and keep pools separate across classes. This is the model most production deployments converge on.

# Allocator scoped to a trust class, not a single tenant.
class ScopedAllocator:
    def __init__(self, num_blocks: int, trust_class: str):
        self.free = list(range(num_blocks))
        self.trust_class = trust_class

    def alloc(self, n: int) -> list[int]:
        if len(self.free) < n:
            raise MemoryError(f"pool exhausted for {self.trust_class}")
        return [self.free.pop() for _ in range(n)]

    def free_blocks(self, blocks: list[int]) -> None:
        self.free.extend(blocks)

Make the prefix key tenant-aware

If you keep a shared prefix cache, the key must include everything that affects the KV: model id, adapter id, quantization config, and a tenant or trust-class id. The hash is cheap. The bug is expensive. A key that omits the tenant is a key that assumes all tenants are equivalent, which is exactly the assumption you cannot make.

def prefix_key(tokens: list[int], model_id: str, adapter_id: str,
               trust_class: str) -> str:
    h = hashlib.sha256()
    h.update(model_id.encode())
    h.update(adapter_id.encode())
    h.update(trust_class.encode())
    for t in tokens:
        h.update(t.to_bytes(4, "little"))
    return h.hexdigest()

The tradeoff is cache hit rate. Tenant-aware keys fragment the cache, and a fragmented cache is a slower cache. Teams typically accept this because the alternative is a cross-tenant read.

Reference-count and fence blocks

A block should not return to the free list until every rank that wrote to it has acknowledged the write. In a tensor-parallel deployment, that means a barrier or a per-rank completion counter before the block is reusable. This is not free. It adds synchronization to the hot path. The alternative is stale KV, which is worse.

class RefCountedBlock:
    def __init__(self, block_id: int, num_ranks: int):
        self.block_id = block_id
        self.refs = 0
        self.pending_writes = num_ranks

    def on_write_complete(self) -> None:
        self.pending_writes -= 1

    def releasable(self) -> bool:
        return self.refs == 0 and self.pending_writes == 0

Separate engines per trust class

The strongest isolation is no sharing. Run one engine per trust class, each with its own process, its own GPU memory, and its own scheduler. This is the architecture to reach for when the tenants are genuinely adversarial. It costs GPU memory and operational complexity, and it removes an entire class of bug. In a RiNET-style stack, this maps cleanly onto separate vLLM processes behind a routing layer, with the WireGuard mesh handling transport and Postgres holding the tenant-to-engine mapping.

Where the LoRA layer changes the calculus

A nightly LoRA pipeline produces adapters that are loaded into the same engine. Two tenants can share a base model and diverge only in adapter weights. If the prefix cache key does not include the adapter id, a request under adapter A can attach to KV computed under adapter B. The tokens match. The attention state does not. This is the most common real-world version of the bug, because adapter sharing is exactly what makes multi-tenant fine-tuning economical.

The fix is the same as above: fold the adapter id into the prefix key and into the block pool scope. If adapter A and adapter B are both loaded, they should not share blocks unless you have explicitly reasoned about why that is safe. In most cases it is not.

A checklist for shared-engine deployments

  • Does the prefix cache key include model, adapter, quantization, and tenant or trust class?
  • Is the block pool scoped per trust class, or is it global?
  • Are freed blocks fenced across all tensor-parallel ranks before reuse?
  • Does eviction expose cross-tenant load information, and does that matter for your threat model?
  • Is there a per-tenant quota on blocks, so one tenant cannot starve another by exhausting the pool?
  • Are adapter switches reflected in the cache key, or only in the model weights?

None of these are exotic. They are the same questions you would ask about any shared-memory system. Paged attention is shared memory with a scheduler on top, and it deserves the same scrutiny.

The architecture decision

The choice is not whether to share. It is what to share. Sharing the GPU is fine. Sharing the block pool is fine within a trust class. Sharing the prefix cache across trust classes is where the bug lives. The cheapest correct design is usually a small number of trust-class-scoped engines, each with a tenant-aware prefix key and a fenced allocator. That is more moving parts than a single global engine, and it is the version you can defend in a review.

The alternative, a single engine with a global cache and a hope that prefixes never collide, works until two agents happen to share a system prompt. At that point the failure is silent, the output is plausible, and the debugging is miserable. Isolation is cheaper than forensics.

#inference#kv-cache#memory-isolation#multi-agent#paged-attention#vllm
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related