mmap vs read/write for vector indexes: pgvector, mlock and the swap storm

Why memory-mapped vector indexes look fast until the kernel decides to reclaim them

by

Most engineers reach for mmap when they need to serve a large vector index from disk. The reasoning is intuitive: map the file, let the kernel page it in, and let the page cache do the caching for you. No buffer pool to size, no read loop to write, no copy from kernel to userspace. This works beautifully in a benchmark that fits in RAM and falls apart the moment the working set exceeds it.

The failure mode is not a slow query. It is a swap storm, followed by the OOM killer, followed by a restart loop that looks like a database problem but is really a virtual memory accounting problem. Understanding why requires looking at what mmap actually promises and what it does not.

What mmap actually buys you

A memory-mapped file gives you a virtual address range backed by a file on disk. Reading from that range triggers a page fault if the page is not resident. The kernel reads the page from disk into the page cache, maps it into your address space, and resumes the faulting instruction. Subsequent reads of the same page are just memory accesses.

The appeal for a vector index is obvious. A flat index of 1M vectors at 768 dimensions and float32 is roughly 3 GB. An IVF or HNSW index adds graph or list structure on top. You do not want to copy that into a heap buffer on every query. Mapping it means the kernel decides what stays hot, and the page cache is shared across all processes on the machine.

int fd = open("index.bin", O_RDONLY);
struct stat st;
fstat(fd, &st);
void *base = mmap(NULL, st.st_size, PROT_READ, MAP_SHARED, fd, 0);
// base now points at the index; reads fault pages in on demand

The catch is that MAP_SHARED file-backed pages are reclaimable. The kernel can evict any clean page at any time, because the backing file is the source of truth. That is the whole point of a page cache: it is a cache, and caches get evicted under pressure.

The swap storm

Here is the sequence that bites teams in production. A vector index is mapped into a long-lived process. The machine has enough RAM for the index plus the rest of the workload, but not much headroom. A burst of queries touches a large fraction of the index, pulling pages into the page cache. Simultaneously, another process allocates anonymous memory, or the same process grows its heap.

The kernel now has two competing demands. It can evict clean file-backed pages, which is cheap and requires no writeback. Or it can push anonymous pages to swap. Under memory pressure, the reclaim path starts scanning, and the LRU lists for file-backed and anonymous pages are managed separately. If the anonymous pages are hot and the file-backed pages are cold, the kernel evicts the file pages and everything is fine. If the anonymous pages are cold and the file pages are hot, the kernel swaps out anonymous memory.

That is the swap storm. The workload is now doing disk I/O for two reasons at once: refaulting evicted index pages and swapping anonymous memory in and out. Latency spikes, the reclaim path runs more aggressively, and the process can spend most of its time in the kernel rather than doing useful work.

A common pattern is to see this diagnosed as a database tuning problem. It is not. It is a memory accounting problem, and the fix is usually to stop pretending that a memory-mapped index is free.

mlock is not a solution

mlock pins pages in physical memory, preventing them from being evicted or swapped. The instinct is to lock the whole index and be done with it.

mlock(base, st.st_size);

This works until it does not. Locking pages consumes the RLIMIT_MEMLOCK budget, which is often small by default. Raising it requires privileges and, more importantly, changes the machine's memory contract. Locked pages are not available for anything else. If you lock a 3 GB index on a 4 GB machine, you have 1 GB left for the kernel, the rest of the process, and every other tenant. The OOM killer does not care that your pages are locked; it will kill something, and it may kill you.

Locking also does not solve the underlying problem. It converts a soft performance problem into a hard capacity problem. You have traded page-fault latency for a guarantee that the index must fit in RAM. If it does not fit, you have made things worse, not better.

There is a narrower use for mlock: pinning a small, hot subset of the index, such as the top-level graph nodes in an HNSW structure or the coarse centroids in an IVF index. That is a defensible optimization. Locking the entire index is usually a sign that the access pattern was never characterized.

read/write with pread

The alternative is to manage the I/O yourself with pread, reading index pages into a buffer pool that you control.

ssize_t n = pread(fd, buf, page_size, offset);
// buf is now in your address space; you decide when to evict it

This is more work. You need a buffer pool, an eviction policy, and a way to map logical index offsets to physical file offsets. You also lose the kernel's readahead heuristics, which are genuinely good for sequential scans and genuinely bad for random access patterns.

What you gain is control. Your buffer pool is anonymous memory that you can size explicitly. You know exactly how much RAM the index will consume. You can implement an eviction policy that matches the access pattern, such as keeping the top levels of a graph index resident and evicting leaf pages. You can instrument hit rates and refault counts directly.

The tradeoff is that you are now writing a cache. That is not a small amount of code, and it is easy to get wrong. But it is a bounded problem with measurable behavior, which is more than can be said for relying on the kernel's global reclaim policy under mixed workloads.

Why pgvector inherits this problem

Postgres already has a buffer pool. It manages shared buffers explicitly, with its own eviction policy and its own accounting. When an extension stores index data in the normal Postgres heap and index files, it goes through that buffer pool. The kernel's page cache is a second layer underneath, but Postgres is not relying on it for correctness or for its primary caching strategy.

The interesting question is what happens when an extension stores large index structures outside the normal buffer pool, or when the index is large enough that the buffer pool cannot hold a meaningful fraction of it. Postgres will read pages through its own I/O path, which means pread-style reads into shared buffers. That is the read/write model, and it has the properties described above: predictable memory usage, explicit eviction, and no swap storm from file-backed pages competing with anonymous memory.

The cost is that Postgres has to manage the working set. If the index is much larger than shared buffers, every query that touches cold pages pays a read from disk. There is no free lunch; there is only a choice about who pays and when.

Choosing between them

The decision is not about which is faster in a microbenchmark. It is about what happens under memory pressure, and that depends on the workload shape.

A few heuristics that hold up in practice:

  • If the index fits comfortably in RAM with headroom for the rest of the workload, mmap is fine. The page cache will hold it, faults will be rare, and the simplicity is worth it.
  • If the index is larger than RAM but the access pattern is skewed, so that a small hot subset serves most queries, mmap plus targeted mlock on the hot subset can work. Measure the refault rate.
  • If the index is larger than RAM and the access pattern is broad, or if the machine runs mixed workloads, explicit buffer management with pread is more predictable. You will pay for it in code complexity, but you will not be surprised by the OOM killer.
  • If you are inside Postgres, prefer the buffer pool. It exists for this reason, and fighting it usually means fighting the rest of the system too.

The common thread is that memory-mapped I/O is a bet that the kernel's global reclaim policy will make good decisions for your workload. Sometimes it does. When it does not, the failure is not graceful degradation; it is a swap storm that takes down the machine.

Instrumentation that actually helps

The metrics that matter are not query latency percentiles. They are the ones that tell you what the kernel is doing.

/proc/vmstat exposes pgmajfault, pgrefill, and the workingset_* counters. A rising major fault rate on a mapped index means pages are being evicted and refaulted. pgscan_kswapd and pgscan_direct tell you how hard the reclaim path is working. If direct reclaim is running during queries, you are already in the bad regime.

For anonymous memory, pgpgin and pgpgout distinguish swap traffic from file I/O. A swap storm shows up as sustained pgpgout with no corresponding application-level write activity.

On the application side, instrument the refault rate directly if you manage your own buffer pool. If you use mmap, you can sample mincore to estimate residency, though it is a coarse tool and the syscall itself has overhead.

The point is to measure the thing that is actually failing. Query latency is a symptom. Page reclaim is the cause.

The uncomfortable conclusion

Memory mapping is not a caching strategy. It is a delegation of caching strategy to the kernel, and the kernel is optimizing for the whole machine, not for your index. That is usually the right call for a general-purpose system and usually the wrong call for a latency-sensitive service with a working set that does not fit in RAM.

The teams that avoid the swap storm are the ones that decide explicitly: either the index fits, and mmap is a reasonable convenience, or it does not, and they manage memory themselves. The teams that get burned are the ones that assume mmap means the index is in RAM, because it does not. It means the index is addressable, and addressable is a very different thing from resident.

If you take one thing from this, take the failure mode. It is not a slow query. It is a machine that is doing two kinds of I/O at once, neither of which is the I/O you designed for, and a kernel that is making a reasonable decision for the system as a whole and a catastrophic one for your service.

#infrastructure#memory-mapping#performance#pgvector#postgres
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.