Qdrant HNSW Rebuilds Under Live Load
Why segment merges can quietly starve the GPU that serves your inference traffic, and how to schedule around them.
Vector databases are usually described as a retrieval concern. In practice, a Qdrant node running HNSW indexing is a compute concern that happens to live next to your inference stack, and it competes for the exact same resources: CPU cycles, memory bandwidth, page cache, and eventually the PCIe path to the GPU. When a segment merge or a full HNSW rebuild fires during peak inference traffic, the symptoms show up as inference latency, not as a database error. That is the part that makes it hard to debug.
This is not a Qdrant bug. It is an architectural property of running an approximate nearest neighbour index that must be periodically rebuilt, on the same host that is serving a latency-sensitive model. Teams typically notice it as a mysterious p99 spike on the inference side, then find the correlation in CPU pressure or memory bandwidth, then discover the merge scheduler was the trigger.
What HNSW rebuilds actually do
HNSW is a layered proximity graph. Every point inserted into a collection has to be linked into the graph at each layer, and the graph has to remain navigable. Qdrant handles this with segments: a collection is a set of immutable-ish segments, each with its own HNSW graph. New writes go into a fresh segment. When segments get small or numerous, Qdrant merges them into larger ones. The merge is where the cost concentrates, because the resulting segment needs a rebuilt graph over the union of points.
Two distinct operations get conflated under "reindex":
- Segment merge: combine N small segments into one. Requires re-inserting every point into a new graph. Cost scales with total points, not with the delta.
- Index build on a new segment: the incremental path for freshly written points. Cheaper per point, but still graph construction.
Both are CPU-bound with a heavy memory-bandwidth component. HNSW construction is essentially a sequence of greedy graph traversals with distance computations; the working set is large and pointer-chasing dominates once the graph exceeds L3. On a host that also runs vLLM, that memory bandwidth is the same bandwidth the model's KV cache and weight streaming depend on.
Why the GPU shows the pain
A common misconception is that indexing is CPU work and therefore isolated from the GPU. That holds only if the two are on separate machines, or if the CPU has enough headroom that the indexing threads never saturate memory controllers or the PCIe/NVLink path.
On a shared host, the coupling is indirect but real:
- Tokenizer and preprocessing threads on the CPU compete with indexing threads for cores. If the inference server is configured to use most cores for its own thread pools, indexing either gets starved (and merges pile up) or steals cycles (and tokenization latency rises).
- Memory bandwidth contention raises the effective latency of every CPU-side operation in the inference path, including request handling, sampling, and logits post-processing.
- Page cache pressure from reading segment files for a merge can evict model weights or tokenizer caches that were warm.
- NUMA effects. If the indexing threads land on the socket that owns the GPU's PCIe root complex, they compete with the data path that feeds the GPU.
The result is that GPU utilization can look normal while end-to-end latency degrades, because the bottleneck moved upstream. That is why "the GPU is fine" and "the service is slow" are not contradictory observations.
Observing it without inventing numbers
You cannot fix what you cannot see. The signals worth collecting on the Qdrant side are:
- Segment count per collection, and the distribution of segment sizes. A growing tail of small segments is the precursor to a merge storm.
- Merge and index-build duration histograms, not just averages. Merges are bursty.
- Thread pool saturation for the optimizer. Qdrant exposes optimizer status per collection; watch for a backlog of pending operations.
- On the host side:
node_cpu_seconds_totalby mode, memory bandwidth counters if your hardware exposes them (uncore/raplon Intel, equivalent PMUs elsewhere), and per-cgroup CPU throttling.
A minimal Prometheus query to catch merge activity is unglamorous but effective:
eyJsYW5nIjoicHJvbXFsIiwidGV4dCI6IiMgcGVuZGluZyBvcHRpbWl6ZXIgb3BlcmF0aW9ucyBwZXIgY29sbGVjdGlvblxubWF4IGJ5IChjb2xsZWN0aW9uKSAocWRyYW50X2NvbGxlY3Rpb25fb3B0aW1pemVyX3BlbmRpbmdfb3BlcmF0aW9ucylcblxuIyBDUFUgdGhyb3R0bGluZyBvbiB0aGUgaW5mZXJlbmNlIGNvbnRhaW5lcidzIGNncm91cFxucmF0ZShjb250YWluZXJfY3B1X2Nmc190aHJvdHRsZWRfc2Vjb25kc190b3RhbHtuYW1lPVwidmxsbVwifVs1bV0pIn0=If container_cpu_cfs_throttled_seconds_total rises in the same windows as optimizer activity, you have your correlation. It is not proof of causation, but it is enough to justify an experiment.
Scheduling merges away from inference peaks
Qdrant does not ship a cron for merges, but you can shape when they happen by controlling write patterns and optimizer thresholds. The practical levers:
Batch your writes. Instead of trickling points in continuously, write in bursts during a maintenance window. Each burst creates one segment; fewer segments mean fewer merges. This is the single highest-leverage change for most teams.
Tune optimizers_config. The defaults are reasonable for general use but aggressive for latency-sensitive co-tenancy. Raising indexing_threshold and memmap_threshold delays when segments get indexed and when they get memory-mapped, trading recall freshness for fewer rebuilds. Lowering max_optimization_threads bounds how much CPU the optimizer can take, at the cost of longer merge windows.
{
"optimizers_config": {
"default_segment_number": 4,
"max_optimization_threads": 2,
"indexing_threshold": 20000,
"memmap_threshold": 50000
}
}Separate the roles. The cleanest fix is architectural: run the indexing-heavy node on a different host from the inference node, and replicate. Qdrant's shard replication lets you put a write-heavy replica on a machine with spare CPU and a read replica next to the GPU. Reads for retrieval then hit the co-located replica, which only ever serves queries and never rebuilds.
Use collection aliases for cutovers. When you do need a full rebuild, build into a new collection and swap the alias. This avoids the in-place merge cost entirely and gives you a rollback path. It costs double storage for the duration, which is usually acceptable.
# build new_collection, then atomically repoint
curl -X POST localhost:6333/collections/aliases -H 'Content-Type: application/json' -d '{
"actions": [
{"delete_alias": {"alias_name": "docs"}},
{"create_alias": {"collection_name": "docs_v2", "alias_name": "docs"}}
]
}'Multitenancy makes it worse
Single-tenant collections hide the problem because segment counts stay low. Multitenant deployments, where one collection holds many tenants separated by a payload filter, are where merges get expensive. Every tenant's writes land in the same segment set, so a merge rebuilds a graph that spans all tenants. The graph is larger, the traversal working set is larger, and the rebuild duration grows with total points rather than with the tenant that wrote.
Two mitigations are worth considering:
- Partition by tenant into separate collections when tenant counts are modest. This trades collection-management overhead for bounded merge cost. Qdrant supports many collections, but each has fixed overhead, so this works best in the tens-to-low-hundreds range.
- Use tenant-aware sharding where available, so a merge only touches the shards belonging to active tenants. This keeps the blast radius proportional to write activity rather than to total corpus size.
The wrong answer is to leave everything in one collection and hope the optimizer behaves. It will not, because the optimizer's cost model does not know about your latency SLO.
A co-tenancy checklist
If you must run indexing and inference on the same host, these are the controls that matter:
- Pin the inference process to a dedicated CPU set with
cpusetcgroups, and leave the rest for indexing. Do not let them share cores. - Pin the GPU's NUMA node. On multi-socket hosts, bind both the inference process and the vector store to the socket that owns the GPU.
- Cap optimizer threads explicitly. Unbounded indexing threads will saturate memory bandwidth.
- Set
vm.swappinesslow and size page cache deliberately, so merges do not evict warm model state. - Run merges during known-low traffic windows by controlling write bursts, not by hoping.
- Alert on optimizer backlog, not just on query latency. Backlog is the leading indicator; latency is the lagging one.
None of this is exotic. It is the same co-tenancy discipline that applies to any two latency-sensitive workloads sharing a machine. The reason it surprises people is that vector search is mentally filed under "database," and databases are assumed to be someone else's problem until they are not.
The takeaway
HNSW rebuilds are not free background maintenance. They are a full graph reconstruction over your corpus, they are CPU and memory-bandwidth bound, and on a shared host they will compete with the inference path in ways that show up as GPU-adjacent latency. The fix is not a single flag. It is a combination of write batching, optimizer tuning, role separation, and in the multitenant case, partitioning so that merge cost scales with activity rather than with total size. Treat the vector index as a first-class compute workload and schedule it accordingly, and the mysterious p99 spikes stop being mysterious.