Silent OOMs on Consumer GPUs: Detecting and Mitigating VRAM Fragmentation in Long-Running Inference
Consumer GPUs, while cost-effective for many inference workloads, present unique challenges for long-running processes due to VRAM fragmentation. Over time, repeated allocation and deallocation of tensors of varying sizes cause the virtual address space to become divided into many small free blocks. A large allocation, such as model weights, then fails silently—the CUDA allocator returns a null pointer, resulting in NaN outputs or driver hangs. This essay examines the fragmentation mechanism, detection techniques, and mitigation patterns.
The Fragmentation Mechanism
CUDA's virtual memory manager on consumer GPUs allocates physical VRAM in pages (e.g., 64 KB or 2 MB) on demand. When intermediate tensors (activations, gradients) are frequently created and destroyed, the free space becomes fragmented. Although total free memory may be high, no single contiguous block is large enough for a subsequent large request. The allocator's default best-fit strategy exacerbates this, leaving holes that accumulate over hours of inference.
Detecting Fragmentation
A key metric is the fragmentation ratio: 1 - (largest free block size / total free memory). A ratio approaching 0.5 or higher indicates danger. One can estimate the largest free block by binary search: try allocating increasingly large dummy tensors until failure, then approximate the largest contiguous space. Monitoring this ratio periodically—say, every few hundred inference steps—allows early detection. Below is a simplified Python skeleton using PyTorch:
import torch
def fragmentation_ratio():
free, total = torch.cuda.mem_get_info()
lo, hi = 0, free
while lo < hi:
mid = (lo + hi + 1) // 2
try:
t = torch.empty(mid, dtype=torch.uint8, device='cuda')
del t
lo = mid
except RuntimeError:
hi = mid - 1
largest = lo
return 1.0 - (largest / free) if free > 0 else 0.0This approach adds overhead but gives a direct measure of address space health.
Mitigation Strategies
Three complementary patterns exist: pre-allocation, defragmentation scheduling, and fragment-aware allocation.
1. Pre-allocate a Memory Pool
Instead of relying on CUDA's dynamic allocator, reserve a large contiguous block at startup and use a sub-allocator for all tensor storage. This eliminates external fragmentation entirely. PyTorch's custom allocator interface or CUDA memory pools (cudaMemPool) can be configured for this purpose, though consumer GPU support may be limited. The trade-off is reduced memory available for other processes or models.
2. Defragmentation Scheduling
If pre-allocation is infeasible, periodic defragmentation can consolidate live tensors. The idea is to copy all active tensors into a freshly allocated contiguous block, then free the old fragmented storage. This operation requires pausing inference during the copy, making it suitable for batch or idle periods. A double-buffer pattern—running two model instances on separate pools and switching traffic—can hide the pause.
3. Fragment-Aware Allocation
Reduce fragmentation by rounding allocation sizes to a multiple of a large alignment (e.g., 2 MB), which encourages reuse of free chunks. Additionally, a slab allocator for common tensor sizes (e.g., activation buffers) can pre-allocate fixed-size blocks and recycle them, preventing small allocations from fragmenting the address space.
Caveats
- Pre-allocation reduces headroom for dynamic memory demands. For multi-model deployments, partition the pool explicitly.
- Defragmentation is not atomic; inference must be paused. The double-buffer approach requires extra VRAM and complexity.
- Consumer GPUs lack ECC and may exhibit subtle memory corruption under fragmentation. Adding checksums on model outputs can detect silent errors early.
Conclusion
VRAM fragmentation is a silent threat to long-running inference on consumer GPUs. By monitoring the fragmentation ratio and adopting a combination of pre-allocation, defragmentation scheduling, and aligned allocation, practitioners can achieve stable runs without crashing or silent data corruption. These patterns require careful engineering but are essential for production reliability on budget hardware.