Silent OOMs on Consumer GPUs: Detecting and Mitigating VRAM Fragmentation in Long-Running Inference

by
Silent OOMs on Consumer GPUs: Detecting and Mitigating VRAM Fragmentation in Long-Running Inference

Consumer GPUs, while cost-effective for many inference workloads, present unique challenges for long-running processes due to VRAM fragmentation. Over time, repeated allocation and deallocation of tensors of varying sizes cause the virtual address space to become divided into many small free blocks. A large allocation, such as model weights, then fails silently—the CUDA allocator returns a null pointer, resulting in NaN outputs or driver hangs. This essay examines the fragmentation mechanism, detection techniques, and mitigation patterns.

The Fragmentation Mechanism

CUDA's virtual memory manager on consumer GPUs allocates physical VRAM in pages (e.g., 64 KB or 2 MB) on demand. When intermediate tensors (activations, gradients) are frequently created and destroyed, the free space becomes fragmented. Although total free memory may be high, no single contiguous block is large enough for a subsequent large request. The allocator's default best-fit strategy exacerbates this, leaving holes that accumulate over hours of inference.

Detecting Fragmentation

A key metric is the fragmentation ratio: 1 - (largest free block size / total free memory). A ratio approaching 0.5 or higher indicates danger. One can estimate the largest free block by binary search: try allocating increasingly large dummy tensors until failure, then approximate the largest contiguous space. Monitoring this ratio periodically—say, every few hundred inference steps—allows early detection. Below is a simplified Python skeleton using PyTorch:

import torch

def fragmentation_ratio():
    free, total = torch.cuda.mem_get_info()
    lo, hi = 0, free
    while lo < hi:
        mid = (lo + hi + 1) // 2
        try:
            t = torch.empty(mid, dtype=torch.uint8, device='cuda')
            del t
            lo = mid
        except RuntimeError:
            hi = mid - 1
    largest = lo
    return 1.0 - (largest / free) if free > 0 else 0.0

This approach adds overhead but gives a direct measure of address space health.

Mitigation Strategies

Three complementary patterns exist: pre-allocation, defragmentation scheduling, and fragment-aware allocation.

1. Pre-allocate a Memory Pool

Instead of relying on CUDA's dynamic allocator, reserve a large contiguous block at startup and use a sub-allocator for all tensor storage. This eliminates external fragmentation entirely. PyTorch's custom allocator interface or CUDA memory pools (cudaMemPool) can be configured for this purpose, though consumer GPU support may be limited. The trade-off is reduced memory available for other processes or models.

2. Defragmentation Scheduling

If pre-allocation is infeasible, periodic defragmentation can consolidate live tensors. The idea is to copy all active tensors into a freshly allocated contiguous block, then free the old fragmented storage. This operation requires pausing inference during the copy, making it suitable for batch or idle periods. A double-buffer pattern—running two model instances on separate pools and switching traffic—can hide the pause.

3. Fragment-Aware Allocation

Reduce fragmentation by rounding allocation sizes to a multiple of a large alignment (e.g., 2 MB), which encourages reuse of free chunks. Additionally, a slab allocator for common tensor sizes (e.g., activation buffers) can pre-allocate fixed-size blocks and recycle them, preventing small allocations from fragmenting the address space.

Caveats

  • Pre-allocation reduces headroom for dynamic memory demands. For multi-model deployments, partition the pool explicitly.
  • Defragmentation is not atomic; inference must be paused. The double-buffer approach requires extra VRAM and complexity.
  • Consumer GPUs lack ECC and may exhibit subtle memory corruption under fragmentation. Adding checksums on model outputs can detect silent errors early.

Conclusion

VRAM fragmentation is a silent threat to long-running inference on consumer GPUs. By monitoring the fragmentation ratio and adopting a combination of pre-allocation, defragmentation scheduling, and aligned allocation, practitioners can achieve stable runs without crashing or silent data corruption. These patterns require careful engineering but are essential for production reliability on budget hardware.

#bare-metal#gpu#hardware-failures#inference#memory-management#vram-fragmentation
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related