Detecting GPU PCIe Link Degradation Before Inference Falls Back
Silent PCIe retraining can quietly halve your inference throughput. Here is how to catch it before it does.
PCIe link degradation is one of the quietest failure modes in GPU inference infrastructure. Unlike an XID error or a thermal throttle, a link that retrains from Gen4 x16 down to Gen3 x8 does not throw an exception. It does not log a warning to dmesg by default. It does not crash the process. It just moves less data, and inference throughput drops. If your monitoring only watches GPU utilization, memory usage, and temperature, you will not see it until the performance regression is already affecting users.
The PCIe link is negotiated at boot between the GPU and the root complex. Under normal conditions it stays at its negotiated speed and width. But a marginal connector, a riser cable with poor signal integrity, a failing retimer, or even a dusty slot can cause the link to retrain at a lower speed or width. The kernel and the GPU driver accept this silently because a narrower link is still a functional link. The GPU works. It is just starved.
Why inference is especially sensitive to PCIe bandwidth
Training workloads are compute-bound. Inference workloads are often memory-bandwidth-bound, but they are also PCIe-bound in specific patterns. Consider a typical setup: a host process loads a model into GPU memory, then streams input data over PCIe for each request. For small models and large batches, the input transfer is negligible. For large models, multi-GPU tensor parallelism, or high request concurrency, the PCIe link becomes a real bottleneck.
When you run tensor parallelism across multiple GPUs, every forward pass requires all-reduce operations between GPUs. On systems without NVLink, those collectives traverse PCIe. A Gen4 x16 link provides roughly 32 GB/s bidirectional. Drop to Gen3 x8 and you are at roughly 8 GB/s. The compute kernels finish faster than the data can move, and the GPUs sit idle waiting on the link. The utilization metric drops, but the root cause is not the GPU.
The same applies to CPU-offloaded models, where weights or KV cache pages are swapped between host memory and GPU memory over PCIe. And it applies to any pipeline where preprocessing happens on the CPU and tensors are copied to the GPU per request.
Reading PCIe link state from sysfs
On Linux, the current link speed and width for a PCIe device are exposed under sysfs. You do not need vendor tools for the basic check. For each GPU, find its PCI bus address, then read the link capabilities and status.
# List NVIDIA GPUs and their PCI bus addresses
nvidia-smi --query-gpu=index,pci.bus_id --format=csv,noheader
# For a GPU at 0000:01:00.0, read current link speed and width
cat /sys/bus/pci/devices/0000:01:00.0/current_link_speed
cat /sys/bus/pci/devices/0000:01:00.0/current_link_width
# Read the maximum capable speed and width
cat /sys/bus/pci/devices/0000:01:00.0/max_link_speed
cat /sys/bus/pci/devices/0000:01:00.0/max_link_widthThe current_link_speed file returns values like 16.0 GT/s PCIe for Gen4, 8.0 GT/s PCIe for Gen3, and 5.0 GT/s PCIe for Gen2. The current_link_width returns an integer like 16, 8, 4, or 1.
The critical comparison is current versus max. If current_link_speed is less than max_link_speed, or current_link_width is less than max_link_width, the link has degraded. This can happen at boot (the link trained low and stayed there) or at runtime (the link retrained after a correctable error threshold was crossed).
A simple shell loop can collect this for all GPUs:
#!/usr/bin/env bash
set -euo pipefail
for bdf in $(nvidia-smi --query-gpu=pci.bus_id --format=csv,noheader | tr 'A-Z' 'a-z'); do
dev="/sys/bus/pci/devices/${bdf}"
cur_speed=$(cat "${dev}/current_link_speed")
max_speed=$(cat "${dev}/max_link_speed")
cur_width=$(cat "${dev}/current_link_width")
max_width=$(cat "${dev}/max_link_width")
echo "${bdf} speed=${cur_speed} (max ${max_speed}) width=x${cur_width} (max x${max_width})"
doneThis gives you a point-in-time snapshot. For detection, you need to run it periodically and compare against a known-good baseline.
Correctable errors and AER
PCIe Advanced Error Reporting (AER) exposes correctable and uncorrectable error counters. Correctable errors are, by definition, recovered from by the hardware. But a rising rate of correctable errors is a leading indicator that the link is marginal and may retrain down or fail entirely.
You can read AER counters from sysfs for each device:
# Correctable error count
cat /sys/bus/pci/devices/0000:01:00.0/aer_dev_correctable
# Uncorrectable error count
cat /sys/bus/pci/devices/0000:01:00.0/aer_dev_uncorrectableThe output is a set of named counters, for example:
RxErr 0
BadTLP 0
BadDLLP 0
Rollover 0
Timeout 0
NonFatalErr 0
CorrIntErr 0
HeaderOF 0A nonzero value in BadTLP, BadDLLP, or RxErr that increases over time indicates physical layer problems. The link may still be at full speed and width, but it is accumulating errors. If the correctable error rate crosses a hardware-defined threshold, the link retrains, and you get the silent downgrade.
Not all kernels expose AER counters for all devices. Some require pci=aer or specific driver support. But when available, they are the earliest warning signal.
dmesg and kernel logs
The kernel does log PCIe link retraining events in some cases. Look for messages like:
pcieport 0000:00:01.0: AER: Corrected error received: 0000:01:00.0
pcieport 0000:00:01.0: PCIe Bus Error: severity=Corrected, type=Physical Layer
nvme 0000:02:00.0: PCIe link downFor GPU-specific events, the NVIDIA driver may log link speed changes or XID errors that correlate with PCIe issues. XID 62, for example, is related to internal micro-controller errors that can be triggered by PCIe problems. XID 79 indicates the GPU has fallen off the bus, which is the extreme end of link failure.
A robust monitoring setup should tail dmesg for PCIe and GPU-related messages and alert on patterns that indicate link instability. But dmesg is not a reliable primary source because many retraining events are not logged at all, especially if they happen during boot before the logging subsystem is fully up.
Building a proactive monitor
The pattern that works is a periodic collector that reads link state and AER counters, stores them as time series, and alerts on deviations from baseline. The collector can be a small Python script run by a systemd timer, or a sidecar container that exposes metrics to Prometheus.
Here is a minimal Python collector that emits Prometheus-style metrics:
import glob
import os
import re
PCI_DEVICES = "/sys/bus/pci/devices"
def read_file(path):
try:
with open(path, "r") as f:
return f.read().strip()
except (OSError, IOError):
return None
def parse_speed(speed_str):
# "16.0 GT/s PCIe" -> 16.0
match = re.match(r"([0-9.]+) GT/s", speed_str)
return float(match.group(1)) if match else None
def collect():
for dev_path in glob.glob(os.path.join(PCI_DEVICES, "*")):
class_file = os.path.join(dev_path, "class")
class_id = read_file(class_file)
if not class_id or not class_id.startswith("0x030000"):
continue # not a display controller
bdf = os.path.basename(dev_path)
cur_speed = parse_speed(read_file(os.path.join(dev_path, "current_link_speed")) or "")
max_speed = parse_speed(read_file(os.path.join(dev_path, "max_link_speed")) or "")
cur_width = read_file(os.path.join(dev_path, "current_link_width"))
max_width = read_file(os.path.join(dev_path, "max_link_width"))
if None in (cur_speed, max_speed, cur_width, max_width):
continue
labels = f'device="{bdf}"'
print(f'pcie_link_speed_current{{{labels}}} {cur_speed}')
print(f'pcie_link_speed_max{{{labels}}} {max_speed}')
print(f'pcie_link_width_current{{{labels}}} {cur_width}')
print(f'pcie_link_width_max{{{labels}}} {max_width}')
# AER correctable errors
aer = read_file(os.path.join(dev_path, "aer_dev_correctable"))
if aer:
for line in aer.splitlines():
parts = line.split()
if len(parts) == 2 and parts[1].isdigit():
err_name = parts[0].lower()
print(f'pcie_aer_correctable{{device="{bdf}",type="{err_name}"}} {parts[1]}')
if __name__ == "__main__":
collect()This script can be run by a textfile collector for node_exporter, or adapted to push to any metrics backend. The key is to store the values over time so you can alert on changes.
Alerting rules that actually catch degradation
The naive alert is current_link_speed < max_link_speed. That catches the case where the link is degraded right now. But it does not catch the case where the link is at full speed but accumulating correctable errors, which is the precursor to a retrain.
A better set of rules:
Link speed or width mismatch — alert immediately if current is less than max. This is a hard fault. The link is degraded and inference is likely affected.
Correctable error rate increase — alert if the derivative of any AER correctable counter is positive over a sustained window. For example, if
BadTLPincreases by more than a small threshold over 15 minutes, something is wrong with the physical link.Link retrain event — if you can capture kernel logs, alert on any message that indicates a link speed change or AER correctable error burst. This catches the moment of retraining, not just the aftermath.
Throughput regression correlated with link state — if you have inference throughput metrics, alert when throughput drops and link state is not at max. This ties the symptom to the cause.
A common pattern is to use rate() on the AER counters and alert if the rate is nonzero for more than a few minutes. The exact threshold depends on the hardware and the acceptable error budget, but any sustained correctable error rate on a healthy link should be zero or near zero.
Self-healing remediation
Detection is only half the problem. Once you know a link is degraded, what do you do? The answer depends on whether the degradation is persistent or transient.
Transient retrain: If the link retrained due to a one-off error burst and is now stable at a lower speed, a device-level reset can sometimes trigger a renegotiation back to full speed. For NVIDIA GPUs, nvidia-smi --gpu-reset -i <index> can reset the GPU, but this requires no active processes on the device and may not affect the PCIe link itself. A PCIe link retrain can be triggered by writing to the link control register, but this is risky and not exposed through standard sysfs. In practice, a host reboot is the reliable way to renegotiate the link.
Persistent degradation: If the link comes back degraded after a reboot, the problem is physical. A reseat of the card, a replacement of the riser cable, or a move to a different slot is required. Software cannot fix a bad connector.
Workload mitigation: In the short term, if a GPU is on a degraded link, you can remove it from the inference pool. If you run a scheduler that is aware of GPU health, you can mark the device as unschedulable and drain existing work. This prevents the degraded GPU from dragging down the throughput of the whole service.
A self-healing loop might look like this:
# Pseudocode for a remediation controller
def reconcile(gpu):
state = read_link_state(gpu)
if state.current_speed == state.max_speed and state.current_width == state.max_width:
return # healthy
if state.degraded_since < now() - PERSISTENT_THRESHOLD:
# Persistent degradation: cordon the GPU and alert a human
cordon_gpu(gpu)
page_oncall("GPU %s has persistent PCIe degradation" % gpu.id)
return
# Recent degradation: try a reset if no workloads are running
if gpu.active_processes == 0:
reset_gpu(gpu.id)
# Re-check after reset
new_state = read_link_state(gpu)
if new_state.current_speed < new_state.max_speed:
cordon_gpu(gpu)
page_oncall("GPU %s link did not recover after reset" % gpu.id)The important part is that the controller does not blindly reset GPUs that are serving traffic. It waits for the device to be idle, or it drains the device first. And it escalates to a human when software remediation fails.
Integrating with the inference stack
If you run inference on vLLM or a similar server, you can expose GPU health as a metric that the load balancer or scheduler consumes. A simple approach is to have the monitoring agent write a health file per GPU, and have the scheduler read it before assigning work.
For example, a scheduler might check /var/run/gpu-health/<gpu-id> and skip any GPU where the file contains degraded. This is a pull-based model that avoids tight coupling between the monitor and the scheduler.
On a WireGuard mesh with multiple inference nodes, the same logic applies. Each node runs its own collector and remediation controller. The mesh carries metrics and health state to a central observability stack, but the remediation decisions are local. This avoids a single point of failure and keeps the control loop fast.
What to watch for in practice
Teams that run GPU inference at scale typically observe a few patterns:
- Link degradation is more common on systems with PCIe risers or backplanes than on systems with GPUs mounted directly in the slot. The extra connector is an extra failure point.
- Gen4 and Gen5 links are more sensitive to signal integrity issues than Gen3. A cable that works fine at Gen3 may retrain to Gen2 or x8 at Gen4.
- Correctable errors often precede visible degradation by hours or days. If you only alert on speed and width, you miss the early warning.
- A single degraded GPU in a tensor-parallel group can slow the entire group. The all-reduce waits for the slowest link.
The cost of monitoring is low: a few sysfs reads per GPU per minute. The cost of not monitoring is a silent performance regression that is hard to diagnose after the fact. By the time someone notices that inference is slow, the link state may have changed again, and the evidence is gone.
Conclusion
PCIe link degradation is not a crash. It is a slow, silent narrowing of the pipe that feeds your GPUs. The fix is straightforward: read the link state and AER counters from sysfs, store them as time series, alert on deviations from baseline, and automate remediation where it is safe. The alternative is discovering the problem through user complaints, long after the link retrained and the logs rolled over.