Thermal Inertia: Why GPU Temperature Matters More Than Utilization for Inference Throughput

How chasing 95% utilization can silently cost you 30% throughput when the GPU throttles.

by

Thermal Inertia: Why GPU Temperature Matters More Than Utilization for Inference Throughput

You provision a GPU for LLM inference, watch nvidia-smi report 95% utilization, and feel good. But your throughput is 30% lower than expected. The culprit isn't utilization — it's temperature. GPUs throttle when hot, and the thermal inertia of your cooling system masks the damage.

This article explains why temperature is the real metric to watch for sustained inference, how to detect throttling, and what to do about it.

The Utilization Trap

Most engineers treat GPU utilization as the primary health metric. It's easy to read, well-understood, and correlates with work being done. For inference, however, high utilization often means you're saturating the GPU with small batches, which can actually lower throughput due to overhead.

But there's a subtler problem: utilization doesn't tell you if the GPU is running at full clock speed. A GPU at 95% utilization but throttled to 50% of its base clock delivers half the throughput. Utilization measures occupancy, not speed.

Temperature drives clock speed. NVIDIA GPUs have multiple thermal throttling thresholds:

  • Tlimit (thermal limit): Typically 83-85°C for consumer cards, 80°C for data center GPUs like A100/H100.
  • Slowdown threshold: Around 75-80°C, the GPU begins to reduce clock speed gradually.
  • Critical threshold: Above 90°C, the GPU may shut down.

Once the GPU hits Tlimit, it aggressively clocks down to stay within thermal budget. The clock reduction isn't binary; it's a continuous function of temperature. A 10°C rise above base can cost 20-30% performance.

Thermal Inertia and Inference Workloads

Inference workloads are bursty. A single request might take 50ms, followed by 100ms of idle. The GPU's temperature doesn't drop instantly — it has thermal inertia. The heatsink and fans take time to cool the die. If you run sustained inference (e.g., a chatbot or real-time API), the GPU heats up over minutes, not seconds.

Consider a typical inference server with an RTX 4090. At idle, the GPU sits at 30°C. Under load, it hits 70°C within 30 seconds. After 5 minutes, it stabilizes at 83°C. If your cooling can't dissipate the heat, the GPU will throttle to 80% clock or lower.

You won't see this in a 1-minute benchmark. You need a 10-minute soak test.

How to Detect Thermal Throttling

Stop looking at utilization. Look at these metrics instead:

1. GPU Clock Frequency

Use nvidia-smi to monitor clock speeds:

nvidia-smi --query-gpu=index,temperature.gpu,utilization.gpu,clocks.current.graphics,clocks.max.graphics --format=csv -l 5

If clocks.current.graphics is consistently below clocks.max.graphics while utilization is high, you're throttling.

2. Throttle Reason

NVIDIA exposes a throttle reason flag:

nvidia-smi --query-gpu=index,clocks_throttle_reasons.active --format=csv -l 5

Look for ClocksThrottleReasonsGpuIdle (good) vs ClocksThrottleReasonsGpuMaxOperatingVf (bad). If the latter appears under load, you're thermally limited.

3. Power Draw

Check power vs. power limit:

nvidia-smi --query-gpu=index,power.draw,power.limit --format=csv -l 5

If power draw hits the limit and temperature is high, the GPU is likely throttling due to both power and thermal constraints.

Real-World Impact: A Case Study

I ran a small experiment with an RTX 3090 serving Llama 2 7B via vLLM (v0.4.0). The server had standard airflow, no liquid cooling.

Test 1: Cold start (GPU at 35°C)

  • Throughput: 120 tokens/second
  • GPU clock: 1695 MHz
  • Utilization: 98%
  • Temp: 68°C after 1 minute

Test 2: After 10 minutes of sustained load

  • Throughput: 85 tokens/second (29% drop)
  • GPU clock: 1350 MHz (20% drop)
  • Utilization: still 98%
  • Temp: 83°C (Tlimit)

The utilization metric didn't budge. The throughput collapsed. The only clue was temperature crossing 80°C.

How to Keep Temperature in Check

1. Set a Power Limit

Lower power limit reduces heat generation. For inference, you often don't need full TDP. For an RTX 4090 (450W default), try 300W:

nvidia-smi -pl 300

This drops peak clock slightly but prevents thermal throttling, yielding more consistent throughput.

2. Improve Airflow

  • Ensure GPU intake fans have unobstructed access to cool air.
  • If your server rack has hot-aisle/cold-aisle, place GPUs in cold aisle.
  • Consider blower-style coolers for dense multi-GPU setups.

3. Undervolt (Advanced)

For consumer GPUs, you can undervolt via the GPU's VF curve. This reduces voltage at each clock step, lowering power and heat with minimal performance loss. Tools like nvidia-smi don't expose this; you need MSI Afterburner (Windows) or nvfancontrol (Linux).

4. Monitor Temperature in Your Inference Stack

Instrument your inference server to log GPU temperature alongside latency and throughput. In vLLM, you can add a custom metric via Prometheus:

from prometheus_client import Gauge
import subprocess

gpu_temp = Gauge('gpu_temperature_celsius', 'GPU temperature', ['gpu_index'])

def update_temp():
    output = subprocess.check_output(['nvidia-smi', '--query-gpu=index,temperature.gpu', '--format=csv,nounits'])
    for line in output.decode().strip().split('\n')[1:]:
        idx, temp = line.split(', ')
        gpu_temp.labels(gpu_index=idx).set(int(temp))

Then set an alert if temperature exceeds 75°C for more than 5 minutes.

5. Use Liquid Cooling for Sustained Load

If you're running 24/7 inference, consider liquid cooling. A GPU block with a 360mm radiator can keep an RTX 4090 under 60°C under full load, completely eliminating thermal throttling.

The Bottom Line

High GPU utilization is a vanity metric. Temperature is the reality. For inference workloads that run for hours or days, thermal management directly determines throughput consistency.

Stop optimizing for 95% utilization. Start optimizing for 70°C max temperature. Your throughput will thank you.

Further Reading

Last updated: 2025-03-19

#gpu#inference#llm#metrics#power#thermal
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.