Cooling a 4-GPU Swarm Without Aircon: A Fan, a Script, and a 30°C Delta

by

Packing four GPUs into a single chassis or a tight rack without dedicated air conditioning is a recipe for thermal throttling. The manufacturer's stock cooling works under standard loads, but sustained AI inference or training pushes junction temperatures past 85°C, often triggering power limits. This article describes a practical, low-cost solution: a high-static-pressure fan controlled by a script that reads GPU temperatures and adjusts fan speed dynamically. The approach can yield a 30°C delta between GPU hotspot and ambient, even in a warm room. No HVAC upgrade required.

The Problem: Dense GPUs, Limited Airflow

Modern GPUs (e.g., NVIDIA A4000, RTX 4090) draw 200–350W each. In a four-GPU setup, that's 800–1400W of heat concentrated in a small volume. Server rooms without air conditioning rely on ambient air movement, but stock chassis fans are often inadequate. Hotspots develop around GPU intake vents, recirculation occurs, and the BMC's fan curve is too conservative. The result: GPUs hover near their thermal limit, clock down, and inference latency spikes.

A common pattern is to remove side panels and point a floor fan at the cards. That helps, but uncontrolled airflow can be noisy and inconsistent. The better approach is a controlled fan that responds to real-time sensor data.

The Setup: Hardware and Fan Choice

The hardware stack is minimal: a single 120mm or 140mm high-static-pressure fan (e.g., Noctua NF-A14 industrialPPC, 3000 RPM variant), a PWM controller (or motherboard header if available), and a simple microcontroller (Arduino Nano or Raspberry Pi Pico) to relay PWM signals if the host lacks a free header. The fan is mounted directly in front of the GPU intake side, either inside the chassis (if space permits) or externally with a shroud. The goal is to force cool air across the GPU heatsinks, not just circulate air inside the room.

The Script: PWM Control from GPU Temperature

A small daemon runs on the host, polling nvidia-smi every 1–2 seconds for temperature metrics. It then maps the highest GPU temperature (typically the hotspot) to a PWM duty cycle. The mapping is a piecewise linear function: below 50°C fan runs at 30% (silent), above 80°C it ramps to 100%. A hysteresis band prevents rapid cycling.

Below is a simplified version in Python. It assumes the fan is connected to a PWM pin on a microcontroller that communicates via serial, but the logic is identical for direct GPIO or sysfs PWM.

import subprocess
import time
import serial

def get_gpu_temp():
    """Return the highest GPU hotspot temperature from nvidia-smi."""
    out = subprocess.check_output([
        "nvidia-smi",
        "--query-gpu=temperature.gpu",
        "--format=csv,noheader,nounits"
    ]).decode().strip().split('\n')
    temps = [int(t) for t in out if t.isdigit()]
    return max(temps) if temps else 60

def temp_to_pwm(temp):
    """Map temperature to PWM duty cycle (0-255)."""
    if temp <= 50:
        return 76  # 30%
    elif temp >= 80:
        return 255  # 100%
    else:
        # linear interpolation
        return int(76 + (temp - 50) * (255 - 76) / (80 - 50))

def main():
    # Adjust serial port and baud rate
    ser = serial.Serial('/dev/ttyUSB0', 9600, timeout=1)
    time.sleep(2)
    try:
        while True:
            temp = get_gpu_temp()
            pwm = temp_to_pwm(temp)
            ser.write(bytes([pwm]))
            time.sleep(1)
    except KeyboardInterrupt:
        ser.close()

if __name__ == '__main__':
    main()

If the fan is connected directly to the motherboard, the script can write to /sys/class/hwmon/hwmon*/pwm1 after enabling manual control. The principle is identical.

Architecture: Why This Works

The key insight is that GPU coolers are designed for open airflow, not recirculating hot air inside a case. By providing a high volume of cool air directly to the intake, the heatsink fins operate closer to ambient. The fan's static pressure pushes through dense fin stacks, overcoming the resistance of neighboring cards. The script ensures the fan only spins fast when needed, saving noise and wear during idle.

This pattern is common in DIY crypto mining rigs, but it applies equally to AI workloads. The 30°C delta (e.g., 35°C ambient → 65°C hotspot, versus 85°C with stock chassis fans) is achievable with a well-positioned fan. The exact delta depends on ambient temperature, GPU load, and fan placement. In practice, engineers see improvements of 20–30°C over passive or stock fan configurations.

Implementation Details and Pitfalls

  • PWM Polarity: Some fans expect a 5V PWM signal, others 3.3V. Check the datasheet or use a level shifter.
  • Fan Start-up: At very low duty cycles (<20%) many fans stall. Set a minimum of 30%.
  • Serial Latency: A microcontroller adds ~5ms per command, negligible for thermal response.
  • Failover: If the script crashes, the fan should default to a safe mid-range (e.g., 50%) via a pull-up resistor or watchdog timer.

Monitoring and Validation

Use nvidia-smi dmon to log temperature over time. Plot a graph before and after implementing the fan control. Typical observations show a 15–30°C reduction in GPU hotspot temperature under sustained 100% load. The fan itself may add 30-40 dB of noise at full speed, so acoustics is a trade-off.

Scaling to More GPUs

For swarms with 4+ GPUs spread across multiple servers, the same per-node script works. A centralized monitoring system (e.g., Prometheus + node_exporter) can track temperatures and fan speeds. The script runs locally on each host, keeping the control loop deterministic and low-latency.

Lessons Learned

  • Airflow direction matters: Exhausting hot air out of the case is as important as intake. Ensure a clear path for hot air to leave.
  • Rack placement: The bottom of a rack tends to be cooler; place GPUs there if possible.
  • Dust filters: High flow fans suck dust. Clean filters every month.
  • Thermal paste: Reapplying high-quality paste to GPU dies can improve heat transfer by several degrees.

This approach is not a replacement for proper HVAC, but for small clusters in home labs or warm server closets, it's a proven, low-cost hack. The combination of a cheap fan and a 50-line script delivers consistent, automatic cooling that keeps GPUs well below their throttle threshold. No air conditioning required.

#bare-metal#fan-control#gpu#infrastructure#monitoring#thermal-management
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.