Bare-metal inference without aircon: how we cool a 4-GPU swarm with a fan and a script

by

We run a 4-GPU inference swarm on three Hetzner servers in a small office closet. No air conditioning. Summer ambient hits 35°C. The GPUs are not datacenter-grade—they are consumer cards with blower fans designed for open-air benches, not stacked in a metal rack. We had to solve thermal throttling without spending on rack cooling.

The solution is not elegant. It is a box fan, a temperature sensor, and a 30-line Python daemon. It works.

Why bare metal at all?

Cloud GPU instances are expensive and come with data sovereignty strings. For our nightly LoRA fine-tuning and Qwen inference on vLLM, we wanted full control. Hetzner dedicated servers are cheap, but their cooling assumes standard CPU loads. Four GPUs pulling 250W each saturate the airflow design.

Our stack: three servers connected via WireGuard mesh. One server hosts two GPUs for inference, the other two each host one GPU for training and vector embedding. All share a Postgres + Qdrant + Neo4j backend. The GPUs are packed in a standard 42U rack with no cold-aisle containment.

The thermal problem

Consumer GPUs throttle at around 85°C junction temperature. Under sustained inference load, they hit 90°C within minutes. The server chassis fans ramp, but the rear exhaust blows hot air into the rack, recirculating through the front intake. Ambient in the closet climbs to 45°C. Thermal runaway is real.

A common pattern in homelab setups is to undervolt the GPUs, but we need full FP16 throughput for LoRA serving. Undervolting drops clocks by 10-15% and introduces instability on some cards. We wanted a solution that does not touch GPU voltage.

The fan

We bought a 20-inch industrial box fan from a hardware store. It sits on a shelf above the rack, pointing downward into the rear of the chassis. The fan moves roughly 3000 CFM on high. It is loud—65 dB—but the office is unattended at night.

The fan alone dropped peak GPU temperature from 90°C to 78°C under full load. That is within spec, but we wanted a margin for summer spikes.

The script

We wrote a Python daemon that reads GPU temperatures via nvidia-smi every 5 seconds and controls the fan speed via a relay board connected over USB serial. The relay board is a cheap 4-channel module with a 5V coil. The fan has three speed settings (low, medium, high) wired to three relays. The fourth relay is unused.

import subprocess
import time
import serial

SERIAL_PORT = '/dev/ttyUSB0'
ser = serial.Serial(SERIAL_PORT, 9600, timeout=1)

def get_gpu_temp():
    output = subprocess.check_output(['nvidia-smi', '--query-gpu=temperature.gpu', '--format=csv,noheader,nounits'])
    temps = [int(x) for x in output.decode().strip().split('\n')]
    return max(temps)

def set_fan_speed(speed):
    # speed: 0=off, 1=low, 2=medium, 3=high
    # relay mapping: ch1=low, ch2=medium, ch3=high
    commands = {
        0: [0,0,0],
        1: [1,0,0],
        2: [0,1,0],
        3: [0,0,1]
    }
    cmd = commands[speed]
    # send serial command to relay board (example protocol)
    ser.write(bytes([0xA0, 0x01, cmd[0], cmd[1], cmd[2], 0xA1]))

while True:
    temp = get_gpu_temp()
    if temp >= 75:
        set_fan_speed(3)
    elif temp >= 65:
        set_fan_speed(2)
    elif temp >= 55:
        set_fan_speed(1)
    else:
        set_fan_speed(0)
    time.sleep(5)

The daemon runs as a systemd service. It logs temperature and fan state to a file for later analysis.

Results

With the fan on automatic, the hottest GPU never exceeds 72°C during sustained inference. The fan only runs on high for about 10 minutes after a heavy batch, then drops to medium. The closet ambient stays under 38°C.

We did not measure power consumption precisely, but the fan draws about 100W on high. That is less than 5% of the total system power. It is a worthwhile trade-off.

What we learned

  • Airflow direction matters. The fan blows downward into the rear exhaust. That pushes hot air away from the intake. A common mistake is to blow air toward the front, which fights the chassis fans.
  • Relay boards are finicky. The cheap ones use a serial protocol that varies by manufacturer. We had to reverse-engineer the command bytes with a logic analyzer. Next time we will buy a board with documented protocol or use a PWM-controlled fan.
  • Temperature polling is cheap. nvidia-smi overhead is negligible. We poll every 5 seconds, but 1 second would also work.
  • Fail-safe is important. If the daemon crashes, the fan stays at its last state. We added a watchdog timer that sets the fan to high if no temperature reading for 30 seconds.

Alternatives we considered

  • Water cooling. Too expensive and risky for a rack with multiple cards. Leaks destroy hardware.
  • Rack-mounted AC unit. Starts at $1000 and requires a dedicated circuit. Our closet has one 15A outlet.
  • Undervolting. Works but reduces throughput. We need every FLOP for LoRA training.
  • Moving the rack to a cooler room. Not an option.

The bigger picture

This is not a production datacenter solution. It is a pragmatic hack for a small-scale sovereign AI infrastructure. We run Qwen 14B on vLLM, BGE-M3 for embeddings, and nightly LoRA fine-tuning—all on three servers with a fan and a script.

The point is: you do not need enterprise cooling for a 4-GPU swarm. You need to understand your thermal envelope and control it cheaply.

Future improvements

  • Replace the relay board with a PWM fan controller for smoother speed ramping.
  • Add a temperature sensor for ambient air to adjust the fan curve dynamically.
  • Integrate with our monitoring stack (Prometheus + Grafana) to alert on thermal anomalies.

For now, the fan runs, the GPUs stay cool, and the models keep serving. That is enough.

#bare-metal#cooling#gpu-thermal#homelab#inference#qwen#thermal-management
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.