Durable Checkpointing Under Fire: Recovering Agents After a GPU Hard Lock Without Lost Steps
How to build agent systems that survive GPU crashes with zero step loss using Postgres and systemd
Durable Checkpointing Under Fire: Recovering Agents After a GPU Hard Lock Without Lost Steps
You're running a swarm of autonomous agents on a single GPU node. Each agent loops: observe, reason, act. The GPU hard locks—maybe a thermal event, a driver bug, or a power glitch. When the machine comes back, your agents are gone. Their in-memory state, their partial reasoning, their pending actions—all vaporized.
Most people treat this as a restart problem. They lose the current step, maybe the last few. But in agent systems, losing a step means losing a decision. That can corrupt downstream state, break invariants, or waste expensive LLM calls. You need durable checkpointing that survives a GPU hard lock with zero step loss.
This article shows how to build that. We'll use Postgres as the checkpoint store, systemd for automatic recovery, and idempotent replay to ensure every step is executed exactly once.
The Problem: GPU Hard Locks Are Not Graceful
A GPU hard lock means the kernel driver stops responding. The GPU can't execute new work, and any running CUDA kernels hang. The OS may eventually recover via a GPU reset, but often the whole machine needs a hard reboot. During that window, any agent state in GPU memory or even in host memory (if tied to CUDA contexts) is lost.
Your agent's step loop typically looks like:
while True:
state = load_state() # in-memory
observation = sense() # might use GPU (vision, etc.)
action = agent.act(observation) # GPU inference
execute(action) # side effect
state = update_state(state, action)
persist_state(state) # periodic, maybe every N stepsIf the GPU locks during agent.act(), the in-memory state is corrupted or lost. The persist_state() call never happens. After reboot, you have no record of the last step. The agent either repeats it (duplicate action) or skips it (lost action). Neither is acceptable.
The Solution: Checkpoint Every Step to Postgres
Checkpointing to Postgres after every step is the only way to guarantee zero step loss. Postgres is crash-safe, transactional, and runs on a separate filesystem (or even a separate machine). We'll use psycopg2 (or asyncpg) to write checkpoints synchronously within the step loop.
Schema
CREATE TABLE agent_checkpoints (
agent_id UUID NOT NULL,
step_number BIGINT NOT NULL,
state JSONB NOT NULL,
created_at TIMESTAMPTZ DEFAULT NOW(),
PRIMARY KEY (agent_id, step_number)
);
CREATE TABLE agent_actions (
agent_id UUID NOT NULL,
step_number BIGINT NOT NULL,
action JSONB NOT NULL,
executed_at TIMESTAMPTZ DEFAULT NOW(),
PRIMARY KEY (agent_id, step_number)
);We store both the state and the action taken. This enables idempotent replay.
Checkpointing Loop
import psycopg2
from uuid import uuid4
conn = psycopg2.connect("dbname=agents host=localhost")
cur = conn.cursor()
agent_id = uuid4()
step = 0
state = initial_state()
while True:
# Before acting, checkpoint the current state and step
cur.execute(
"""
INSERT INTO agent_checkpoints (agent_id, step_number, state)
VALUES (%s, %s, %s)
ON CONFLICT (agent_id, step_number) DO NOTHING
""",
(agent_id, step, json.dumps(state))
)
conn.commit()
observation = sense() # may use GPU
action = agent.act(observation) # GPU inference
execute(action) # side effect
# Record the action taken
cur.execute(
"""
INSERT INTO agent_actions (agent_id, step_number, action)
VALUES (%s, %s, %s)
""",
(agent_id, step, json.dumps(action))
)
conn.commit()
state = update_state(state, action)
step += 1Key points:
- We checkpoint before acting. If the GPU locks during
agent.act(), the checkpoint for the current step already exists. After recovery, we can detect that this step's action is missing and redo it. - We use
ON CONFLICT DO NOTHINGto make the insert idempotent. This matters when we replay. - Every step produces exactly one checkpoint and optionally one action record. The action record is written only after successful execution.
Recovery: systemd + Idempotent Replay
When the machine reboots, systemd starts our agent service. The service runs a recovery script that:
- Connects to Postgres.
- Finds the last checkpoint for each agent.
- Checks if the corresponding action exists. If not, the step was interrupted—re-execute it.
- Resume from the last fully completed step.
systemd Unit
[Unit]
Description=Agent Service
After=network.target postgresql.service
Requires=postgresql.service
[Service]
Type=simple
ExecStart=/usr/local/bin/agent-runner
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.targetWe set Restart=always so systemd automatically relaunches the agent after a crash. The RestartSec=10 gives the GPU driver time to reset.
Recovery Logic
def recover_agent(agent_id):
cur.execute(
"""
SELECT step_number, state
FROM agent_checkpoints
WHERE agent_id = %s
ORDER BY step_number DESC
LIMIT 1
""",
(agent_id,)
)
row = cur.fetchone()
if row is None:
return initial_state(), 0
last_checkpoint_step, state = row
# Check if action for this step exists
cur.execute(
"""
SELECT 1 FROM agent_actions
WHERE agent_id = %s AND step_number = %s
""",
(agent_id, last_checkpoint_step)
)
action_exists = cur.fetchone() is not None
if not action_exists:
# The step was interrupted. Re-execute.
observation = sense()
action = agent.act(observation)
execute(action)
# Record the action
cur.execute(
"""
INSERT INTO agent_actions (agent_id, step_number, action)
VALUES (%s, %s, %s)
""",
(agent_id, last_checkpoint_step, json.dumps(action))
)
conn.commit()
state = update_state(state, action)
# Now step is complete
# Resume from next step
return state, last_checkpoint_step + 1After recovery, the agent continues the loop from the next step. The re-executed action is idempotent because we check for its existence before writing. If the system crashes again during replay, the same logic applies—the checkpoint is already there, and we'll try again.
Handling GPU Dependencies
sense() and agent.act() may use GPU resources (e.g., vision models, LLM inference). After a hard lock, the GPU driver may be in a bad state. Our recovery script should explicitly reset the GPU:
#!/bin/bash
# Reset GPU before starting agent
nvidia-smi --gpu-reset
# Or for AMD: rocm-smi --resetWe can run this in a ExecStartPre hook in systemd:
[Service]
ExecStartPre=/usr/local/bin/reset-gpu.sh
ExecStart=/usr/local/bin/agent-runnerThis ensures the GPU is in a clean state before the agent tries to use it.
Testing the Recovery
Simulate a hard lock by killing the agent process mid-step:
# Run agent in background
./agent-runner &
AGENT_PID=$!
sleep 5 # let it run a few steps
kill -9 $AGENT_PID # simulate hard lock
# Wait for systemd to restart
sleep 15
# Check logs: agent should recover and continue from last checkpointCheck Postgres to verify no duplicate steps:
SELECT agent_id, step_number, COUNT(*)
FROM agent_actions
GROUP BY agent_id, step_number
HAVING COUNT(*) > 1;Should return zero rows.
Performance Considerations
Checkpointing every step to Postgres adds latency—typically 1-5ms per step for a local Postgres. For most agent loops (which involve LLM calls taking seconds), this is negligible. If your steps are sub-millisecond, batch checkpoints every N steps, but accept the risk of losing up to N steps.
Use pgbouncer or connection pooling to avoid connection overhead. Use asyncpg for non-blocking writes if your agent is async.
Alternatives and Why They Fail
- File-based checkpoints: Not crash-safe. A power loss during write corrupts the file. Postgres's WAL guarantees atomicity.
- In-memory replication: Still lost on reboot.
- Checkpointing only on graceful shutdown: Hard locks are not graceful.
- External orchestration (Kubernetes): Adds complexity; still need durable storage. Postgres is simpler for a single-node setup.
Conclusion
GPU hard locks are inevitable in production. By checkpointing every step to Postgres and using systemd for automatic restart with idempotent replay, you can build agent systems that survive crashes without losing a single step. The pattern is simple, reliable, and uses tools you already know.
Next time your GPU locks, your agents won't skip a beat.