Durable Checkpointing Under Fire: Recovering Agents After a GPU Hard Lock Without Lost Steps

How to build agent systems that survive GPU crashes with zero step loss using Postgres and systemd

by

Durable Checkpointing Under Fire: Recovering Agents After a GPU Hard Lock Without Lost Steps

You're running a swarm of autonomous agents on a single GPU node. Each agent loops: observe, reason, act. The GPU hard locks—maybe a thermal event, a driver bug, or a power glitch. When the machine comes back, your agents are gone. Their in-memory state, their partial reasoning, their pending actions—all vaporized.

Most people treat this as a restart problem. They lose the current step, maybe the last few. But in agent systems, losing a step means losing a decision. That can corrupt downstream state, break invariants, or waste expensive LLM calls. You need durable checkpointing that survives a GPU hard lock with zero step loss.

This article shows how to build that. We'll use Postgres as the checkpoint store, systemd for automatic recovery, and idempotent replay to ensure every step is executed exactly once.

The Problem: GPU Hard Locks Are Not Graceful

A GPU hard lock means the kernel driver stops responding. The GPU can't execute new work, and any running CUDA kernels hang. The OS may eventually recover via a GPU reset, but often the whole machine needs a hard reboot. During that window, any agent state in GPU memory or even in host memory (if tied to CUDA contexts) is lost.

Your agent's step loop typically looks like:

while True:
    state = load_state()          # in-memory
    observation = sense()         # might use GPU (vision, etc.)
    action = agent.act(observation)  # GPU inference
    execute(action)               # side effect
    state = update_state(state, action)
    persist_state(state)          # periodic, maybe every N steps

If the GPU locks during agent.act(), the in-memory state is corrupted or lost. The persist_state() call never happens. After reboot, you have no record of the last step. The agent either repeats it (duplicate action) or skips it (lost action). Neither is acceptable.

The Solution: Checkpoint Every Step to Postgres

Checkpointing to Postgres after every step is the only way to guarantee zero step loss. Postgres is crash-safe, transactional, and runs on a separate filesystem (or even a separate machine). We'll use psycopg2 (or asyncpg) to write checkpoints synchronously within the step loop.

Schema

CREATE TABLE agent_checkpoints (
    agent_id UUID NOT NULL,
    step_number BIGINT NOT NULL,
    state JSONB NOT NULL,
    created_at TIMESTAMPTZ DEFAULT NOW(),
    PRIMARY KEY (agent_id, step_number)
);

CREATE TABLE agent_actions (
    agent_id UUID NOT NULL,
    step_number BIGINT NOT NULL,
    action JSONB NOT NULL,
    executed_at TIMESTAMPTZ DEFAULT NOW(),
    PRIMARY KEY (agent_id, step_number)
);

We store both the state and the action taken. This enables idempotent replay.

Checkpointing Loop

import psycopg2
from uuid import uuid4

conn = psycopg2.connect("dbname=agents host=localhost")
cur = conn.cursor()

agent_id = uuid4()
step = 0
state = initial_state()

while True:
    # Before acting, checkpoint the current state and step
    cur.execute(
        """
        INSERT INTO agent_checkpoints (agent_id, step_number, state)
        VALUES (%s, %s, %s)
        ON CONFLICT (agent_id, step_number) DO NOTHING
        """,
        (agent_id, step, json.dumps(state))
    )
    conn.commit()

    observation = sense()  # may use GPU
    action = agent.act(observation)  # GPU inference
    execute(action)  # side effect

    # Record the action taken
    cur.execute(
        """
        INSERT INTO agent_actions (agent_id, step_number, action)
        VALUES (%s, %s, %s)
        """,
        (agent_id, step, json.dumps(action))
    )
    conn.commit()

    state = update_state(state, action)
    step += 1

Key points:

  • We checkpoint before acting. If the GPU locks during agent.act(), the checkpoint for the current step already exists. After recovery, we can detect that this step's action is missing and redo it.
  • We use ON CONFLICT DO NOTHING to make the insert idempotent. This matters when we replay.
  • Every step produces exactly one checkpoint and optionally one action record. The action record is written only after successful execution.

Recovery: systemd + Idempotent Replay

When the machine reboots, systemd starts our agent service. The service runs a recovery script that:

  1. Connects to Postgres.
  2. Finds the last checkpoint for each agent.
  3. Checks if the corresponding action exists. If not, the step was interrupted—re-execute it.
  4. Resume from the last fully completed step.

systemd Unit

[Unit]
Description=Agent Service
After=network.target postgresql.service
Requires=postgresql.service

[Service]
Type=simple
ExecStart=/usr/local/bin/agent-runner
Restart=always
RestartSec=10

[Install]
WantedBy=multi-user.target

We set Restart=always so systemd automatically relaunches the agent after a crash. The RestartSec=10 gives the GPU driver time to reset.

Recovery Logic

def recover_agent(agent_id):
    cur.execute(
        """
        SELECT step_number, state
        FROM agent_checkpoints
        WHERE agent_id = %s
        ORDER BY step_number DESC
        LIMIT 1
        """,
        (agent_id,)
    )
    row = cur.fetchone()
    if row is None:
        return initial_state(), 0
    last_checkpoint_step, state = row

    # Check if action for this step exists
    cur.execute(
        """
        SELECT 1 FROM agent_actions
        WHERE agent_id = %s AND step_number = %s
        """,
        (agent_id, last_checkpoint_step)
    )
    action_exists = cur.fetchone() is not None

    if not action_exists:
        # The step was interrupted. Re-execute.
        observation = sense()
        action = agent.act(observation)
        execute(action)
        # Record the action
        cur.execute(
            """
            INSERT INTO agent_actions (agent_id, step_number, action)
            VALUES (%s, %s, %s)
            """,
            (agent_id, last_checkpoint_step, json.dumps(action))
        )
        conn.commit()
        state = update_state(state, action)
        # Now step is complete
    # Resume from next step
    return state, last_checkpoint_step + 1

After recovery, the agent continues the loop from the next step. The re-executed action is idempotent because we check for its existence before writing. If the system crashes again during replay, the same logic applies—the checkpoint is already there, and we'll try again.

Handling GPU Dependencies

sense() and agent.act() may use GPU resources (e.g., vision models, LLM inference). After a hard lock, the GPU driver may be in a bad state. Our recovery script should explicitly reset the GPU:

#!/bin/bash
# Reset GPU before starting agent
nvidia-smi --gpu-reset
# Or for AMD: rocm-smi --reset

We can run this in a ExecStartPre hook in systemd:

[Service]
ExecStartPre=/usr/local/bin/reset-gpu.sh
ExecStart=/usr/local/bin/agent-runner

This ensures the GPU is in a clean state before the agent tries to use it.

Testing the Recovery

Simulate a hard lock by killing the agent process mid-step:

# Run agent in background
./agent-runner &
AGENT_PID=$!
sleep 5  # let it run a few steps
kill -9 $AGENT_PID  # simulate hard lock
# Wait for systemd to restart
sleep 15
# Check logs: agent should recover and continue from last checkpoint

Check Postgres to verify no duplicate steps:

SELECT agent_id, step_number, COUNT(*)
FROM agent_actions
GROUP BY agent_id, step_number
HAVING COUNT(*) > 1;

Should return zero rows.

Performance Considerations

Checkpointing every step to Postgres adds latency—typically 1-5ms per step for a local Postgres. For most agent loops (which involve LLM calls taking seconds), this is negligible. If your steps are sub-millisecond, batch checkpoints every N steps, but accept the risk of losing up to N steps.

Use pgbouncer or connection pooling to avoid connection overhead. Use asyncpg for non-blocking writes if your agent is async.

Alternatives and Why They Fail

  • File-based checkpoints: Not crash-safe. A power loss during write corrupts the file. Postgres's WAL guarantees atomicity.
  • In-memory replication: Still lost on reboot.
  • Checkpointing only on graceful shutdown: Hard locks are not graceful.
  • External orchestration (Kubernetes): Adds complexity; still need durable storage. Postgres is simpler for a single-node setup.

Conclusion

GPU hard locks are inevitable in production. By checkpointing every step to Postgres and using systemd for automatic restart with idempotent replay, you can build agent systems that survive crashes without losing a single step. The pattern is simple, reliable, and uses tools you already know.

Next time your GPU locks, your agents won't skip a beat.

#agent-infrastructure#checkpointing#gpu-failure#recovery#self-healing
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.