Checkpointing at the Action Level: Why We Store Every Agent Step and How It Saves Runs

Stop losing hours of agent work to a single crash. Action-level checkpointing is the safety net your autonomous system needs.

by
Checkpointing at the Action Level: Why We Store Every Agent Step and How It Saves Runs

The Problem with Task-Level Checkpointing

Most agent orchestrators default to checkpointing at task granularity. A task might involve multiple sub-steps; if a sub-step fails, the entire task's progress is lost. The agent re-executes all sub-steps, including expensive LLM calls and API calls. Task-level checkpointing also hides the real error—the root cause may be a malformed output from a previous sub-step, but retrying the whole task obscures it.

Action-Level Checkpointing: A Minimal Definition

An action is the smallest unit of work that changes the agent's state. Examples include an LLM completion (prompt + response), a tool invocation (function name + arguments + result), a conditional branch decision (which path was taken), or a variable assignment. Each action is stored as a row in a checkpoints table with a simple schema:

CREATE TABLE checkpoints (
    run_id      UUID NOT NULL,
    step_index  INT  NOT NULL,
    action_type TEXT NOT NULL,
    input       JSONB NOT NULL,
    output      JSONB,
    created_at  TIMESTAMPTZ DEFAULT now(),
    PRIMARY KEY (run_id, step_index)
);

No fancy graph database or event sourcing framework is needed—just Postgres with a composite primary key. Batch inserts in transactions ensure consistency.

How Recovery Works

When an agent crashes and restarts, it queries the most recent checkpoint for its run. If the last action is completed (output not null), the next action is started. If incomplete, only that action is re-executed. Full input and output are stored so recovery can reconstruct the exact state. For LLM completions, the entire message history is implicit and can be replayed by reading all checkpoints in order.

Why Postgres?

Postgres provides transactional guarantees, point-in-time recovery via WAL archiving, and no extra infrastructure if already used for metadata. Querying the checkpoint table allows analyzing failed runs—e.g., "show all runs where the last action was a tool call returning a 429 status code."

Real-World Impact

A monitoring agent that runs a loop of scraping metrics, detecting outliers, writing summaries, and sending alerts benefits from action-level checkpointing. Before, a crash would lose the current loop and debugging was difficult without per-action logs. After implementation, pinpointing the exact failing API call allowed targeted retry logic, reducing crash frequency. Another agent running a multi-hour research pipeline used to lose everything on a crash; now it resumes from the last completed action, saving hours of compute per incident.

Performance Considerations

Storing each action is cheap. An action record is typically a few kilobytes. Active agents generating many actions per hour produce modest data volumes that Postgres handles easily. Batching writes in groups with a single INSERT keeps write latency low. Reads by primary key are always fast. For large outputs (e.g., LLM completions), compression may be considered but is often unnecessary for typical throughput.

Lessons Learned

Lesson 1: Make actions idempotent. If a tool call creates a resource, replaying the action must not duplicate it. Use idempotency keys—each action carries a unique ID, and the downstream service deduplicates based on that.

Lesson 2: Store the full input, not just a reference. Inlining everything avoids dependencies on external state that may be cleared. Storage is cheap; debugging time is not.

Lesson 3: Prune old checkpoints. Keep checkpoints for a retention period (e.g., 30 days) then archive or delete. Failed runs may be kept longer for debugging, perhaps compressed after a certain period.

Lesson 4: Monitor checkpoint throughput. Track insert and read rates to detect agent stalls or recovery events.

When Not to Use Action-Level Checkpointing

If agents run for seconds and failures are cheap, the overhead of writing a database row per action may not be worth it. Also, actions with side effects that cannot be made idempotent (e.g., launching a rocket) require different patterns like two-phase commit or saga.

For most autonomous agents—research bots, monitoring systems, data pipelines—action-level checkpointing is a net win. The cost is small, and the benefits in resilience and debuggability are significant.

Implementation Guide

A minimal implementation can use a decorator that writes an incomplete checkpoint before executing the action and updates with the output after. Connection pooling and retry logic should be added in production. The core idea is simple: write before, update after.

The Future: Checkpointing as a Service

A generic checkpointing service could support multiple backends and provide a consistent API for checkpointing, recovery, and replay. The goal is to make action-level checkpointing the default.

Until then, the pattern is proven: stop losing runs, start checkpointing every action.

#agent-orchestration#checkpointing#failure-modes#postgres#recovery
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related