State Loss in Agent Swarms: Why We Checkpoint Every Action and How It Saved a 3-Hour Run

Every uncheckpointed action is a gamble; here's how we made recovery the default.

by
State Loss in Agent Swarms: Why We Checkpoint Every Action and How It Saved a 3-Hour Run

Agent swarms are notoriously fragile. You launch a multi-hour pipeline—dozens of agents, each making calls to LLMs, vector stores, and external APIs—and then one node hiccups. The whole run is gone. This article is about why we now checkpoint every single action, how we built it with Postgres and pgvector, and the principles that guided the design.

The Problem: State Loss in Long-Running Swarms

A swarm run can involve hundreds of actions per agent. If you only save state at the end of an agent’s lifecycle, a crash mid-run destroys all intermediate work. The cost is not just lost compute—it’s lost time, lost context, and lost trust in the system. The root cause is insufficient granularity in checkpointing.

The Architecture of Checkpointing

Our solution is straightforward: every action an agent takes is recorded as a row in a Postgres table. An action is any atomic unit of work—calling an LLM, querying a vector store, writing to a file, sending a message to another agent. Each action has a unique ID, a timestamp, the agent ID, the action type, input parameters, output result, and a status (pending, running, completed, failed).

We use Postgres with pgvector for storing embeddings of action inputs/outputs, enabling similarity search across past actions. This is not just for recovery—it also allows agents to reuse previous results, saving cost and time.

The Table Schema

CREATE TABLE swarm_actions (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    agent_id TEXT NOT NULL,
    action_type TEXT NOT NULL,  -- e.g., 'llm_call', 'vector_search', 'file_write'
    input_params JSONB,
    output_result JSONB,
    embedding VECTOR(768),  -- from an embedding model
    status TEXT DEFAULT 'pending',
    created_at TIMESTAMPTZ DEFAULT now(),
    completed_at TIMESTAMPTZ
);

CREATE INDEX idx_agent_status ON swarm_actions (agent_id, status);
CREATE INDEX idx_embedding ON swarm_actions USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);

Every action is inserted with status = 'pending' before execution. Once the action completes, we update the row with the output and set status to completed. If it fails, we set failed and store the error.

Recovery Flow

When the orchestrator restarts, it queries all actions with status = 'running' or status = 'pending' for each agent. Those are the actions that were in flight at the time of crash. For each, it can either retry (if idempotent) or skip (if the action was actually completed but the status wasn't updated). We use a simple rule: if the action is idempotent and the input is identical, we retry. Otherwise, we mark it as failed and notify the agent.

This recovery happens in seconds, not hours. The agents pick up where they left off, using the checkpointed outputs of completed actions as context.

Checkpointing Every Action: The Trade-Offs

You might worry about overhead. Writing a row to Postgres for every LLM call or vector search adds latency. In practice, with a local Postgres instance and batched inserts, the overhead is small relative to the action’s own duration. For fast actions like local embedding lookups, batching mitigates the cost.

Storage is also manageable. Each row is a few KB (JSONB for inputs/outputs, plus the embedding vector). Archiving old runs to S3-compatible storage after a week keeps the table lean.

Lessons Learned

  1. Checkpoint at the action level, not the agent level. An agent’s lifecycle can include dozens of actions. If you only save state when the agent finishes, you lose everything if it crashes mid-way.
  2. Use Postgres as the control plane. It’s battle-tested, supports JSONB, and with pgvector you get similarity search for free. Use pgbouncer for connection pooling to handle many concurrent connections.
  3. Idempotency is key. Design actions so that retrying them produces the same result. For LLM calls, use deterministic parameters (temperature=0, same seed). For file writes, use content-hash filenames.
  4. Recovery must be fast. Index on (agent_id, status) to quickly find incomplete actions.
  5. Embeddings enable reuse. Storing embeddings of action inputs allows agents to find previous similar actions and reuse results, reducing LLM costs.

Implementation Details

We use a lightweight Python library (not yet open-sourced) that wraps any async function with checkpointing. The decorator looks like:

@checkpoint_action(action_type="llm_call")
async def call_llm(prompt: str, model: str = "gpt-4") -> str:
    # actual LLM call
    return response

Under the hood, it generates a unique action ID, inserts a pending row, executes the function, and updates the row. If the process crashes before the update, the action remains running and is retried on recovery.

For the orchestrator, we have a simple systemd service that runs the main loop. On restart, it executes a recovery routine before dispatching new tasks. The recovery routine queries swarm_actions for the current run ID and rebuilds the agent state from completed actions.

The Bigger Picture

State loss is the silent killer of agent swarms. When you scale from a demo to a production pipeline handling real data, crashes are inevitable. Network blips, OOM kills, power failures—they all happen. If you don’t checkpoint every action, you’re gambling with compute time and user trust.

Our approach is not novel—it’s the same principle used in databases (WAL) and stream processing (Kafka offsets). But applying it to agent swarms with Postgres and pgvector is practical and effective.

What’s Next

We’re working on checkpointing the internal state of agents (their conversation history, tool call stacks) using the same action table. This will allow full snapshot-and-restore of an agent’s context. We’re also exploring using Qdrant as an alternative vector store for even faster similarity search across actions.

But for now, the simple Postgres table is our workhorse. It’s reliable and efficient.

#agent-swarms#orchestration#pgvector#postgres#recovery#state
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related