Checkpointing Agent Conversations to Postgres: The Serialization Format That Survived a GPU Crash

by
Checkpointing Agent Conversations to Postgres: The Serialization Format That Survived a GPU Crash

Checkpointing Agent Conversations to Postgres: The Serialization Format That Survived a GPU Crash

You're running a swarm of autonomous agents. Each agent has a multi-turn conversation with tools, memory, and state. Then your GPU crashes—power spike, driver panic, or just a thermal event. The inference server goes down. When it comes back, every agent conversation is gone. That's not just an annoyance; it's data loss. In a self-hosted, sovereign setup, you don't get to blame a cloud provider. You own the failure.

This article is about a serialization format that survived exactly that scenario. It's not clever. It's not novel. It's the minimum viable schema that, when checkpointed to Postgres after every turn, let us restore a 47-turn agent conversation after a GPU crash with zero context loss. The format handles tool calls, embeddings, and state transitions. It's built on Postgres 16, pgvector 0.7.0, and a few conventions.

Why Postgres?

Postgres is the control plane for everything. If your agents are stateless, you don't need this. But agents are stateful by nature: they carry conversation history, tool outputs, and internal state. If you're already using Postgres for metadata, user data, and job queues, adding checkpointing is natural. You get ACID, point-in-time recovery, and no extra infrastructure.

We use Postgres 16 with pgvector for embedding storage, and pgbouncer for connection pooling. The checkpoint table is separate from the vector store, but both live in the same database. That's intentional: a transaction can atomically update both state and embeddings.

The Serialization Format

The goal is to store every turn of an agent conversation as a single row, but with enough structure to reconstruct the full context. After experimenting with JSON blobs and flat tables, we settled on a normalized schema that balances queryability and write speed.

Table: agent_checkpoints

CREATE TABLE agent_checkpoints (
    id BIGSERIAL PRIMARY KEY,
    conversation_id UUID NOT NULL,
    turn_index INT NOT NULL,
    agent_id TEXT NOT NULL,
    parent_turn_id BIGINT REFERENCES agent_checkpoints(id),
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    messages JSONB NOT NULL,
    tool_calls JSONB,
    tool_results JSONB,
    state JSONB,
    embedding vector(768),
    UNIQUE(conversation_id, turn_index)
);

CREATE INDEX idx_checkpoints_conversation ON agent_checkpoints(conversation_id, turn_index);
CREATE INDEX idx_checkpoints_agent ON agent_checkpoints(agent_id);

Key decisions:

  • conversation_id is a UUID generated once per conversation. It's the top-level grouping key.
  • turn_index is a monotonically increasing integer per conversation. It allows ordering without timestamps.
  • parent_turn_id creates a linked list for branching. If an agent forks (e.g., retries a tool call), the new turn references the parent. This is optional but useful for debugging.
  • messages is a JSONB array of message objects. Each message has role, content, and optional metadata. The content can be text or a structured object (e.g., tool output).
  • tool_calls and tool_results are separate JSONB arrays. This avoids mixing tool call structures with user messages.
  • state is a JSONB object for any agent-specific state (e.g., remaining steps, variables, flags).
  • embedding is a pgvector column for the conversation embedding at that turn. Useful for similarity search across conversations.

Message Format

Each message in the messages array follows this structure:

{
  "role": "user" | "assistant" | "system" | "tool",
  "content": "string or object",
  "metadata": {
    "tool_name": "search_web",
    "tool_call_id": "call_abc123",
    "timestamp": "2025-03-21T12:34:56Z"
  }
}

For tool calls, the assistant message includes a tool_calls array in the metadata. For tool results, the role is "tool" and the content is the result.

State Serialization

The state column stores the agent's internal state as a JSONB object. This includes:

  • step_count: number of steps taken so far
  • max_steps: maximum allowed steps
  • current_tool: name of the tool being executed (if any)
  • variables: a map of variable name to value (e.g., "query": "latest news")
  • flags: boolean flags like is_final, requires_human_input

We keep state minimal. Anything that can be recomputed from messages is not stored. The state is a snapshot at the end of the turn.

Checkpointing Flow

Every turn, the agent writes a checkpoint. The flow is:

  1. Agent receives input (user message or tool result).
  2. Agent processes and generates a response (may include tool calls).
  3. Before executing any tool, agent writes a checkpoint with the current messages, tool_calls (if any), and state.
  4. After tool execution, agent writes another checkpoint with the tool result appended to messages and updated state.
  5. If the agent crashes mid-turn, the last checkpoint is loaded and the turn is replayed from the last safe point.

This means each turn may produce multiple checkpoints if there are tool calls. The turn_index is incremented only at the start of a new turn, not per sub-step. Sub-steps are distinguished by a sub_step field in the state.

Surviving the GPU Crash

Here's what happened: our inference server (vLLM 0.6.0 on a single A100) crashed due to a GPU driver issue. At that moment, 12 agent conversations were in progress. When the server restarted, we loaded the latest checkpoint for each conversation from Postgres. The agents resumed exactly where they left off—including the last message that was sent before the crash.

How? Because the checkpoint is written before the inference call that might crash. The sequence is:

  1. Write checkpoint (Postgres transaction commits).
  2. Send inference request to vLLM.
  3. If vLLM crashes, the checkpoint is already saved.
  4. On restart, load the checkpoint. The agent sees the last user message and regenerates the assistant response.

This means we never lose a user message. The worst case is a duplicate assistant response, which is handled by idempotency keys on the client side.

Performance Considerations

Writing a checkpoint per turn adds latency. In our tests, a single checkpoint write (including the embedding calculation) takes about 50ms. For a 10-turn conversation, that's 500ms of overhead. Acceptable for most use cases.

We batch embedding calculations: the embedding is computed asynchronously after the checkpoint is written, using a separate worker that reads from a queue. This avoids blocking the agent.

For high-throughput scenarios, consider using pg_stat_statements to monitor write performance. We also use INSERT ... ON CONFLICT to handle retries safely.

Restoring a Conversation

To restore a conversation after a crash, we query the last checkpoint:

SELECT * FROM agent_checkpoints
WHERE conversation_id = $1
ORDER BY turn_index DESC, created_at DESC
LIMIT 1;

This returns the most recent checkpoint. We then reconstruct the agent's state from the state column and replay the messages array into the agent's context window.

If the conversation branched (parent_turn_id is not null), we walk the linked list to get the full history. But for linear conversations, the messages array already contains the entire history up to that point.

Trade-offs

  • Storage: Each checkpoint duplicates the entire message history up to that point. For long conversations, this can be significant. We mitigate by archiving old checkpoints after a conversation ends. For a 50-turn conversation, storage is about 500KB per checkpoint, so 25MB total. Manageable.
  • Consistency: The checkpoint is eventually consistent with the agent's internal state if the embedding is computed asynchronously. But the critical path (messages and state) is synchronous.
  • Complexity: This schema is more complex than a simple JSON blob. But it allows querying by turn index, agent, and embedding.

Alternatives Considered

  • Flat key-value store: Redis or etcd. Faster writes but no ACID across multiple keys. Postgres gives us atomicity.
  • Event sourcing: Store every event separately. More flexible but more complex to reconstruct state. Our format is a snapshot-per-turn, which is simpler.
  • In-memory only: Fast but crashes lose everything. Not acceptable for sovereign infrastructure.

Conclusion

This serialization format is not revolutionary. It's a pragmatic choice that worked in production. Postgres as the control plane for agent state is a natural fit when you're already self-hosting. The schema survived a GPU crash, and that's the only test that matters.

If you're building agent swarms, don't overlook checkpointing. It's not glamorous, but it's the difference between a resilient system and a fragile demo. Start with this schema, adapt it to your tool calls and state, and test it with a simulated crash. Your future self—and your users—will thank you.

#agent-swarms#failover#postgres#serialization#state
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related