Deterministic Replay: Reproducing Agent Failures Inside Postgres Transactions

How to capture and replay agent execution traces using Postgres transaction isolation

by

Deterministic Replay: Reproducing Agent Failures Inside Postgres Transactions

Every autonomous agent developer has been there: a swarm fails in production, you check the logs, and the state is gone. The agent made a decision based on a context window that no longer exists. You can't reproduce it because the LLM call, the tool output, and the exact database state are lost. This is the single biggest obstacle to building reliable agent systems.

The fix is to treat agent execution like a database transaction. Record every input, every output, every state change, and then replay it all inside a Postgres transaction with serializable isolation. If you can replay the exact sequence of events, you can reproduce any failure deterministically.

The Problem: Non-Deterministic Agent Failures

Agents are state machines driven by LLM calls and tool outputs. Each step depends on the previous context, which includes:

  • The system prompt and conversation history
  • Results from tool calls (database queries, API responses, file reads)
  • The internal state of the agent (memory, vector store, etc.)
  • Randomness from LLM sampling (temperature, top_p)

When something goes wrong, you typically see a stack trace or an error message, but the context that led to it is gone. You can't ask the LLM "what were you thinking?" because the exact prompt is lost.

Postgres as the Replay Engine

Postgres is not just a database; it's a transaction log with snapshot isolation. We can use it to record every agent step as a row in a table, and then replay those steps inside a single transaction. Because Postgres guarantees serializable isolation, we can ensure that the replay sees exactly the same data as the original run.

Here's the schema:

CREATE TABLE agent_steps (
    id BIGSERIAL PRIMARY KEY,
    run_id UUID NOT NULL,
    step_number INT NOT NULL,
    step_type TEXT NOT NULL, -- 'llm_call', 'tool_call', 'state_update'
    input JSONB,
    output JSONB,
    created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE INDEX idx_agent_steps_run ON agent_steps(run_id, step_number);

During execution, every agent step writes to this table. The key is that the write happens before the step is applied to the actual system state. This gives us a write-ahead log for the agent.

Recording a Run

Let's say we have an agent that calls an LLM, then executes a tool. The recording looks like this:

import psycopg2
from uuid import uuid4

run_id = uuid4()
conn = psycopg2.connect("dbname=agents")
cur = conn.cursor()

def record_step(step_type, input_data, output_data):
    cur.execute(
        """INSERT INTO agent_steps (run_id, step_number, step_type, input, output)
           VALUES (%s, %s, %s, %s, %s)""",
        (run_id, step_number, step_type, psycopg2.extras.Json(input_data), psycopg2.extras.Json(output_data))
    )
    conn.commit()

Each step is committed immediately. This is important because if the agent crashes, we still have the steps up to the crash point.

Replaying a Run

To replay, we open a new transaction with SERIALIZABLE isolation level. This ensures that all reads within the transaction see a consistent snapshot of the database as of the start of the transaction. We then iterate through the recorded steps and feed them to the agent, but instead of making actual LLM calls or tool executions, we replay the recorded outputs.

from psycopg2.extras import Json

def replay_run(run_id):
    conn = psycopg2.connect("dbname=agents")
    conn.set_session(isolation_level='SERIALIZABLE')
    cur = conn.cursor()
    
    cur.execute(
        """SELECT step_number, step_type, input, output
           FROM agent_steps
           WHERE run_id = %s
           ORDER BY step_number""",
        (run_id,)
    )
    
    for step_number, step_type, input_data, output_data in cur.fetchall():
        if step_type == 'llm_call':
            # Instead of calling the LLM, use the recorded output
            agent.llm_response = output_data
        elif step_type == 'tool_call':
            # Instead of executing the tool, use the recorded output
            agent.tool_result = output_data
        agent.step()

Because the transaction is serializable, any database reads the agent makes during replay will see the same data as the original run (assuming no other writes happened in between). This is the key to determinism.

Handling Non-Determinism in LLM Calls

LLM calls are inherently non-deterministic due to sampling. To replay them, we need to record the exact input and output. But what if the failure is caused by the LLM's output? Then replaying the same output is exactly what we want. However, if we want to debug by trying a different output, we can modify the recorded output before feeding it to the agent.

For full determinism, you can also record the random seed used for the LLM call. Some LLM APIs allow you to set a seed. If you do, record it:

ALTER TABLE agent_steps ADD COLUMN seed BIGINT;

Then during replay, set the same seed.

Real-World Example: Debugging a Tool Loop

I had a swarm where an agent kept calling a search tool with the same query in an infinite loop. The logs showed the loop but not why the agent thought it needed to search again. By replaying the run, I could step through each decision and see that the LLM was ignoring the tool results because of a subtle prompt formatting issue. The replay made it obvious: the tool output was being truncated by a max token limit, so the agent never saw the full result.

Without replay, I would have had to guess. With replay, I fixed the prompt in 10 minutes.

Performance Considerations

Recording every step adds latency. In practice, the overhead is minimal because the write is a simple INSERT. For high-throughput swarms, you can batch writes or use async commits. But for debugging, you don't need to record every run; only record runs that are flagged as interesting (e.g., based on error rates or user reports).

Replay is slower than real-time because you're reading from the database and simulating steps. But that's fine; you're debugging, not running production.

Limitations

  • External dependencies: If the agent calls an external API that changes state (e.g., sends an email), replay won't re-send it. You need to mock those calls during replay.
  • Time-dependent logic: If the agent uses current time, replay will see the recorded time, not the actual time. You can override time functions during replay.
  • Concurrent state changes: If other processes modify the database between the original run and replay, the serializable snapshot may differ. To avoid this, replay in a dedicated environment with no other writes.

Conclusion

Deterministic replay using Postgres transactions is a powerful debugging technique for agent swarms. By recording every step and replaying inside a serializable transaction, you can reproduce any failure exactly. This turns debugging from guesswork into a systematic process. Next time your swarm does something inexplicable, don't just stare at logs — replay it.

Tools used: Postgres 15+, psycopg2 2.9, Python 3.11.

#agent-orchestration#postgres#reliability#swarm-behavior#testing
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.