Deterministic Replay from Structured Logs: Reproducing Agent Failures Without the Swarm

by

When an agent swarm goes off the rails, the first instinct is to replay the scenario. But swarms are non-deterministic by nature: LLM temperature, network jitter, race conditions, and external API latency all conspire to make the exact same sequence of events unrepeatable. Without a time machine, you can't step through the failure in a debugger.

We've adopted a different approach: structured logging combined with a deterministic replay simulator. Instead of trying to reproduce the exact runtime conditions, we log every decision point, every LLM response, every tool call, and every environment observation in a structured format. Then we build a lightweight replay engine that feeds these logs back into the agent's core logic in a controlled manner, effectively simulating the original execution without needing the swarm, the network, or even the LLM.

This post walks through the architecture, the log schema, and the replay loop, with concrete code snippets that you can adapt to your own agent framework.

Why Deterministic Replay Matters

Debugging agent failures is fundamentally different from debugging a traditional program. In a typical application, the same input always produces the same output (ignoring concurrency). But agents are stateful, reactive, and often call external services. A failure might depend on:

  • A specific LLM response that is non-deterministic due to temperature or model version drift.
  • A race condition between two agents reading from the same queue.
  • A transient network timeout that only happens under load.
  • An environment change that is not under your control.

When you try to reproduce the failure live, you get a different execution path. Deterministic replay cuts through this by capturing the minimal set of inputs that drive the agent's decision loop, and then replaying them in a sandboxed simulator.

The Structured Log Schema

Every agent in our swarm writes to a shared log stream (over UDP to avoid blocking the agent). Each log entry is a JSON object with a timestamp, a unique request ID, an agent ID, an event type, and a payload. We define a small set of event types that cover all sources of non-determinism:

  • llm_request: the prompt sent to the LLM, including system prompt, conversation history, and tool definitions.
  • llm_response: the raw response from the LLM, including finish reason and token usage.
  • tool_request: the tool name and arguments before execution.
  • tool_response: the tool's output (or error) after execution.
  • env_observation: any sensor reading or external state fetch (e.g., database query result, file contents).
  • agent_decision: the agent's internal state transition (e.g., which branch it took in a conditional).

Here's an example llm_request entry:

{
  "timestamp": "2025-03-15T10:23:45.123Z",
  "request_id": "req-abc123",
  "agent_id": "planner-01",
  "event_type": "llm_request",
  "payload": {
    "model": "qwen2.5:32b",
    "temperature": 0.1,
    "messages": [
      {"role": "system", "content": "You are a planner..."},
      {"role": "user", "content": "Deploy the new service..."}
    ],
    "tools": [
      {
        "name": "run_shell",
        "description": "Execute a shell command...",
        "parameters": {"type": "object", "properties": {...}}
      }
    ]
  }
}

The key is that the payload must be sufficient to exactly reproduce the agent's next action. For llm_request, that means the full prompt (including dynamic context). For tool_response, that means the exact return value. For env_observation, that means the raw data as seen by the agent.

Building the Replay Simulator

The replay simulator is a standalone process that reads a log file (or a stream of logs) and replays them through the same agent logic, but with all external calls replaced by mock functions that return the logged values. The simulator does not call the LLM, does not execute shell commands, and does not query databases. Instead, it matches each incoming log entry to the next expected event and feeds the recorded response back to the agent.

Here's a simplified Python implementation of the replay loop:

import json
from collections import defaultdict

class ReplaySimulator:
    def __init__(self, log_path):
        with open(log_path) as f:
            self.logs = [json.loads(line) for line in f]
        self.index = 0
        self.pending = defaultdict(list)  # request_id -> list of events

    def next_event(self, request_id):
        """Return the next event for a given request_id, or block."""
        while self.index < len(self.logs):
            entry = self.logs[self.index]
            self.index += 1
            if entry["request_id"] == request_id:
                return entry
            # Events for other request_ids are queued for later
            self.pending[entry["request_id"]].append(entry)
        # Check queue
        if self.pending[request_id]:
            return self.pending[request_id].pop(0)
        return None

    def replay(self, agent_factory):
        agent = agent_factory()
        agent.register_llm_handler(self.mock_llm)
        agent.register_tool_handler(self.mock_tool)
        agent.register_env_handler(self.mock_env)
        # Start the agent's main loop
        agent.run()

    def mock_llm(self, request_id, prompt):
        # Wait for the matching llm_response event
        while True:
            event = self.next_event(request_id)
            if event and event["event_type"] == "llm_response":
                return event["payload"]
            # Log any other events that arrive out of order
            self.stash(event)

    def mock_tool(self, request_id, tool_name, args):
        # Similar: wait for tool_response event
        ...

    def mock_env(self, request_id, observation_key):
        # Wait for env_observation event
        ...

The simulator uses a single-threaded, blocking approach: it advances through the log file sequentially, but it can interleave events for different request IDs by queuing them. This preserves the exact ordering of events as they happened in the original swarm, including any interleaving between agents.

Handling Non-Deterministic Internal State

Some agent frameworks use random number generators for exploration (e.g., epsilon-greedy in RL agents) or for sampling from a distribution. To achieve full determinism, you must also seed the RNG with a value that was logged during the original run. We add an rng_seed field to the agent's initialization log entry, and the replay simulator sets the RNG seed before calling the agent's main loop.

Similarly, any source of randomness in the agent's decision logic (e.g., shuffling a list) must be replaced with a deterministic version that reads from a logged sequence of random values. This is a bit more invasive but usually only affects a few places in the codebase.

Validating Replay Fidelity

How do you know the replay is faithful? The ultimate test is that the agent produces the exact same sequence of decisions and tool calls as the original run. We compare the replayed log (generated by the simulator) with the original log. If they match exactly, the replay is correct. In practice, small differences can arise due to floating-point rounding or unlogged internal state. We tolerate minor numeric differences but flag any divergence in control flow.

We also run a suite of unit tests that replay known failure scenarios and assert that the agent reaches the same final state (e.g., a specific error message or a deadlock). Over time, this becomes a regression test suite for the agent's core logic.

Limitations and Trade-offs

Deterministic replay is not a silver bullet. Some limitations:

  • Log volume: Every LLM request and response must be logged, which can be large (especially for long conversations). We use a circular buffer in memory and flush to disk only on failure, reducing overhead.
  • External state: If the agent mutates external state (e.g., writes to a database), replaying with the same inputs may not produce the same outputs because the database state has changed. We handle this by snapshotting the relevant external state at the start of the replay (e.g., a DB dump) and restoring it.
  • Timing-dependent bugs: If the failure depends on real-time delays (e.g., a timeout that triggers only after 5 seconds), the replay must simulate time. We add a clock abstraction that can be advanced manually during replay, but this adds complexity.
  • Non-reproducible hardware: Bugs that depend on hardware characteristics (e.g., GPU kernel launches) cannot be captured by logging alone. In practice, these are rare in agent logic.

Despite these limitations, deterministic replay has saved us countless hours of debugging. It turns an unreproducible heisenbug into a reproducible, step-through-able failure.

Putting It All Together

Here's the workflow we follow:

  1. Instrument the agent framework to emit structured logs for every source of non-determinism.
  2. Run the swarm normally, logging to a file or a central log server.
  3. When a failure occurs, grab the log file for the relevant time window.
  4. Run the replay simulator with the log file and the same agent code (but with mocked externals).
  5. Step through the replay using a debugger or by inserting breakpoints in the mock functions.
  6. Fix the bug, then add the log file to the regression test suite.

The replay simulator itself is stateless and can be run in CI on every commit, ensuring that past failures never regress.

Conclusion

Deterministic replay from structured logs is a powerful technique for debugging agent swarms. By capturing the minimal set of inputs that drive the agent's behavior, you can reproduce failures offline without the swarm, the network, or the LLM. The approach requires careful instrumentation and a dedicated replay simulator, but the payoff is a debuggable, testable, and reproducible agent system.

We've used this technique to fix race conditions, prompt injection vulnerabilities, and infinite loops that were impossible to reproduce live. It's now a standard part of our development workflow.

Next time, we'll discuss how to extend the replay simulator to support differential testing between agent versions.

#agent#debugging#determinism#replay#simulation#structured-logging#testing
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.