Partial Rollback vs Full Replay: Recovering from Poisoned Agent State Without Data Loss

by

When an autonomous agent operates over a long horizon, its internal state accumulates decisions, observations, and learned parameters. A single poisoned action — a hallucinated tool result, a corrupted embedding, a misapplied LoRA update — can cascade, corrupting downstream state. Recovering without losing legitimate progress is a hard engineering problem. Two dominant strategies exist: partial rollback and full replay. Neither is universally superior; each has a place depending on state granularity, dependency structure, and acceptable downtime.

The Shape of Agent State

Before choosing a recovery strategy, we must define what "state" means in an agentic system. Typically, state falls into three categories:

  • Ephemeral context: conversation history, intermediate tool outputs, current plan steps. Often stored in-memory or in a short-lived cache.
  • Persistent knowledge: embeddings in a vector store, graph nodes in a knowledge graph, fine-tuned LoRA adapters on a base model.
  • Control state: agent loop counters, retry budgets, routing decisions, and flags that influence future behavior.

A poisoned state can originate from any of these. For example, a tool returning incorrect data might poison the ephemeral context, which then gets embedded and stored persistently. Alternatively, a faulty LoRA update might corrupt the model's behavior for all subsequent inferences.

Partial Rollback: Undo Selective Changes

Partial rollback aims to revert only the components that were directly affected by the poison, leaving the rest of the state intact. This is analogous to a database savepoint or a git revert of a single commit.

How It Works

To perform a partial rollback, the system must maintain a journal of state mutations. Each mutation is tagged with a causal dependency — which action or input triggered it. When a poison is detected, the system walks backward through the journal, identifying mutations that depend (transitively) on the poisoned source. Only those mutations are undone.

class StateJournal:
    def __init__(self):
        self.entries = []  # list of (timestamp, mutation_fn, undo_fn, dependency_id)
        self.dependency_graph = {}

    def record(self, mutation, undo, depends_on=None):
        entry_id = len(self.entries)
        self.entries.append((entry_id, mutation, undo, depends_on))
        if depends_on:
            self.dependency_graph.setdefault(depends_on, []).append(entry_id)

    def rollback_to(self, checkpoint_id):
        # Collect all entries after checkpoint that depend on poisoned path
        to_undo = set()
        def collect_dependents(eid):
            for dep in self.dependency_graph.get(eid, []):
                if dep not in to_undo:
                    to_undo.add(dep)
                    collect_dependents(dep)
        for eid in range(checkpoint_id + 1, len(self.entries)):
            if eid not in to_undo:
                continue
            # Apply undo in reverse order
        for eid in reversed(sorted(to_undo)):
            self.entries[eid][2]()  # call undo function
        # Truncate journal
        self.entries = self.entries[:checkpoint_id + 1]

When to Use

Partial rollback shines when the poison is localized and the dependency graph is a DAG (no cycles). It minimizes data loss because only tainted mutations are discarded. It is fast — no need to reprocess legitimate actions.

Risks

  • Hidden dependencies: If the dependency tracking is incomplete, some tainted state may survive the rollback, leaving the agent still poisoned.
  • Stateful side effects: If a mutation had irreversible side effects (e.g., sent an email, wrote to an external API), rolling back internally doesn't undo the real-world impact.
  • Clock skew: Timestamp-based rollback can be tricky in distributed systems; logical clocks or vector clocks are safer.

Full Replay: Start Over from a Known Good Snapshot

Full replay discards all state after a known good checkpoint and re-executes every action from that point forward. This is the nuclear option, but it guarantees consistency.

How It Works

The system periodically snapshots the entire agent state — including model weights (or LoRA adapters), vector store contents, and control state. On poison detection, it restores the latest clean snapshot and replays the action log (a sequence of inputs and tool calls) up to the current point, but with the poison removed.

class ReplayManager:
    def __init__(self, snapshot_store, action_log):
        self.snapshot_store = snapshot_store
        self.action_log = action_log  # list of (action_id, input, timestamp)

    def recover(self, poison_action_id):
        # Find the snapshot taken before poison_action_id
        snapshot = self.snapshot_store.get_snapshot_before(poison_action_id)
        self.restore_state(snapshot)
        # Replay actions from after snapshot up to poison_action_id - 1
        for action in self.action_log:
            if action.action_id >= poison_action_id:
                break
            if action.action_id < snapshot.action_id:
                continue
            self.execute(action.input)
        # Skip the poisoned action and continue with subsequent actions
        for action in self.action_log:
            if action.action_id <= poison_action_id:
                continue
            self.execute(action.input)

When to Use

Full replay is appropriate when:

  • The poison is widespread or the dependency graph is cyclic.
  • The cost of missed data loss is higher than the cost of reprocessing.
  • The action log is deterministic and idempotent.

Risks

  • Amplified latency: Replaying hours of actions can take significant time, especially if actions involve expensive LLM calls or external API calls.
  • Non-idempotent actions: If an action cannot be safely re-executed (e.g., a payment charge), replay may cause duplicates.
  • Snapshot storage overhead: Frequent snapshots consume space; infrequent snapshots increase replay length.

Hybrid Approaches

In practice, a hybrid strategy often works best. For example:

  1. Journal-based rollback for ephemeral context (fast, low risk).
  2. Snapshot + replay for persistent knowledge (slow but safe).
  3. Selective replay: Only re-run actions that are known to be deterministic; for non-deterministic or side-effecting actions, use a compensating action instead of replay.

Another pattern is to maintain multiple independent state partitions. If a partition is poisoned, only that partition is rolled back/replayed, while others continue unaffected. This requires careful design of state boundaries.

Practical Considerations

Detecting Poison

Recovery is useless if you don't detect poison early. Common detection methods:

  • Anomaly detection on embeddings: Sudden shifts in embedding norms or distances.
  • Consistency checks: Validate tool outputs against expected schemas or historical patterns.
  • Human-in-the-loop: For critical decisions, require human approval before state mutation is committed.

Testing Recovery Procedures

Regularly test both rollback and replay paths in staging. A common failure is that the undo function itself has a bug, or the replay logic doesn't handle edge cases (e.g., actions that depend on wall-clock time).

Tradeoff Summary

Aspect Partial Rollback Full Replay
Data loss Minimal (only tainted mutations) Potentially large (all post-snapshot state)
Recovery speed Fast (seconds to minutes) Slow (minutes to hours)
Implementation complexity High (dependency tracking) Moderate (snapshot + replay)
Safety guarantee Low (hidden dependencies) High (full consistency)

Conclusion

There is no silver bullet for poisoned agent state. Partial rollback is elegant when dependencies are well-understood; full replay is brute-force reliable. The right choice depends on your system's state model, latency tolerance, and tolerance for data loss. In practice, design for recoverability from day one: journal mutations, take periodic snapshots, and test your recovery paths as rigorously as your primary paths. The cost of a bad recovery is often higher than the cost of the original bug.

#agent#data-loss#poisoned#recovery#replay#rollback#state
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related