The Poisoned Agent Dilemma: Partial Rollback vs. Full Replay After Corrupted State
How to recover an autonomous agent when its internal state gets corrupted without rebuilding from scratch.
Every autonomous agent operates on an evolving state — a snapshot of memory, tool outputs, intermediate inferences, and accumulated context. That state can become poisoned: a hallucinated fact gets embedded into long-term memory, a vector store index gets corrupted by a bad embedding, or a sub-agent returns malformed JSON that cascades into downstream decisions. Once you detect the corruption, you face a dilemma: do you surgically remove the poisoned parts and continue (partial rollback), or do you rewind the entire agent to a known-good checkpoint and replay the task from there (full replay)? Both approaches have deep engineering trade-offs that depend on state granularity, task length, and the cost of recomputation.
The Anatomy of Agent State
To understand recovery strategies, we must first decompose what "agent state" actually means in a typical stack. A production agent running on a stack like RiNET — Qwen on vLLM for inference, BGE-M3 for embeddings, Postgres for relational logs, Qdrant for vector memory, Neo4j for knowledge graph state — holds state across several layers:
- Conversation history: raw message turns, tool calls, and responses. Usually append-only, but can be truncated.
- Short-term working memory: a sliding window of recent context, often held in-memory within the agent loop.
- Long-term memory: vector embeddings stored in Qdrant, representing facts or summaries that persist across sessions.
- Knowledge graph state: nodes and edges in Neo4j that encode relationships extracted during the task.
- Tool call caches: results of expensive tool invocations (e.g., API responses, database queries) that may be reused.
- Agent execution graph: the DAG of sub-tasks, decisions, and branches that the agent has traversed.
Corruption can strike any of these. A poisoned long-term memory entry might cause the agent to retrieve a wrong fact on every future query. A corrupted knowledge graph edge might misdirect the reasoning path. A bad tool call cache might return stale or wrong data.
Partial Rollback: Surgical Correction
Partial rollback attempts to remove or correct only the corrupted substates while preserving the rest of the agent's progress. The key challenge is identifying exactly which state elements are tainted and ensuring that downstream dependencies are also invalidated.
Implementation Pattern
A common pattern is to maintain a provenance graph alongside the agent state. Each state element carries a list of its ancestors — the inputs, tool calls, and reasoning steps that produced it. When corruption is detected (e.g., a fact fails a consistency check), you walk the provenance graph backward to mark all derived elements as invalid. Then you either recompute only those elements or delete them and let the agent re-derive them lazily.
class ProvenanceNode:
def __init__(self, id, data, parents=None):
self.id = id
self.data = data
self.parents = parents or []
self.children = []
self.valid = True
def invalidate(self):
self.valid = False
for child in self.children:
child.invalidate()This works well when the state graph is acyclic and the corruption is localized. For example, if a single vector embedding is found to be corrupted (e.g., its distance to a known centroid is an outlier), you can delete that embedding from Qdrant and invalidate any knowledge graph nodes that depended on it. The agent's next retrieval will simply miss that entry and may re-ingest the source data.
Pitfalls
Partial rollback fails when corruption is not localized. A subtle bias in the agent's reasoning, such as consistently favoring a certain action due to a poisoned training example in the context window, may not be traceable to a single state element. The provenance graph only tracks explicit dependencies, not implicit influences like prompt ordering or token-level biases. Additionally, if the state graph has cycles (e.g., an agent that mutates its own memory), invalidation can loop or miss indirect effects.
Another practical issue: in a system with nightly LoRA fine-tuning (as in the RiNET stack), a corrupted state might have already been used as training data. Rolling back the agent state won't undo the fine-tuned weights. You would need to revert the LoRA adapter to a previous checkpoint, which is a form of rollback at the model level.
Full Replay: Clean Slate
Full replay discards the entire agent state and restarts the task from a known-good checkpoint. The agent re-executes all steps, potentially using cached results for idempotent tool calls. This is the nuclear option: it guarantees no residual corruption, but at the cost of recomputing everything.
Checkpointing Strategy
To make full replay feasible, you need periodic checkpoints of the entire agent state. A checkpoint includes the conversation history up to that point, the contents of all memory stores (or at least snapshots), and the execution graph. Checkpointing every N turns or every M minutes creates a trade-off between recovery granularity and storage overhead.
import pickle
from datetime import datetime
def save_checkpoint(agent_state, checkpoint_dir):
timestamp = datetime.utcnow().isoformat()
path = f"{checkpoint_dir}/checkpoint_{timestamp}.pkl"
with open(path, 'wb') as f:
pickle.dump(agent_state, f)
return path
def load_checkpoint(path):
with open(path, 'rb') as f:
return pickle.load(f)In practice, serializing large vector stores or knowledge graphs can be expensive. Many teams instead use append-only logs and replay the log from the checkpoint onward. For example, you can store every tool call and its result in Postgres, then replay them in order, skipping those that are idempotent.
Deterministic Replay
Full replay assumes that the agent's behavior is deterministic given the same inputs and state. This is rarely true in practice: LLM outputs are stochastic, tool calls may have side effects, and external APIs may return different results over time. To achieve deterministic replay, you must log the random seeds for each LLM call and re-seed the RNG during replay. For external tool calls, you can cache all responses keyed by (tool_name, arguments, timestamp) and replay the cached response instead of re-invoking the tool.
import random
class DeterministicReplayAgent:
def __init__(self, seed):
self.rng = random.Random(seed)
self.tool_cache = {}
def llm_call(self, prompt):
seed = self.rng.randint(0, 2**32)
# In replay, use the cached seed and response
return cached_llm_response(prompt, seed)
def tool_call(self, tool, args):
key = (tool, tuple(sorted(args.items())))
if key in self.tool_cache:
return self.tool_cache[key]
response = actual_tool_call(tool, args)
self.tool_cache[key] = response
return responseEven with caching, some tools are inherently non-idempotent (e.g., sending an email, creating a database record). For those, you must either skip them during replay or use a sandbox that intercepts and logs side effects.
When to Use Which
Choosing between partial rollback and full replay depends on three factors: the cost of recomputation, the precision of corruption localization, and the risk of latent corruption.
Use partial rollback when: the corruption is clearly isolated to a few state elements, the provenance graph is acyclic and well-maintained, and recomputing those elements is cheap. Also prefer it when the task has taken many steps and the cost of replaying from scratch is high (e.g., a multi-hour data pipeline).
Use full replay when: the corruption is diffuse or its source is unknown, the agent state has cycles or implicit dependencies, or the cost of a small residual corruption outweighs the cost of recomputation. Full replay is also simpler to implement and debug.
Hybrid approach: Some systems implement a tiered recovery. First attempt partial rollback. If the agent's behavior still appears anomalous after a few steps (e.g., confidence scores drop or consistency checks fail), escalate to full replay from the last checkpoint.
Forensic Analysis After Recovery
Regardless of the recovery method, the corrupted state should not be discarded without analysis. The root cause of the corruption — whether a model hallucination, a tool returning malformed data, or a bug in the agent's reasoning loop — must be identified to prevent recurrence. This is where forensic logging comes in. Every state mutation should be logged with a timestamp, source component, and a hash of the input that produced it. After recovery, you can replay the corruption path in a sandboxed environment to reproduce the issue.
import hashlib
import json
def log_mutation(state_id, mutation, source):
record = {
'state_id': state_id,
'mutation': mutation,
'source': source,
'hash': hashlib.sha256(json.dumps(mutation).encode()).hexdigest(),
'timestamp': datetime.utcnow().isoformat()
}
append_to_forensic_log(record)This forensic log becomes invaluable for debugging and for improving the agent's resilience. For example, if a pattern emerges where a particular tool's response format causes parsing errors, you can add input validation or a retry with a different prompt.
Conclusion
The poisoned agent dilemma has no universal answer. Partial rollback offers efficiency but risks leaving invisible corruption. Full replay offers safety but wastes compute. The right choice depends on your agent's state architecture, the cost of failure, and the quality of your provenance tracking. Invest in checkpointing, provenance graphs, and deterministic replay infrastructure before you need them — because when corruption strikes, you won't have time to build them.
This article is part of a series on agent resilience. Previous: "Checkpointing Strategies for Long-Running Agents". Next: "Sandboxing Tool Calls to Prevent State Contamination".