The 3-Signal Rule: Why We Log Every Agent Action Three Ways Before Trusting It
A single log line is a liar. When an autonomous agent takes an action—calls a tool, writes a file, or sends a message—you get one line in a file. If something goes wrong, that line is all you have. And it's never enough.
We built our agent swarm on a simple premise: before we trust any agent output, we must have three independent signals that agree. We call this the 3-signal rule. It's not about redundancy for the sake of it. It's about creating enough evidence to reconstruct causality when a swarm of agents interacts in ways you didn't anticipate.
The Anatomy of a Single Signal
Most logging systems capture a single structured event: timestamp, agent ID, action type, payload. That's signal one. It tells you what happened, but not why or how it relates to everything else.
Consider a typical log line:
{
"timestamp": "2025-03-21T14:32:10.123Z",
"agent": "planner-7",
"action": "call_tool",
"tool": "vector_search",
"query": "latest Qwen fine-tuning results",
"result_count": 3
}This is useful. But if the planner later makes a bad decision, this line alone can't tell you whether the search returned wrong results, the agent misinterpreted them, or something else entirely. You need context. You need the other two signals.
Signal Two: The Vector Embedding Trail
Every agent action that involves semantic content—queries, responses, tool outputs—gets embedded into a vector space using BGE-M3. These embeddings are stored in Qdrant with the same event ID as the structured log. This gives us a second view: instead of reading a timestamp, you can search for semantically similar events.
Here's the pattern:
# Pseudocode for logging signal two
event_id = generate_uuid()
structured_log(event_id, agent_id, action, payload)
# Also embed and store
embedding = bge_m3.encode(payload["query"])
qdrant_client.upsert(
collection="agent_traces",
points=[{
"id": event_id,
"vector": embedding,
"payload": {
"agent": agent_id,
"action": action,
"timestamp": now()
}
}]
)Why does this matter? When debugging, you often don't know which event caused a failure. You have a symptom—an agent produced a strange output. You can embed that output and search for similar past events. If you find three events with similar embeddings but different outcomes, you can trace the divergence. The vector trail lets you ask "what else looked like this?"
Signal Three: The Causal Trace
The third signal is the hardest to generate but the most valuable: a causal trace that records why the agent chose this action. This isn't a stack trace. It's a graph of dependencies: which prior events influenced this decision.
We store causal traces in Neo4j. Each node is an event (log entry), and edges represent influence: "agent-7 used output from tool call X to decide action Y." The trace is built incrementally as the agent runs:
// Example: recording a causal edge
MATCH (prev:Event {id: $previous_event_id})
MATCH (current:Event {id: $current_event_id})
CREATE (prev)-[:INFLUENCED]->(current)When we log signal one (structured event), we also emit a causal edge if the agent explicitly references a prior result. This requires the agent to be instrumented to declare its dependencies. In our swarm, each agent call includes a context_ids field listing the events it consumed.
Why Three? Why Not Two or Four?
Two signals can collude in a bug. If your structured log and your vector embedding both use the same serialization code, a bug in that code corrupts both. Three signals with independent code paths—structured JSON, vector embedding, and graph edge—means a bug would have to simultaneously manifest in three different subsystems. That's unlikely.
Four signals would be more robust, but the overhead is real. Each signal costs compute and storage. For a swarm running nightly LoRA training loops, the log volume is already high. Three is the minimum number that gives you a cross-check without drowning in data.
Implementing the Rule in Practice
Our stack runs on three Hetzner servers with a WireGuard mesh. Each server runs a logging agent that collects events from all local agents and writes them to a shared Postgres database (signal one), a Qdrant collection (signal two), and a Neo4j instance (signal three). The logging agent itself is stateless; if it crashes, events are buffered on disk and replayed.
Here's the core loop:
async def log_agent_action(agent_id, action, payload, context_ids):
event_id = str(uuid.uuid4())
timestamp = datetime.utcnow().isoformat()
# Signal 1: structured log to Postgres
await pg.execute(
"INSERT INTO agent_events (id, agent_id, action, payload, timestamp) VALUES ($1, $2, $3, $4, $5)",
event_id, agent_id, action, json.dumps(payload), timestamp
)
# Signal 2: vector embedding to Qdrant
text = f"{agent_id} {action} {str(payload)}"
embedding = embed_model.encode(text)
await qdrant_client.upsert(
collection="agent_traces",
points=[{
"id": event_id,
"vector": embedding.tolist(),
"payload": {"agent": agent_id, "action": action, "timestamp": timestamp}
}]
)
# Signal 3: causal edges to Neo4j
for ctx_id in context_ids:
await neo4j_session.run(
"""
MATCH (prev:Event {id: $prev_id})
MERGE (curr:Event {id: $curr_id, agent: $agent, action: $action})
CREATE (prev)-[:INFLUENCED]->(curr)
""",
prev_id=ctx_id, curr_id=event_id, agent=agent_id, action=action
)This runs synchronously before the agent's action is committed to the outside world. If any signal fails, the action is rolled back. The agent never acts on incomplete evidence.
Debugging with Three Signals
When something goes wrong, the debugging workflow is:
- Signal one: Query Postgres for events around the failure time. Get a list of actions.
- Signal two: Take the problematic output, embed it, and search Qdrant for similar events. This often reveals a pattern: the agent repeatedly made the same mistake under similar semantic conditions.
- Signal three: Walk the causal graph in Neo4j to find the root influence. Often it's a tool call that returned subtly wrong data three hops back.
Without the third signal, you'd be guessing. With it, you can trace exactly which prior event caused the agent to go off the rails.
Tradeoffs and Caveats
The 3-signal rule adds latency and storage. Each action now requires three writes to three different data stores. On our Hetzner servers, this adds about 5–15 ms per action, which is acceptable for most agent workflows. For high-frequency tool calls, we batch signals and flush every 100 ms.
Storage grows fast. A single agent action might generate 1 KB of structured log, a 768-dimension vector (3 KB), and a few graph nodes (negligible). For a swarm of 50 agents making 10 actions per second, that's about 2 GB per day. We rotate logs weekly and keep vector indexes for 30 days.
The Bottom Line
Treating a single log line as truth is a recipe for silent corruption. The 3-signal rule forces you to build independent evidence paths. When an agent does something unexpected, you don't ask "what happened?"—you ask "which signal is lying?" And because you have three, you can always find the truth.