Agent Swarm State Machines: Why We Replaced Generic Loops with DAGs and What Broke

by
Agent Swarm State Machines: Why We Replaced Generic Loops with DAGs and What Broke

Agent Swarm State Machines: Why We Replaced Generic Loops with DAGs and What Broke

When you first wire up an agent swarm, the natural instinct is to throw a while loop around a list of agents and call them in sequence. It works for two agents. It works for five. Then you add branching, retries, and parallel sub-tasks, and suddenly your loop is a tangled mess of flags, timeouts, and race conditions.

We hit that wall six months ago. Our swarm of 12 specialized agents (planner, coder, reviewer, tester, deployer, etc.) was supposed to autonomously ship features. Instead, we spent more time debugging loop logic than building features. So we ripped it out and replaced it with a Directed Acyclic Graph (DAG) state machine. Here's what we learned, what broke, and why you should consider the same switch.

The Problem with Generic Loops

Our first implementation looked like this (simplified):

agents = [planner, coder, reviewer, tester, deployer]
max_iterations = 10
for i in range(max_iterations):
    for agent in agents:
        result = agent.run(context)
        if result.status == "failed":
            # retry logic
            pass
        if result.status == "blocked":
            # wait and retry later
            pass

This worked until we added parallel agents (e.g., run unit tests and integration tests simultaneously). Suddenly we needed threading, shared state, and error propagation across loops. The loop became a god function that knew about every agent's internal state. Adding a new agent meant modifying the loop logic. Code reviews turned into archaeology.

Key failures we observed:

  1. No explicit state transitions – The loop implicitly assumed a linear flow, but real swarms have conditional paths (e.g., if tests fail, go back to coder; if security scan fails, block deploy).
  2. Error handling was an afterthought – Retries were nested inside the loop, leading to exponential backoff spaghetti. A single agent crash could hang the entire loop.
  3. Parallelism was bolted on – We used Python's concurrent.futures but had no way to express dependencies (e.g., "deploy waits for both test suites to pass").
  4. Observability was poor – Logging told us what agent ran, but not why it ran or what state it expected.

DAGs to the Rescue

A DAG-based state machine models each agent as a node, and edges represent state transitions. Each node has a single responsibility: execute a task, then emit a signal (success, failure, skip). A central orchestrator reads the graph and advances the state based on those signals.

We built ours on top of a Postgres-backed state store (using pgvector for embedding similarity checks, but that's another story). The core data structures:

CREATE TABLE dag_nodes (
    id UUID PRIMARY KEY,
    swarm_id UUID NOT NULL,
    agent_type TEXT NOT NULL,  -- e.g., 'planner', 'coder'
    status TEXT DEFAULT 'pending',  -- pending, running, success, failed, skipped
    depends_on UUID[] DEFAULT '{}',
    retry_count INT DEFAULT 0,
    max_retries INT DEFAULT 3,
    created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE dag_edges (
    id UUID PRIMARY KEY,
    from_node UUID REFERENCES dag_nodes(id),
    to_node UUID REFERENCES dag_nodes(id),
    condition TEXT DEFAULT 'success',  -- 'success', 'failure', 'always'
    created_at TIMESTAMPTZ DEFAULT NOW()
);

Each agent, when invoked, receives the full context (stored as JSONB in a contexts table) and returns a status. The orchestrator then walks the graph:

async def run_dag(swarm_id):
    while True:
        ready_nodes = await get_ready_nodes(swarm_id)  # all dependencies satisfied
        if not ready_nodes:
            break
        results = await asyncio.gather(*[run_node(node) for node in ready_nodes])
        for node, result in zip(ready_nodes, results):
            await update_node_status(node.id, result.status)
            if result.status == "failed" and node.retry_count < node.max_retries:
                await reset_node(node.id)
            else:
                await advance_edges(node.id, result.status)

This removed the loop entirely. The graph defines the flow; the orchestrator just follows the edges.

What Broke in Practice

Switching to DAGs wasn't painless. Here are the concrete failures we encountered, and how we fixed them.

1. Cyclic Dependencies (The "Loop" Returns)

We naively assumed all workflows would be acyclic. But some workflows require iterative refinement: coder produces code, reviewer suggests changes, coder updates. That's a cycle. Our DAG rejected it.

Fix: We introduced a max_iterations per node and a loopback edge type. When a reviewer returns "changes requested", the DAG creates a new instance of the coder node with a fresh ID, linked back from the reviewer. The old coder node is marked superseded. This keeps the graph acyclic while allowing iterative loops.

2. State Explosion

With 12 agents and conditional branches, the number of possible states grew combinatorially. Our Postgres queries for "get ready nodes" became slow (seconds) when a swarm had hundreds of nodes.

Fix: We added a materialized view that precomputes readiness for each node based on its dependencies. Refreshed on each status change. Also, we limited the DAG depth to 50 nodes per swarm (rarely hit in practice).

3. Orphaned Nodes

If an agent crashes without returning a status, its dependents wait forever. The DAG stalls.

Fix: Each node gets a TTL (time-to-live). If no status update within 5 minutes (configurable), the orchestrator marks it as timed_out and follows failure edges. We also added a watchdog that sweeps for stale nodes every 30 seconds.

4. Context Bloat

Agents share context via a JSONB blob. As the DAG progresses, the context grows (logs, artifacts, intermediate results). Postgres TOAST tables kicked in, and queries slowed.

Fix: We partitioned context into two tables: context_metadata (always loaded) and context_blobs (loaded on demand). Agents declare what blobs they need upfront. The orchestrator passes only the required blobs.

5. Debugging the Graph

When something went wrong, we had to trace through dozens of nodes. Logs were scattered.

Fix: We built a small web UI (Flask + D3.js) that renders the DAG in real-time, color-coding nodes by status. Also, every node logs its input and output context to a dedicated node_logs table. We use pgvector to search logs by similarity when debugging.

When to Use DAGs vs. Loops

Not every swarm needs a DAG. If your agents run in a strict linear order with no branching or parallelism, a loop is simpler. But as soon as you have:

  • Conditional paths (if A fails, run B else run C)
  • Parallel execution with dependencies
  • Retry logic that varies per agent
  • Need for audit trails and state replay

...a DAG state machine is worth the complexity.

Our Current Setup

We run the orchestrator as a systemd service on a single VM (4 vCPU, 16 GB RAM). Postgres 15 with pgvector 0.5.1 stores the graph and context. For inference, we use vLLM 0.3.3 for LLM agents (like planner and reviewer) and llama.cpp for smaller models (like coder and tester). The DAG overhead (graph traversal, state updates) adds about 10ms per node, negligible compared to LLM inference times (seconds).

We've been running this in production for 3 months. Our median feature-to-deploy time dropped from 45 minutes to 12 minutes. Debug time per incident went down by 60%. The graph makes failures visible: we can see exactly which agent failed and why, without digging through logs.

What's Next

We're extending the DAG to support dynamic node insertion: agents that can add new nodes to the graph at runtime (e.g., a security agent that spawns a vulnerability scan node). This requires careful validation to prevent infinite loops, but it's the next step toward truly autonomous swarms.

We're also experimenting with using the DAG as a training signal for reinforcement learning: reward the orchestrator for completing swarms quickly, penalize it for excessive retries. Early results suggest the DAG structure simplifies credit assignment – you can trace which node caused the delay.

If you're building agent swarms, skip the loop. Design your state machine as a DAG from day one. Your future self will thank you.

#agent-design#error-handling#orchestration
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related