Why We Replaced Generic Agent Loops with DAGs: A State Machine Autopsy

by

Every agent swarm starts with a simple loop: while True: observe, think, act. It feels natural—like a cognitive cycle. But once you have more than a handful of agents, or agents with branching decisions, that loop turns into a tangled mess of race conditions, lost context, and unkillable runaway processes. We've been through that pain, and the fix was to throw out the generic loop and model everything as a directed acyclic graph (DAG) of state machines.

The Failure Modes of Generic Loops

A generic agent loop looks elegant in a demo. Pseudocode:

while True:
    observation = sense_environment()
    thought = llm_reason(observation, memory)
    action = decide_action(thought)
    execute(action)
    memory.update(observation, thought, action)

This works for a single agent with a single linear task. But introduce multiple agents, parallel branches, conditional retries, or human-in-the-loop pauses, and the loop becomes a liability:

  • State explosion: Each agent's memory becomes a kitchen sink. Branching decisions pile up as nested conditionals. The loop's single continuation point makes it impossible to isolate a failing branch.
  • No fault isolation: If one agent's llm_reason hangs or returns garbage, the entire loop stalls. There's no way to replace that agent mid-flight.
  • Non-deterministic replay: Debugging requires replaying the exact sequence of observations and decisions. A while-true loop with external side effects (API calls, database writes) is nearly impossible to replay deterministically.
  • Scaling overhead: To run agents in parallel, you need to wrap the loop in threads or async tasks. Now you have shared memory, locks, and the joy of debugging deadlocks.

The DAG State Machine Pattern

Instead of a loop, we model each agent's lifecycle as a state machine. The overall orchestration is a DAG where nodes are state machines and edges are data dependencies or control flow. This is not new—workflow engines have used DAGs for decades—but applying it to agent swarms changes how you think about reasoning and action.

Each agent state machine has a fixed set of states: IDLE, OBSERVING, REASONING, ACTING, WAITING, ERROR, DONE. Transitions are triggered by events (e.g., observation ready, LLM response received, timeout). The DAG defines which agents run, in what order, and how data flows between them.

from enum import Enum

class AgentState(Enum):
    IDLE = 0
    OBSERVING = 1
    REASONING = 2
    ACTING = 3
    WAITING = 4
    ERROR = 5
    DONE = 6

class AgentStateMachine:
    def __init__(self, agent_id, config):
        self.id = agent_id
        self.state = AgentState.IDLE
        self.config = config
        self.data = {}

    def transition(self, event, payload=None):
        # event-driven transitions; no while loop
        if self.state == AgentState.IDLE and event == 'start':
            self.state = AgentState.OBSERVING
            self.data['observation'] = payload
        elif self.state == AgentState.OBSERVING and event == 'observation_ready':
            self.state = AgentState.REASONING
            self.data['observation'] = payload
        elif self.state == AgentState.REASONING and event == 'reasoning_complete':
            self.state = AgentState.ACTING
            self.data['action'] = payload
        elif self.state == AgentState.ACTING and event == 'action_executed':
            self.state = AgentState.DONE
            self.data['result'] = payload
        elif event == 'error':
            self.state = AgentState.ERROR
            self.data['error'] = payload
        # ... more transitions for WAITING, retries, etc.
        return self.state

Orchestration as a DAG

The DAG itself is a declarative structure. We define it in YAML or Python dicts, not in imperative code. Each node is an agent state machine, each edge is a data dependency. The orchestrator walks the DAG, starting nodes whose dependencies are satisfied, and propagates data along edges.

# dag-definition.yaml
nodes:
  - id: researcher
    agent_type: web_search
    config: { engine: "google", max_results: 5 }
  - id: summarizer
    agent_type: llm_summarize
    config: { model: "qwen2.5:7b" }
  - id: fact_checker
    agent_type: llm_fact_check
    config: { model: "qwen2.5:7b", temperature: 0.0 }
  - id: writer
    agent_type: llm_generate
    config: { model: "qwen2.5:14b", max_tokens: 2000 }
edges:
  - from: researcher
    to: summarizer
    data_key: search_results
  - from: summarizer
    to: fact_checker
    data_key: summary
  - from: fact_checker
    to: writer
    data_key: verified_summary
  - from: researcher
    to: writer
    data_key: raw_search_results  # optional side channel

The orchestrator reads this DAG, instantiates the state machines, and starts them as events flow. A simple topological sort ensures execution order. Because each node is isolated, you can replace the researcher agent with a different implementation without touching the rest of the DAG.

Fault Isolation and Retries

In a while-true loop, a single LLM timeout poisons the entire agent. In a DAG, each node has its own error state. The orchestrator can retry the node, skip it, or route to a fallback node—all without affecting sibling nodes.

def handle_node_error(node, error):
    if node.config.get('retry_on_error', False):
        node.reset_state()
        orchestrator.schedule_retry(node, delay=5)
    elif node.config.get('fallback_node'):
        fallback = node.config['fallback_node']
        orchestrator.reroute(node, fallback)
    else:
        orchestrator.mark_node_failed(node)
        # propagate error downstream? depends on DAG semantics

Deterministic Replay

Because each state machine transition is event-driven and the DAG structure is static, you can log every event and replay the entire execution. This is invaluable for debugging: you can feed the exact same sequence of LLM responses, API results, and timeouts to a local replica of the DAG and see if the outcome matches.

class EventLog:
    def __init__(self):
        self.events = []

    def record(self, node_id, event, payload):
        self.events.append({
            'timestamp': time.time(),
            'node_id': node_id,
            'event': event,
            'payload': payload
        })

    def replay(self, dag, event_log):
        for entry in event_log:
            node = dag.get_node(entry['node_id'])
            node.transition(entry['event'], entry['payload'])

Hot-Swappable Branches

In production, you may want to A/B test different agent strategies. With a DAG, you can define multiple sub-DAGs for a single task and route traffic based on user_id, latency budget, or random split. The orchestrator simply selects which DAG to instantiate per request.

branch_a = load_dag('fast_cheap.yaml')
branch_b = load_dag('thorough_expensive.yaml')

if user.is_premium:
    dag = branch_b
else:
    dag = branch_a

orchestrator.run(dag, input_data)

Tradeoffs and When Not to DAG

DAG-based orchestration adds upfront complexity. You need to define states, transitions, and edges explicitly. For a single, linear agent that never branches, a while-true loop is simpler and faster to write. The DAG pattern shines when you have:

  • Multiple agents with data dependencies
  • Conditional branching (if fact check fails, re-run researcher)
  • Human-in-the-loop pauses (WAITING state)
  • Need for observability and replay

We also found that the DAG approach encourages a more disciplined separation of concerns: each node is a pure function of its inputs and state. This makes testing trivial.

Implementation Notes

We run this on a modest stack: three Hetzner servers with WireGuard mesh, Qwen models served via vLLM, BGE-M3 for embeddings, and Postgres + Qdrant + Neo4j for storage. The orchestrator itself is a lightweight Python process that reads DAG definitions from a Git repository (version-controlled orchestration!). The state machines are pure Python classes; we serialize their state to Postgres for persistence across restarts.

Nightly LoRA training updates the Qwen models used by the fact_checker and writer nodes. The DAG definition doesn't change—only the model weights do—so we can hot-swap the model behind a node without touching the orchestration logic.

The Autopsy Verdict

We didn't throw out the while-true loop because it was broken for all use cases. We threw it out because it was broken for the use cases that matter: multi-agent, branching, fault-tolerant, observable swarms. The DAG state machine pattern gave us deterministic behavior, clean fault isolation, and the ability to reason about the system as a whole rather than as a soup of competing loops.

If your agent swarm is still a single while-true loop and you're adding more agents, you're building a house of cards. Consider the DAG. Your future self—debugging a production incident at 2 AM—will thank you.

#agent-swarms#dag#engineering#orchestration#state-machine
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related