Why We Replaced Generic Agent Loops with DAGs: A State Machine Autopsy
Every agent swarm starts with a simple loop: while True: observe, think, act. It feels natural—like a cognitive cycle. But once you have more than a handful of agents, or agents with branching decisions, that loop turns into a tangled mess of race conditions, lost context, and unkillable runaway processes. We've been through that pain, and the fix was to throw out the generic loop and model everything as a directed acyclic graph (DAG) of state machines.
The Failure Modes of Generic Loops
A generic agent loop looks elegant in a demo. Pseudocode:
while True:
observation = sense_environment()
thought = llm_reason(observation, memory)
action = decide_action(thought)
execute(action)
memory.update(observation, thought, action)This works for a single agent with a single linear task. But introduce multiple agents, parallel branches, conditional retries, or human-in-the-loop pauses, and the loop becomes a liability:
- State explosion: Each agent's
memorybecomes a kitchen sink. Branching decisions pile up as nested conditionals. The loop's single continuation point makes it impossible to isolate a failing branch. - No fault isolation: If one agent's
llm_reasonhangs or returns garbage, the entire loop stalls. There's no way to replace that agent mid-flight. - Non-deterministic replay: Debugging requires replaying the exact sequence of observations and decisions. A while-true loop with external side effects (API calls, database writes) is nearly impossible to replay deterministically.
- Scaling overhead: To run agents in parallel, you need to wrap the loop in threads or async tasks. Now you have shared memory, locks, and the joy of debugging deadlocks.
The DAG State Machine Pattern
Instead of a loop, we model each agent's lifecycle as a state machine. The overall orchestration is a DAG where nodes are state machines and edges are data dependencies or control flow. This is not new—workflow engines have used DAGs for decades—but applying it to agent swarms changes how you think about reasoning and action.
Each agent state machine has a fixed set of states: IDLE, OBSERVING, REASONING, ACTING, WAITING, ERROR, DONE. Transitions are triggered by events (e.g., observation ready, LLM response received, timeout). The DAG defines which agents run, in what order, and how data flows between them.
from enum import Enum
class AgentState(Enum):
IDLE = 0
OBSERVING = 1
REASONING = 2
ACTING = 3
WAITING = 4
ERROR = 5
DONE = 6
class AgentStateMachine:
def __init__(self, agent_id, config):
self.id = agent_id
self.state = AgentState.IDLE
self.config = config
self.data = {}
def transition(self, event, payload=None):
# event-driven transitions; no while loop
if self.state == AgentState.IDLE and event == 'start':
self.state = AgentState.OBSERVING
self.data['observation'] = payload
elif self.state == AgentState.OBSERVING and event == 'observation_ready':
self.state = AgentState.REASONING
self.data['observation'] = payload
elif self.state == AgentState.REASONING and event == 'reasoning_complete':
self.state = AgentState.ACTING
self.data['action'] = payload
elif self.state == AgentState.ACTING and event == 'action_executed':
self.state = AgentState.DONE
self.data['result'] = payload
elif event == 'error':
self.state = AgentState.ERROR
self.data['error'] = payload
# ... more transitions for WAITING, retries, etc.
return self.stateOrchestration as a DAG
The DAG itself is a declarative structure. We define it in YAML or Python dicts, not in imperative code. Each node is an agent state machine, each edge is a data dependency. The orchestrator walks the DAG, starting nodes whose dependencies are satisfied, and propagates data along edges.
# dag-definition.yaml
nodes:
- id: researcher
agent_type: web_search
config: { engine: "google", max_results: 5 }
- id: summarizer
agent_type: llm_summarize
config: { model: "qwen2.5:7b" }
- id: fact_checker
agent_type: llm_fact_check
config: { model: "qwen2.5:7b", temperature: 0.0 }
- id: writer
agent_type: llm_generate
config: { model: "qwen2.5:14b", max_tokens: 2000 }
edges:
- from: researcher
to: summarizer
data_key: search_results
- from: summarizer
to: fact_checker
data_key: summary
- from: fact_checker
to: writer
data_key: verified_summary
- from: researcher
to: writer
data_key: raw_search_results # optional side channelThe orchestrator reads this DAG, instantiates the state machines, and starts them as events flow. A simple topological sort ensures execution order. Because each node is isolated, you can replace the researcher agent with a different implementation without touching the rest of the DAG.
Fault Isolation and Retries
In a while-true loop, a single LLM timeout poisons the entire agent. In a DAG, each node has its own error state. The orchestrator can retry the node, skip it, or route to a fallback node—all without affecting sibling nodes.
def handle_node_error(node, error):
if node.config.get('retry_on_error', False):
node.reset_state()
orchestrator.schedule_retry(node, delay=5)
elif node.config.get('fallback_node'):
fallback = node.config['fallback_node']
orchestrator.reroute(node, fallback)
else:
orchestrator.mark_node_failed(node)
# propagate error downstream? depends on DAG semanticsDeterministic Replay
Because each state machine transition is event-driven and the DAG structure is static, you can log every event and replay the entire execution. This is invaluable for debugging: you can feed the exact same sequence of LLM responses, API results, and timeouts to a local replica of the DAG and see if the outcome matches.
class EventLog:
def __init__(self):
self.events = []
def record(self, node_id, event, payload):
self.events.append({
'timestamp': time.time(),
'node_id': node_id,
'event': event,
'payload': payload
})
def replay(self, dag, event_log):
for entry in event_log:
node = dag.get_node(entry['node_id'])
node.transition(entry['event'], entry['payload'])Hot-Swappable Branches
In production, you may want to A/B test different agent strategies. With a DAG, you can define multiple sub-DAGs for a single task and route traffic based on user_id, latency budget, or random split. The orchestrator simply selects which DAG to instantiate per request.
branch_a = load_dag('fast_cheap.yaml')
branch_b = load_dag('thorough_expensive.yaml')
if user.is_premium:
dag = branch_b
else:
dag = branch_a
orchestrator.run(dag, input_data)Tradeoffs and When Not to DAG
DAG-based orchestration adds upfront complexity. You need to define states, transitions, and edges explicitly. For a single, linear agent that never branches, a while-true loop is simpler and faster to write. The DAG pattern shines when you have:
- Multiple agents with data dependencies
- Conditional branching (if fact check fails, re-run researcher)
- Human-in-the-loop pauses (WAITING state)
- Need for observability and replay
We also found that the DAG approach encourages a more disciplined separation of concerns: each node is a pure function of its inputs and state. This makes testing trivial.
Implementation Notes
We run this on a modest stack: three Hetzner servers with WireGuard mesh, Qwen models served via vLLM, BGE-M3 for embeddings, and Postgres + Qdrant + Neo4j for storage. The orchestrator itself is a lightweight Python process that reads DAG definitions from a Git repository (version-controlled orchestration!). The state machines are pure Python classes; we serialize their state to Postgres for persistence across restarts.
Nightly LoRA training updates the Qwen models used by the fact_checker and writer nodes. The DAG definition doesn't change—only the model weights do—so we can hot-swap the model behind a node without touching the orchestration logic.
The Autopsy Verdict
We didn't throw out the while-true loop because it was broken for all use cases. We threw it out because it was broken for the use cases that matter: multi-agent, branching, fault-tolerant, observable swarms. The DAG state machine pattern gave us deterministic behavior, clean fault isolation, and the ability to reason about the system as a whole rather than as a soup of competing loops.
If your agent swarm is still a single while-true loop and you're adding more agents, you're building a house of cards. Consider the DAG. Your future self—debugging a production incident at 2 AM—will thank you.