Agent Swarms vs. Single-Prompt Chains: When Orchestration Overhead Actually Pays Back (and When It Is Theater)
A practical guide to choosing between flat prompts and multi-agent systems.
Agent Swarms vs. Single-Prompt Chains: When Orchestration Overhead Actually Pays Back (and When It Is Theater)
Every few months a new framework drops that promises to turn your LLM into a "swarm of agents" — each one specialized, communicating, delegating. The demos are impressive: a researcher agent, a coder agent, a reviewer agent, all chaining outputs in a beautiful DAG. But when you try to use this for a real task — say, summarizing your company's quarterly reports — you end up with 15 API calls, 30 seconds of latency, and a result that's barely better than a single well-crafted prompt.
This is the central tension: orchestration overhead vs. raw prompt power. When does building a multi-agent system actually pay back? And when is it just theater?
The Case for Single-Prompt Chains
A single-prompt chain is the simplest form of LLM usage: one prompt, one completion, one output. It's what you get when you call llama.cpp with a system prompt and a user message. No loops, no branching, no sub-agents.
When it works:
- Simple extraction tasks: "Extract all dates and amounts from this invoice." A single prompt with a structured output format (e.g., JSON schema) works perfectly.
- Summarization: "Summarize this 10-page document in 3 bullet points." Modern models (Llama 3.1 70B, GPT-4o) handle this with high fidelity.
- Classification: "Is this email urgent? Reply 'yes' or 'no'."
- Translation: Straightforward language mapping.
When it breaks:
- Multi-step reasoning with external tools: The model needs to call a database, then check a file, then compose a response. A single prompt can't interleave tool calls naturally — you end up with a chain of separate prompts.
- Tasks requiring iterative refinement: The model needs to generate code, run it, see errors, fix them. A single prompt can't loop.
- Tasks with multiple conflicting objectives: "Summarize this document, but also extract all email addresses and check if they are valid against our CRM." A single prompt might do both, but quality suffers as the context window fills with instructions.
The overhead: Essentially zero. One HTTP call, one response. Latency is the model's inference time plus a round trip.
The Case for Agent Swarms
An agent swarm is a system where multiple LLM instances (agents) collaborate, each with a specific role, memory, and tool access. They communicate via messages, often orchestrated by a central controller or a protocol like OpenAI's function calling or Microsoft's AutoGen.
When it pays back:
- Complex, multi-step workflows with branching logic: Example: A customer support ticket that requires looking up the user's account, checking order history, reading a knowledge base article, and composing a personalized response. A single agent can do this, but a swarm with a triage agent, a lookup agent, and a response agent can parallelize and specialize.
- Tasks requiring external tool use with state: Each agent can maintain its own context. A retrieval agent can query a vector store (e.g., Qdrant) and pass results to a reasoning agent without bloating the reasoning agent's context window.
- Long-running, fault-tolerant processes: If one agent fails (e.g., a timeout calling an external API), the orchestrator can retry or escalate. In a single chain, a failure means starting over.
- Regulatory or audit requirements: Each agent's actions can be logged independently. For EU AI Act compliance, you might need to prove that the "data minimization agent" never passed PII to the "summarization agent." A swarm's message log provides that trace.
The overhead: Significant. Each agent call adds latency. Orchestration (e.g., a systemd service managing a Celery-like task queue) adds complexity. You need to handle message serialization, agent discovery, and failure modes. Tooling like crewai or langgraph abstracts some of this, but you still pay in infrastructure.
The Theater: When Swarms Are Overkill
I've seen teams build swarms for tasks that could be done with two prompts. Common red flags:
- The "demo-itis" trap: You saw a video of a swarm writing a full app, so you think your email classification system needs six agents. It doesn't.
- Premature modularization: You break a simple task into agents because you think it's "more scalable." But scalability means nothing if the task is trivial.
- Ignoring model capability: Modern models (e.g., Llama 3.1 405B, Claude 3.5 Sonnet) can handle complex instructions in a single prompt. If you're using a small model (e.g., Llama 3.2 3B) that can't follow complex chains, you might need agents — but the better fix is to use a larger model.
- Over-engineering for edge cases: "What if the user asks for a refund AND a replacement?" You build a branching agent tree. But 95% of requests are simple. Handle the edge cases with a fallback prompt, not a swarm.
A concrete example: I once saw a team build a three-agent swarm to answer "What's the weather in Berlin?" — one agent to parse the query, one to call a weather API, one to format the response. The single-prompt alternative: a system prompt that says "You are a weather assistant. To get weather data, call the function get_weather(city) with the city name." That's one call, one function execution. The swarm added 2x latency and 3x code complexity for zero benefit.
When Orchestration Overhead Earns Its Keep
Here are the scenarios where the extra complexity is justified:
1. Multi-source data integration with validation
You need to pull data from a PostgreSQL database, a Qdrant vector store, and a REST API, then combine them with business rules. A single agent can't hold all those connections open. A swarm with a database agent, a retrieval agent, and a validation agent can work in parallel and cross-check results.
Example: A compliance report that checks transactions against sanctions lists (vector search), customer profiles (SQL), and real-time exchange rates (API). The validation agent cross-references all three and flags discrepancies.
2. Long-running processes with human-in-the-loop
A document review workflow: an initial agent drafts a summary, a reviewer agent checks for PII, a human approves, then a final agent formats the output. Each step might take minutes. A single chain would block the entire pipeline. A swarm with a message queue (e.g., RabbitMQ, Redis) allows async handoffs.
3. Iterative code generation and testing
An agent writes code, another agent runs it in a sandbox (e.g., Docker container), another checks test coverage. If the code fails, the coder agent gets the error and fixes it. This loop can run 5-10 times. A single prompt can't loop without external orchestration.
Tooling: vLLM for inference, Docker for sandboxing, a simple systemd timer to retry failed steps.
4. Data sovereignty and audit trails
Under the EU AI Act, you may need to prove that certain data never left a specific agent's context. For example, a "customer data agent" that only sees PII and a "reporting agent" that only sees anonymized aggregates. A swarm's explicit message routing provides an audit log that a single prompt cannot.
How to Decide: A Practical Framework
Before you build a swarm, ask:
- Can a single prompt with function calling do it? If yes, stop. Use that.
- Does the task require looping or conditional branching? If yes, consider a chain (not a swarm) — e.g., a Python script that calls the LLM in a loop.
- Do you need parallel execution of independent sub-tasks? If yes, a swarm might help, but first try a single prompt with parallel function calls (if the model supports it).
- Do you need different context windows or tool access per sub-task? If yes, a swarm is likely justified.
- Is the task latency-sensitive? If yes, avoid swarms. Each agent adds at least one round trip.
- Do you have the infrastructure to monitor and debug a distributed system? If not, you'll spend more time debugging than the swarm saves.
The Bottom Line
Agent swarms are a powerful tool, but they come with real costs: latency, complexity, and infrastructure. Use them when the task genuinely requires specialization, parallelism, or async handoffs. For everything else, a single-prompt chain — or at most a simple chain of two or three prompts — is faster, cheaper, and easier to maintain.
Don't let the theater of multi-agent demos trick you into over-engineering. The best system is the one that solves your problem with the least moving parts. Often, that's just one prompt.