Running an Agent Swarm: Orchestration, OAuth, and the Failure Modes Nobody Warns You About

Lessons from building a multi-agent system on-prem

by
Running an Agent Swarm: Orchestration, OAuth, and the Failure Modes Nobody Warns You About

Running an Agent Swarm: Orchestration, OAuth, and the Failure Modes Nobody Warns You About

Agent swarms are the hot new pattern: multiple LLM-powered agents collaborating on tasks. The demos look slick. But when you actually run them in production—on your own iron—you hit a wall of operational reality. This article covers three areas that will bite you: orchestration, OAuth, and silent failure modes. I'll skip the theory and tell you what broke for me.

Orchestration: The Control Plane Is Not Optional

You can't just spin up 20 Python scripts and call it a swarm. Agents need a control plane that manages lifecycle, communication, and state. I've seen teams reach for Celery or plain Redis queues. That works for simple pipelines, but agent swarms have loops, branching, and conditional handoffs. You need something that can model a directed acyclic graph (DAG) or at least a state machine.

My Stack: Postgres + pgque + systemd

I run a lightweight orchestrator built on Postgres with the pgque extension for job queues. Each agent is a systemd unit. The orchestrator writes tasks to a queue table, agents poll for work, and report status back. It's boring, it's reliable, and it doesn't require Kubernetes.

Key design decisions:

  • Idempotent tasks: Every task has a unique ID. If an agent crashes and restarts, it picks up where it left off. Use ON CONFLICT DO NOTHING in Postgres to avoid duplicates.
  • Heartbeat timeouts: Agents write a heartbeat timestamp to Postgres every 30 seconds. The orchestrator has a background worker that kills and restarts agents that miss three heartbeats.
  • Backpressure: The queue has a max depth. If it's full, the orchestrator refuses new tasks and returns a 503. This prevents cascading failures.

A common failure: agents that get stuck in an infinite loop because the LLM keeps generating the same output. I added a cycle detector that tracks task signatures (input hash + agent ID). If the same signature appears more than 3 times, the orchestrator sends a kill signal and logs an alert.

Why Not Kubernetes?

Kubernetes adds complexity you don't need for a small swarm. If you have fewer than 50 agents, systemd + Postgres is simpler to debug. When an agent fails, you SSH in and check the journal. No YAML spelunking.

OAuth: Token Management Is a Nightmare

Agents often need to call external APIs—GitHub, Slack, Google Drive, whatever. OAuth flows assume a browser and a human. Agents don't have those. You have to store refresh tokens and handle expiry silently. Get it wrong and your swarm silently stops working.

The Pattern: Long-Lived Refresh Tokens + Postgres Vault

I store OAuth tokens in a Postgres table encrypted at rest using pgcrypto. Each agent has a service account with a refresh token that never expires (if the provider allows it). For providers that expire refresh tokens (looking at you, Google), I run a cron job that re-authenticates via a stored service account credential.

Critical points:

  • Token refresh is a background task: Agents never refresh tokens themselves. They fetch a valid access token from a token service that runs as a separate systemd unit. The token service checks expiry and refreshes if needed. This prevents race conditions where two agents refresh the same token simultaneously and invalidate each other.
  • Rate limiting: The token service has a rate limiter (token bucket) to avoid hitting provider rate limits on refresh endpoints.
  • Logging: Every refresh attempt is logged with token ID, success/failure, and error message. I can't stress this enough—without logs, you'll waste hours wondering why an agent suddenly can't access a resource.

Failure Mode: Token Revocation Without Notice

A provider can revoke a token at any time (user deletes app, security policy changes). Your agent gets a 401, retries a few times, then gives up. If you don't monitor 401 rates, you'll discover the outage days later. I now have a Prometheus alert on 401 counts per agent.

The Failure Modes Nobody Warns You About

These are the silent killers that don't show up in demo notebooks.

1. LLM Hallucination of Tool Calls

Your agent is supposed to call search_database(query) but instead calls search_database(query, extra_param=True) because the LLM invented a parameter. If your tool validation is loose, the call succeeds with wrong results. I now validate all tool call arguments against a JSON schema before execution. If validation fails, the agent gets an error message and must retry. Schema validation is done with Python's jsonschema library on the orchestrator side, not the agent.

2. Deadlock from Circular Dependencies

Agent A waits for Agent B, Agent B waits for Agent C, Agent C waits for Agent A. This happens when you have a task that requires multiple outputs. I added a timeout to every inter-agent dependency: if an agent waits more than 60 seconds for a dependency, it times out and logs a warning. The orchestrator then inspects the dependency graph and breaks the cycle by killing the youngest agent in the chain.

3. Token Context Bloat

Every step in a conversation adds tokens. After 10 rounds, the context window is full and the agent starts forgetting. The naive fix is to truncate the oldest messages, but that loses context. I use a sliding window with summarization: when the context reaches 70% of the model's limit, a summarizer agent condenses the conversation history into a short summary that replaces the oldest messages. The summarizer runs as a separate, faster model (e.g., Mistral 7B) to avoid slowing down the main agent.

4. Resource Exhaustion from Parallel Agents

If you run 10 agents that each use vLLM with 2 GPUs, you'll OOM. I pin each agent to a specific GPU via CUDA_VISIBLE_DEVICES and limit the number of concurrent agents per GPU to a value derived from the model size and GPU memory. For example, on an A100 80GB running Llama 3 70B, I allow at most 2 concurrent agents. This is configured in a YAML file that the orchestrator reads.

5. Silent Data Corruption

Agents write intermediate results to Postgres. If two agents write to the same row concurrently, you get lost updates. I use SELECT ... FOR UPDATE with row-level locking for critical writes. For high-throughput writes, I use a separate append-only table and a materialized view for reads.

Operational Checklist

If you're building an agent swarm today, here's what I'd put in place from day one:

  • Centralized logging with structured JSON. Every agent logs to stdout, which is captured by journald and forwarded to Loki. I search by agent ID, task ID, and error type.
  • Health endpoint on every agent: /health returns 200 with uptime, last task timestamp, and queue depth. The orchestrator polls this every 30 seconds.
  • Graceful shutdown: systemd sends SIGTERM, agent finishes current task, writes state to Postgres, then exits. If it doesn't exit in 10 seconds, systemd sends SIGKILL.
  • Canary agent: One agent runs a synthetic task every minute. If it fails, the whole swarm is considered degraded and alerts fire.

Conclusion

Agent swarms are powerful, but they're not magic. The orchestrator is the unsung hero. OAuth will break at 2 AM. And silent failures—hallucinated tool calls, deadlocks, token bloat—will rot your system from the inside. Treat your swarm like a distributed system, not a bunch of clever scripts. Use boring technology, validate everything, and monitor aggressively. Your future self will thank you.

#agent-swarms#oauth#ops#orchestration#self-hosted
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related