Agent Swarms in Production: What Actually Breaks
War stories from running Postgres-coordinated swarms on RTX 4090s.
We run agent swarms coordinated through Postgres 16 on bare metal with pgbouncer on port 6432. No Redis, no Kafka. Each agent is a process that polls a tasks table, picks a row, runs inference via vLLM (Llama 3.3 70B INT4 on 2x RTX 4090), and writes results back. It sounds clean. It is not. Here are the things that broke—and what we did about them.
Postgres Deadlocks When Agents Compete for Rows
The simplest pattern—SELECT ... FOR UPDATE SKIP LOCKED—works fine at low concurrency. At 50+ agents polling every 100ms, we started seeing deadlock detected errors daily. Agents would hold row locks while calling vLLM (which can take seconds), and other agents waiting on those rows would block each other in lock queues. The fix was twofold: (1) use SKIP LOCKED religiously so agents never wait on locked rows—they just skip to the next task, and (2) add a locked_at timestamp with a 30-second TTL so stuck agents can be reaped by a watchdog. We also moved to advisory locks for critical sections (e.g., task assignment) to reduce row-level contention. Since then, zero deadlocks in 4 months.
vLLM OOM on Bad Prompts
A single prompt with 32k tokens of repeated "A" characters caused vLLM to allocate a key-value cache that exceeded our 24 GB VRAM per GPU. The process died silently, leaving agents hanging on pending tasks. We now enforce a hard token limit of 8192 on the application side—truncate or reject before sending to vLLM. We also set --max-model-len 8192 in vLLM and monitor VRAM usage via nvidia-smi every 5 seconds. If VRAM exceeds 22 GB, we drain the GPU and failover to DeepSeek (paid fallback). The cost of DeepSeek is painful ($0.80 per million tokens) but better than a silent OOM that corrupts a batch of 200 tasks.
Runaway Loops Eating Tokens
An agent was tasked with summarizing a conversation. It returned a summary that was longer than the original. Another agent then summarized the summary. This looped 14 times before we noticed the token bill on DeepSeek ($23 in one hour). The root cause: no guardrail on output length relative to input. We now enforce: if output tokens > 2x input tokens, the agent must explain why in a structured log field. We also cap total agent steps per task at 5 (configurable per task type). The runaway loop now hits the step limit and the task is flagged for human review. We log every step input/output to Postgres—costly in storage (~2 GB/day) but invaluable for debugging.
Retry Storms After Partial Failures
vLLM occasionally returns a partial response (e.g., connection reset mid-stream). Our first retry logic was naive: retry up to 3 times with exponential backoff. But when 20 agents all hit the same GPU at the same time, a transient spike would cause 60 retries in 2 seconds, overwhelming the GPU further. We now use a centralized retry budget: a Postgres table retry_budget with a per-task-type token bucket. Each retry consumes a token; if the bucket is empty, the task is deferred for 60 seconds. We also randomize retry jitter between 1 and 5 seconds. This smoothed out the retry storm and improved overall throughput by 40%.
Partial-Failure Ambiguity: Did the Agent Finish?
An agent writes a status column: pending, running, done, failed. But what if the agent crashes after writing running but before writing the result? The task is stuck. We added a heartbeat column last_heartbeat_at that agents update every 5 seconds while working. A separate watchdog process (running every 10 seconds) resets any task with status = 'running' and last_heartbeat_at < now() - 30s back to pending. This introduced a new problem: if the agent is slow but alive, the watchdog steals the task. We tuned the heartbeat interval to 5 seconds and the timeout to 15 seconds—aggressive but safe for our typical inference times (2-10 seconds).
Observability Gaps: You Can't Fix What You Can't See
We had no tracing. When a task failed, we had to grep logs across 10 front nodes (Hetzner EX130-S). We now log every state transition to a task_events table: task_id, from_status, to_status, agent_id, timestamp, and a free-form reason text. We also emit structured logs (JSON) to stdout, which we ship to a simple Loki instance. The cost is minimal (~100 MB/day) and the debugging time dropped from hours to minutes. We also added a Grafana dashboard showing: number of active agents, queue depth, average inference time, retry rate, and OOM count. The OOM count graph is our favorite—it's been flat for 3 weeks.
What We'd Do Differently
Start with a task-level timeout from day one. We didn't, and a single stuck task once blocked a downstream pipeline for 2 hours. Use SKIP LOCKED from the start—it's not optional at any concurrency above 10. Implement a centralized retry budget before you need it; retrofitting it is painful. And for the love of all that is holy, log everything to a structured table, not just stdout. We spent two weeks replaying logs from disk after a node crashed.
We're still running this stack. It works, but it's held together with Postgres constraints and careful tuning. If you're evaluating self-hosted AI infrastructure, expect to write your own guardrails. We'll share our patterns—just ask.