Silent Backpressure: How One Stalled Inference Blocked a 50-Instance Swarm Mesh
When your agent swarm freezes, look for the hidden queue that nobody monitors.
Silent Backpressure: How One Stalled Inference Blocked a 50-Instance Swarm Mesh
You've scaled your agent swarm to 50 instances. Each agent is a self-contained loop: listen, reason, act. They share a Postgres-backed work queue and use a local vLLM endpoint for inference. It's beautiful—until it stops.
No crash. No error. Just silence. Agents that were processing tasks now sit idle. New tasks pile up. The UI shows "healthy" but throughput is zero.
This is the story of how one stalled inference call brought down an entire swarm mesh, and how you can prevent it.
The Setup
Our swarm uses a simple architecture:
- Agents (Python, asyncio) poll a Postgres table for work items.
- Each agent picks a task, calls the LLM (vLLM 0.5.3 on a single GPU), processes the response, and writes results.
- A central scheduler (Celery-like, but custom) assigns tasks to agents via the DB.
- Monitoring: Prometheus scrapes per-agent metrics (tasks processed, queue depth).
50 agents, each with its own event loop. No shared state except the database.
The Incident
At 14:23, throughput dropped to zero. Agents were alive—heartbeats registered—but no tasks completed. The work queue grew unbounded. Restarting a single agent didn't help; the new agent also stalled.
We checked:
- Postgres: idle, no locks.
- vLLM: no errors, queue depth low.
- Network: fine.
Then we looked at the agent logs. Every agent was stuck on the same line:
response = await llm_client.complete(prompt)No timeout. The complete method never returned. One agent had started a long inference (a 4096-token generation) and the vLLM worker was busy. But why did all agents block? vLLM is concurrent.
The Root Cause
We had configured vLLM with --max-num-seqs 1. This limits the number of concurrent sequences the engine processes. Normally, that shouldn't cause blocking—other requests queue up in vLLM's internal scheduler. But our HTTP client had no timeout and no connection pooling limit. Every agent made a new HTTP connection to vLLM. The server's backlog filled, and eventually the OS TCP backlog limit was hit. New connections were dropped silently. The agents' aiohttp sessions retried indefinitely, never failing.
The result: all 50 agents blocked on a single stalled inference, waiting for a connection that would never be established.
This is silent backpressure. The system looked healthy but was completely deadlocked.
Why It's Insidious
Traditional backpressure propagates upstream: if the LLM is slow, the agent slows, then the queue grows, and the scheduler stops pushing. That's visible. Here, the backpressure was invisible:
- vLLM reported no errors (it was busy, not failing).
- Agents reported no errors (they were waiting, not crashing).
- The database showed no contention.
Only by instrumenting the HTTP layer could we see the stalled connections.
How to Prevent It
1. Always Set Timeouts
Every external call must have a timeout. For LLM inference, set a generous one (e.g., 300s) but enforce it.
async with aiohttp.ClientSession() as session:
try:
async with session.post(url, json=data, timeout=aiohttp.ClientTimeout(total=300)) as resp:
result = await resp.json()
except asyncio.TimeoutError:
logger.error("LLM inference timed out")
# handle gracefully: retry, skip, or alert2. Limit Connection Pool
Don't let agents open unlimited connections. Use a connection pool with a max size:
connector = aiohttp.TCPConnector(limit=10, limit_per_host=5)
session = aiohttp.ClientSession(connector=connector)This prevents the OS backlog from flooding.
3. Monitor Connection States
Prometheus can track active connections per host. If connections to vLLM spike, alert.
# HELP llm_http_connections_active Active connections to LLM server
# TYPE llm_http_connections_active gauge
llm_http_connections_active{host="vllm:8000"} 42If that number plateaus at the pool limit, you have backpressure.
4. Implement Circuit Breakers
If the LLM endpoint fails repeatedly, open the circuit. Don't let agents hammer a dead service.
circuit = CircuitBreaker(failure_threshold=5, recovery_timeout=30)
@circuit
async def call_llm(prompt):
...5. Use Backpressure-Aware Queues
Instead of a simple Postgres table, use a queue that exposes its depth and limits producers. For example, PgBouncer's queue depth or a dedicated message broker like RabbitMQ with prefetch limits.
The Fix
We added:
- Timeout of 300s on all LLM calls.
- Connection pool limit of 10 per agent.
- A circuit breaker after 3 consecutive timeouts.
- A Prometheus alert on
llm_http_connections_active > 50.
After deploying, the swarm survived a 1200-second inference without blocking.
Key Takeaways
- Monitor the network layer, not just application metrics. Connection pools and TCP backlogs are invisible until they break.
- Set timeouts everywhere. An infinite wait is a bug.
- Limit concurrency at the client. Don't let each agent open unlimited connections.
- Test failure modes. Simulate a slow inference and see what happens.
Silent backpressure is the most dangerous failure mode in distributed systems because it looks like health. Don't let one stalled inference kill your swarm.