Chaos Engineering for Agent Swarms: Killing Postgres Connections Safely
A methodology for fault injection that validates resilience without collateral damage.
Agent systems that coordinate through a shared Postgres instance inherit a single point of failure that is easy to ignore during development. When a swarm of agents reads task state, writes results, and claims work via row locks, a dropped connection can cascade into duplicated work, stuck transactions, or silent data loss. Chaos engineering offers a way to find those failure modes before production does, but only if the experiments are designed to be safe, bounded, and reversible.
Why Postgres connection loss is a special case
Most fault injection targets process crashes or network partitions. Postgres connection loss sits in a middle ground: the database stays up, but a specific client's session dies. That distinction matters because the failure is asymmetric. One agent loses its connection while others continue. The surviving agents may observe stale locks, phantom rows, or partial writes from the failed peer.
A common pattern in agent swarms is to use advisory locks or SELECT ... FOR UPDATE SKIP LOCKED to claim tasks. When a connection drops mid-transaction, Postgres rolls back the transaction and releases locks. But the agent process may not know this immediately. It might retry with a stale task ID, or worse, assume the write succeeded and move on.
-- Typical task claim pattern in agent swarms
BEGIN;
SELECT id, payload FROM tasks
WHERE status = 'pending'
ORDER BY created_at
FOR UPDATE SKIP LOCKED
LIMIT 1;
-- agent processes task, then:
UPDATE tasks SET status = 'done' WHERE id = $1;
COMMIT;If the connection dies between the SELECT and the UPDATE, the lock is released and another agent can claim the same task. Without idempotency, that task runs twice. Chaos experiments should expose exactly this window.
Designing safe experiments
Safety in chaos engineering means three things: the blast radius is bounded, the experiment is reversible, and the system under test is isolated from real users. For agent swarms, that usually means a dedicated test namespace with synthetic tasks and a Postgres instance that mirrors production schema but holds no real data.
The experiment itself should be defined as a hypothesis. For example: "If an agent loses its Postgres connection during task processing, the swarm will not duplicate work and will recover within two retry cycles." The hypothesis forces you to specify observable outcomes before you inject anything.
A minimal fault injector for Postgres connections can be built with a proxy that sits between the agent and the database. The proxy terminates connections on command, optionally after a delay or when a specific query pattern is seen.
# Simplified TCP proxy that can kill Postgres connections on demand
import asyncio
import signal
class FaultProxy:
def __init__(self, listen_port, target_host, target_port):
self.listen_port = listen_port
self.target = (target_host, target_port)
self.connections = set()
self.kill_switch = False
async def handle_client(self, reader, writer):
self.connections.add(writer)
try:
target_reader, target_writer = await asyncio.open_connection(*self.target)
async def pipe(src, dst):
while True:
data = await src.read(4096)
if not data:
break
if self.kill_switch:
dst.close()
return
dst.write(data)
await dst.drain()
await asyncio.gather(
pipe(reader, target_writer),
pipe(target_reader, writer)
)
finally:
self.connections.discard(writer)
writer.close()
def kill_all(self):
self.kill_switch = True
for writer in list(self.connections):
writer.close()This proxy is deliberately simple. It does not parse the Postgres wire protocol, so it cannot target specific queries. That is a feature for a first experiment: you want to kill connections indiscriminately and observe how the swarm reacts. Later experiments can add protocol awareness to kill only during specific transaction phases.
Instrumenting the swarm for observability
Fault injection without observability is just breakage. Before running any experiment, the agent swarm needs to emit structured events for connection state, task claims, and retry attempts. A common approach is to use a correlation ID per task and log every state transition.
{
"event": "task_claim_attempt",
"task_id": "t-1234",
"agent_id": "agent-7",
"correlation_id": "c-5678",
"timestamp": "2025-01-15T10:00:00Z",
"connection_state": "active"
}When a connection is killed, the logs should show a gap: the agent attempted a claim, then no completion event, then a retry with the same correlation ID. If the retry uses a new correlation ID, you have a traceability problem that will make debugging harder.
The observability layer should also capture Postgres-side state. Query pg_stat_activity before and after the experiment to see if connections are properly closed. Check for idle-in-transaction sessions that might indicate a leaked connection.
SELECT pid, state, query_start, state_change
FROM pg_stat_activity
WHERE application_name = 'agent-swarm'
ORDER BY state_change DESC;Running the experiment
Start with a single agent and a single connection. Kill the connection while the agent is idle, then while it is mid-transaction. Observe how the agent detects the failure. Does it rely on a TCP timeout, a Postgres error code, or a health check? Each detection method has different latency and reliability characteristics.
Next, scale to a small swarm: three to five agents processing synthetic tasks. Kill all connections simultaneously. This tests whether the swarm has a thundering herd problem when every agent reconnects at once. A common mitigation is exponential backoff with jitter, but the experiment should confirm that the backoff is actually implemented and effective.
# Exponential backoff with jitter for reconnection
import random
import time
def reconnect_with_backoff(attempt, base_delay=0.1, max_delay=10.0):
delay = min(max_delay, base_delay * (2 ** attempt))
jitter = random.uniform(0, delay * 0.1)
time.sleep(delay + jitter)
# attempt connectionIf the swarm uses a connection pool, the experiment should also test pool exhaustion. Kill connections faster than the pool can replenish them. Observe whether the pool blocks, throws errors, or silently drops requests. Pool behavior under stress is often undocumented and varies between libraries.
Guardrails and abort conditions
Every chaos experiment needs an abort condition. For Postgres connection loss, the abort condition might be: if more than half of the agents fail to recover within a defined window, stop the experiment and restore connections. The window should be based on the system's service level objective, not on guesswork.
The abort mechanism should be independent of the system under test. A separate control plane that monitors agent health and can trigger the proxy to stop killing connections is safer than relying on the agents themselves to signal distress.
# Experiment definition with abort condition
experiment:
name: postgres-connection-kill
duration: 300s
target:
type: tcp-proxy
port: 5433
fault:
action: kill-connections
interval: 30s
abort:
condition: "healthy_agents < 50%"
window: 60s
rollback:
action: restore-connectionsThis YAML is illustrative; the actual implementation depends on your orchestration tooling. The key point is that the abort condition is evaluated externally and the rollback is automatic.
What to look for in results
The most valuable outcome of a Postgres connection chaos experiment is not a pass or fail. It is a map of failure modes. Look for:
- Duplicate task execution: two agents completing the same task ID.
- Orphaned locks: advisory locks held by dead connections.
- Retry storms: reconnect attempts that overwhelm the database.
- Silent data loss: writes that the agent believed succeeded but were rolled back.
- Partial state: tasks stuck in a processing state with no owner.
Each of these maps to a specific code path. Duplicate execution usually means missing idempotency keys. Orphaned locks mean the agent is not using session-level advisory locks correctly. Retry storms mean backoff is missing or misconfigured.
Iterating on the experiment
Once the basic experiment is stable, add complexity. Kill connections during specific transaction phases by parsing the Postgres wire protocol. Introduce network latency alongside connection kills to simulate a degraded network. Test recovery when the database itself restarts, not just the client connections.
The goal is not to break everything at once. It is to build a library of fault injection scenarios that cover the realistic failure space. Each scenario should have a clear hypothesis, a bounded blast radius, and an automated rollback. Over time, these scenarios become regression tests that run in CI against a disposable Postgres instance.
Agent swarms are distributed systems, and distributed systems fail in distributed ways. Chaos engineering for Postgres connections is a narrow but high-value slice of that failure space. Done safely, it turns unknown-unknowns into documented, tested behavior.
A note on tooling
There are existing chaos engineering tools that can inject network faults, but few understand the semantics of Postgres connections or agent task claims. A custom proxy is often the right level of abstraction because it can be extended to understand the specific protocol patterns your swarm uses. Start simple, instrument heavily, and expand the fault surface only after the basic experiment is reliable.
The methodology matters more than the tool. Define the hypothesis, bound the blast radius, instrument the system, run the experiment, and feed the results back into the design. That loop is what makes chaos engineering useful rather than destructive.