Self-Healing Infrastructure: Keeping a Fleet of Services Alive Without an On-Call Human
Automated recovery loops that eliminate pager fatigue
Self-Healing Infrastructure: Keeping a Fleet of Services Alive Without an On-Call Human
You know the drill: 3 AM, your phone buzzes, a service is down. You SSH in, restart a process, go back to sleep. Repeat next week. This is not engineering — it’s janitorial work. The fix is not more humans; it’s self-healing infrastructure.
I run a fleet of on-prem services — inference engines, embedding pipelines, vector databases, agent orchestrators — all without a dedicated on-call rotation. When something fails, the system fixes itself. Here’s exactly how.
The Three Pillars of Self-Healing
- Health-aware supervision — The init system must know when a service is dead, not just when its PID disappears.
- Graceful degradation — Downstream services must handle upstream failures without cascading.
- Automated recovery with backoff — Blind restarts make things worse. Exponential backoff and circuit breakers prevent thundering herds.
Let’s walk through each with real tools.
Pillar 1: systemd Health Checks Beyond PID Files
systemd’s default Type=simple only tracks the PID. If the process hangs, systemd thinks everything is fine. You need Type=notify with WatchdogSec and a health check script.
Here’s a unit file for a Python inference service using sd_notify:
[Unit]
Description=LLM Inference Service
After=network.target
[Service]
Type=notify
ExecStart=/usr/local/bin/llm-server --config /etc/llm/config.yaml
WatchdogSec=30
Restart=on-failure
RestartSec=5
StartLimitBurst=3
StartLimitInterval=60
[Install]
WantedBy=multi-user.targetIn your service code, call sd_notify("WATCHDOG=1") every 10 seconds. If it misses two intervals, systemd kills and restarts the process. No custom scripts, no cron jobs.
For services that can’t be modified (e.g., closed-source binaries), use an ExecStartPost script that checks the service’s HTTP health endpoint and exits non-zero if unhealthy:
ExecStartPost=/usr/local/bin/health-check.sh http://localhost:8080/healthIf the health check fails, systemd marks the unit as failed and triggers Restart=on-failure.
Pillar 2: Graceful Degradation with Circuit Breakers
A restart storm happens when a downstream database goes down and every upstream service crashes on connection failure. Then they all restart simultaneously, hit the still-down database, and crash again. The result: thundering herd and total outage.
Solution: circuit breakers in every client. I use pybreaker for Python services and hystrix-go for Go services. Configuration example with pybreaker:
import pybreaker
import requests
breaker = pybreaker.CircuitBreaker(fail_max=3, reset_timeout=30)
def fetch_embeddings(text):
@breaker
def call():
resp = requests.post("http://embedding-service:8000/embed", json={"text": text}, timeout=5)
return resp.json()
try:
return call()
except pybreaker.CircuitBreakerError:
# Return stale embeddings or empty list
return []When the circuit opens, the caller immediately returns a fallback instead of hammering a dead service. The service stays up, logs the degraded state, and waits for the circuit to half-open after 30 seconds. This prevents cascading failures.
Pillar 3: Exponential Backoff and Jitter
systemd’s RestartSec can be static, but for services that depend on others (e.g., a vector DB that needs postgres), you want backoff. Use RestartSec with a script that sleeps exponentially:
#!/bin/bash
# /usr/local/bin/backoff-restart.sh
MAX_BACKOFF=60
sleep $(( RANDOM % MAX_BACKOFF + 1 ))Then in the unit:
ExecStartPre=/usr/local/bin/backoff-restart.shThis adds jitter. Without jitter, all services restart at the same time, causing load spikes. With jitter, they trickle back.
Observability: The Feedback Loop
Self-healing without visibility is blind. You need to know when a recovery happened and why. I ship all recovery events to a local Loki instance via systemd’s journal:
journalctl -u llm-inference.service -n 50 --since "5 minutes ago" | grep -E "(watchdog|failed|restart)" | logger -t self-healThen alert on self-heal log entries only if the same service restarts more than 5 times in an hour. That’s a real problem, not a transient blip.
Real-World Example: Embedding Pipeline Recovery
My embedding pipeline has three stages: (1) text chunker, (2) BGE-M3 inference on GPU, (3) Qdrant upsert. If the GPU OOMs, systemd restarts the inference service. The chunker has a circuit breaker: if it can’t reach inference, it buffers chunks to disk and retries with backoff. The Qdrant client also has a circuit breaker: if upsert fails, it retries up to 3 times with exponential backoff, then falls back to a local SQLite cache.
Result: the pipeline never fully stops. Chunks accumulate, services recover, and the backlog catches up. I get a single log line: "Self-heal: inference restarted after OOM, 124 chunks replayed." No pager.
What Not to Automate
Some failures should page a human:
- Disk full (automated cleanup is risky — you might delete data)
- Repeated restarts (more than 5 in an hour — indicates code bug)
- Certificate expiry (automated renewal is fine, but if it fails, you need a human)
For these, use a separate alert route (e.g., email, Slack) with a low priority. The goal is to eliminate 95% of pages, not all of them.
The Stack
All tools mentioned are open-source and self-hosted:
- systemd (v247+) — init system with watchdog
- pybreaker / hystrix-go — circuit breakers
- Loki + Promtail — log aggregation
- Postgres — used as a control plane for storing circuit breaker state and health check configs
No SaaS, no vendor lock-in. Your infrastructure stays sovereign.
Final Thoughts
Self-healing infrastructure is not about AI magic. It’s about engineering discipline: health checks, circuit breakers, backoff, and observability. Implement these three pillars, and you can sleep through the night while your fleet fixes itself.
Stop being a janitor. Start building systems that heal.