Self-Healing Infrastructure: Keeping a Fleet of Services Alive Without an On-Call Human

Automated recovery loops that eliminate pager fatigue

by
Self-Healing Infrastructure: Keeping a Fleet of Services Alive Without an On-Call Human

Self-Healing Infrastructure: Keeping a Fleet of Services Alive Without an On-Call Human

You know the drill: 3 AM, your phone buzzes, a service is down. You SSH in, restart a process, go back to sleep. Repeat next week. This is not engineering — it’s janitorial work. The fix is not more humans; it’s self-healing infrastructure.

I run a fleet of on-prem services — inference engines, embedding pipelines, vector databases, agent orchestrators — all without a dedicated on-call rotation. When something fails, the system fixes itself. Here’s exactly how.

The Three Pillars of Self-Healing

  1. Health-aware supervision — The init system must know when a service is dead, not just when its PID disappears.
  2. Graceful degradation — Downstream services must handle upstream failures without cascading.
  3. Automated recovery with backoff — Blind restarts make things worse. Exponential backoff and circuit breakers prevent thundering herds.

Let’s walk through each with real tools.

Pillar 1: systemd Health Checks Beyond PID Files

systemd’s default Type=simple only tracks the PID. If the process hangs, systemd thinks everything is fine. You need Type=notify with WatchdogSec and a health check script.

Here’s a unit file for a Python inference service using sd_notify:

[Unit]
Description=LLM Inference Service
After=network.target

[Service]
Type=notify
ExecStart=/usr/local/bin/llm-server --config /etc/llm/config.yaml
WatchdogSec=30
Restart=on-failure
RestartSec=5
StartLimitBurst=3
StartLimitInterval=60

[Install]
WantedBy=multi-user.target

In your service code, call sd_notify("WATCHDOG=1") every 10 seconds. If it misses two intervals, systemd kills and restarts the process. No custom scripts, no cron jobs.

For services that can’t be modified (e.g., closed-source binaries), use an ExecStartPost script that checks the service’s HTTP health endpoint and exits non-zero if unhealthy:

ExecStartPost=/usr/local/bin/health-check.sh http://localhost:8080/health

If the health check fails, systemd marks the unit as failed and triggers Restart=on-failure.

Pillar 2: Graceful Degradation with Circuit Breakers

A restart storm happens when a downstream database goes down and every upstream service crashes on connection failure. Then they all restart simultaneously, hit the still-down database, and crash again. The result: thundering herd and total outage.

Solution: circuit breakers in every client. I use pybreaker for Python services and hystrix-go for Go services. Configuration example with pybreaker:

import pybreaker
import requests

breaker = pybreaker.CircuitBreaker(fail_max=3, reset_timeout=30)

def fetch_embeddings(text):
    @breaker
    def call():
        resp = requests.post("http://embedding-service:8000/embed", json={"text": text}, timeout=5)
        return resp.json()
    try:
        return call()
    except pybreaker.CircuitBreakerError:
        # Return stale embeddings or empty list
        return []

When the circuit opens, the caller immediately returns a fallback instead of hammering a dead service. The service stays up, logs the degraded state, and waits for the circuit to half-open after 30 seconds. This prevents cascading failures.

Pillar 3: Exponential Backoff and Jitter

systemd’s RestartSec can be static, but for services that depend on others (e.g., a vector DB that needs postgres), you want backoff. Use RestartSec with a script that sleeps exponentially:

#!/bin/bash
# /usr/local/bin/backoff-restart.sh
MAX_BACKOFF=60
sleep $(( RANDOM % MAX_BACKOFF + 1 ))

Then in the unit:

ExecStartPre=/usr/local/bin/backoff-restart.sh

This adds jitter. Without jitter, all services restart at the same time, causing load spikes. With jitter, they trickle back.

Observability: The Feedback Loop

Self-healing without visibility is blind. You need to know when a recovery happened and why. I ship all recovery events to a local Loki instance via systemd’s journal:

journalctl -u llm-inference.service -n 50 --since "5 minutes ago" | grep -E "(watchdog|failed|restart)" | logger -t self-heal

Then alert on self-heal log entries only if the same service restarts more than 5 times in an hour. That’s a real problem, not a transient blip.

Real-World Example: Embedding Pipeline Recovery

My embedding pipeline has three stages: (1) text chunker, (2) BGE-M3 inference on GPU, (3) Qdrant upsert. If the GPU OOMs, systemd restarts the inference service. The chunker has a circuit breaker: if it can’t reach inference, it buffers chunks to disk and retries with backoff. The Qdrant client also has a circuit breaker: if upsert fails, it retries up to 3 times with exponential backoff, then falls back to a local SQLite cache.

Result: the pipeline never fully stops. Chunks accumulate, services recover, and the backlog catches up. I get a single log line: "Self-heal: inference restarted after OOM, 124 chunks replayed." No pager.

What Not to Automate

Some failures should page a human:

  • Disk full (automated cleanup is risky — you might delete data)
  • Repeated restarts (more than 5 in an hour — indicates code bug)
  • Certificate expiry (automated renewal is fine, but if it fails, you need a human)

For these, use a separate alert route (e.g., email, Slack) with a low priority. The goal is to eliminate 95% of pages, not all of them.

The Stack

All tools mentioned are open-source and self-hosted:

  • systemd (v247+) — init system with watchdog
  • pybreaker / hystrix-go — circuit breakers
  • Loki + Promtail — log aggregation
  • Postgres — used as a control plane for storing circuit breaker state and health check configs

No SaaS, no vendor lock-in. Your infrastructure stays sovereign.

Final Thoughts

Self-healing infrastructure is not about AI magic. It’s about engineering discipline: health checks, circuit breakers, backoff, and observability. Implement these three pillars, and you can sleep through the night while your fleet fixes itself.

Stop being a janitor. Start building systems that heal.

#infrastructure#observability#on-prem#self-healing#systemd
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.