The No-Ops Threshold: Designing Self-Healing Infrastructure That Doesn't Page a Human

Where to draw the line between automated recovery and human intervention

by
The No-Ops Threshold: Designing Self-Healing Infrastructure That Doesn't Page a Human

The No-Ops Threshold: Designing Self-Healing Infrastructure That Doesn't Page a Human

Every operator knows the pain of a 3 AM page for a service that automatically recovered five minutes later. The problem isn't the failure—it's the noise. When infrastructure pages humans for transient blips, trust erodes and fatigue sets in. The goal of this article is to define a clear no-ops threshold: the point at which a system should heal itself without human intervention. We'll explore concrete patterns using systemd, health checks, and circuit breakers, and discuss where to draw the line.

What Is the No-Ops Threshold?

The no-ops threshold is the boundary between automated recovery and human escalation. Below this threshold, infrastructure should self-heal without paging. Above it, a human must be notified. The threshold is not a single metric but a combination of failure duration, frequency, and impact.

For example, a single HTTP 503 for 2 seconds due to a network hiccup is below the threshold. A database connection pool exhaustion lasting 10 minutes is above. The art is in defining these limits for each component.

Patterns for Self-Healing

1. systemd Restart Policies

systemd is the init system on most Linux distributions. Its Restart= directive is the simplest self-healing mechanism. For a service that should always be running, use:

[Service]
Restart=always
RestartSec=3
StartLimitInterval=60
StartLimitBurst=5

This restarts the service after a 3-second pause, but if it crashes 5 times in 60 seconds, systemd stops trying and leaves the unit in a failed state. That's the threshold: transient failures are healed, but persistent crashes escalate.

2. Health Checks with Grace Periods

Health checks are essential, but they must include a grace period to avoid flapping. For a load balancer health check, don't mark a backend as unhealthy on the first failure. Use a sliding window:

  • 3 consecutive failures in 10 seconds → mark unhealthy
  • 2 consecutive successes → mark healthy

This prevents a brief spike from triggering a full rebalance. Implement this in HAProxy or Nginx:

eyJsYW5nIjoiaGFwcm94eSIsInRleHQiOiJiYWNrZW5kIGFwcFxuICAgIG9wdGlvbiBodHRwY2hrIEdFVCAvaGVhbHRoXG4gICAgZGVmYXVsdC1zZXJ2ZXIgaW50ZXIgM3MgZmFsbCAzIHJpc2UgMlxuICAgIHNlcnZlciBhcHAxIDEwLjAuMC4xOjgwODAgY2hlY2sifQ==

3. Circuit Breakers

A circuit breaker protects downstream services from cascading failures. When error rates exceed a threshold, the breaker opens and all requests fail fast. After a timeout, it transitions to half-open, allowing a probe request. If that succeeds, it closes again. This is a classic self-healing pattern.

In code (using a library like Hystrix or a simple implementation):

class CircuitBreaker:
    def __init__(self, threshold=5, timeout=30):
        self.failures = 0
        self.threshold = threshold
        self.timeout = timeout
        self.state = 'closed'
        self.last_failure_time = None

    def call(self, func):
        if self.state == 'open':
            if time() - self.last_failure_time > self.timeout:
                self.state = 'half-open'
            else:
                raise CircuitBreakerOpen()
        try:
            result = func()
            if self.state == 'half-open':
                self.state = 'closed'
                self.failures = 0
            return result
        except Exception as e:
            self.failures += 1
            self.last_failure_time = time()
            if self.failures >= self.threshold:
                self.state = 'open'
            raise

4. Retry with Backoff

Transient failures (network timeouts, temporary resource exhaustion) should be retried with exponential backoff. The key is to cap the number of retries and the total time. For example, retry up to 5 times with backoff starting at 100ms, doubling each time, with jitter. If all retries fail, escalate.

import time
import random

def retry_with_backoff(func, max_retries=5, base_delay=0.1):
    for attempt in range(max_retries):
        try:
            return func()
        except Exception as e:
            if attempt == max_retries - 1:
                raise
            delay = base_delay * (2 ** attempt) + random.uniform(0, 0.1)
            time.sleep(delay)

Where to Draw the Line

Not every failure should self-heal. Some require human judgment:

  • Resource exhaustion: Disk full, memory leak. Self-healing via restart may work temporarily, but root cause needs investigation. Page after 2 restarts in 1 hour.
  • Data corruption: A corrupted database index cannot be healed by restart. Page immediately.
  • Security breaches: Any sign of intrusion must page a human.
  • Configuration errors: A service that fails due to a config change should not auto-restart; it should stay down and page.

A good rule of thumb: if the recovery action is deterministic and safe (restart, retry, failover), automate it. If it requires analysis or changes state (data repair, config update), page.

Tooling Integration

Combine these patterns into a cohesive system:

  • Use systemd for service-level restarts.
  • Use Prometheus with alerting rules that have for: duration to avoid flapping.
  • Use HAProxy or Envoy for health check-based load balancing.
  • Use Consul or etcd for distributed health checks and leader election.

Example Prometheus alert rule with duration:

groups:
  - name: self-healing
    rules:
      - alert: HighErrorRate
        expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05
        for: 2m
        annotations:
          summary: "High HTTP error rate"

The for: 2m ensures that a brief spike doesn't fire the alert. If the spike persists, it's a real problem.

Case Study: Reducing Pager Load by 80%

At a previous company, we ran a fleet of 200 microservices. Pages were frequent (10-15 per week). We implemented:

  • systemd restart policies for all stateless services.
  • Circuit breakers for all HTTP clients.
  • Health check grace periods in the load balancer.
  • Prometheus alerting with for: durations.

Result: pages dropped to 2-3 per week. Those remaining were genuine issues: disk full, bug in new release, config error. The team regained trust in alerts.

Final Thoughts

The no-ops threshold is a design decision, not a fixed number. Start conservative—page more—then tune down as you gain confidence. Document every automated recovery action. Monitor the self-healing itself: if a service restarts 100 times a day, that's a symptom of a deeper problem. The goal is not zero pages, but meaningful pages.

Self-healing infrastructure buys you sleep, but it also buys you time to focus on architecture, not firefighting. Draw your threshold wisely.

#autonomous-systems#devops#infrastructure#monitoring#self-healing
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.