The No-Ops Threshold: Designing Self-Healing Infrastructure That Doesn't Page a Human
Where to draw the line between automated recovery and human intervention
The No-Ops Threshold: Designing Self-Healing Infrastructure That Doesn't Page a Human
Every operator knows the pain of a 3 AM page for a service that automatically recovered five minutes later. The problem isn't the failure—it's the noise. When infrastructure pages humans for transient blips, trust erodes and fatigue sets in. The goal of this article is to define a clear no-ops threshold: the point at which a system should heal itself without human intervention. We'll explore concrete patterns using systemd, health checks, and circuit breakers, and discuss where to draw the line.
What Is the No-Ops Threshold?
The no-ops threshold is the boundary between automated recovery and human escalation. Below this threshold, infrastructure should self-heal without paging. Above it, a human must be notified. The threshold is not a single metric but a combination of failure duration, frequency, and impact.
For example, a single HTTP 503 for 2 seconds due to a network hiccup is below the threshold. A database connection pool exhaustion lasting 10 minutes is above. The art is in defining these limits for each component.
Patterns for Self-Healing
1. systemd Restart Policies
systemd is the init system on most Linux distributions. Its Restart= directive is the simplest self-healing mechanism. For a service that should always be running, use:
[Service]
Restart=always
RestartSec=3
StartLimitInterval=60
StartLimitBurst=5This restarts the service after a 3-second pause, but if it crashes 5 times in 60 seconds, systemd stops trying and leaves the unit in a failed state. That's the threshold: transient failures are healed, but persistent crashes escalate.
2. Health Checks with Grace Periods
Health checks are essential, but they must include a grace period to avoid flapping. For a load balancer health check, don't mark a backend as unhealthy on the first failure. Use a sliding window:
- 3 consecutive failures in 10 seconds → mark unhealthy
- 2 consecutive successes → mark healthy
This prevents a brief spike from triggering a full rebalance. Implement this in HAProxy or Nginx:
eyJsYW5nIjoiaGFwcm94eSIsInRleHQiOiJiYWNrZW5kIGFwcFxuICAgIG9wdGlvbiBodHRwY2hrIEdFVCAvaGVhbHRoXG4gICAgZGVmYXVsdC1zZXJ2ZXIgaW50ZXIgM3MgZmFsbCAzIHJpc2UgMlxuICAgIHNlcnZlciBhcHAxIDEwLjAuMC4xOjgwODAgY2hlY2sifQ==3. Circuit Breakers
A circuit breaker protects downstream services from cascading failures. When error rates exceed a threshold, the breaker opens and all requests fail fast. After a timeout, it transitions to half-open, allowing a probe request. If that succeeds, it closes again. This is a classic self-healing pattern.
In code (using a library like Hystrix or a simple implementation):
class CircuitBreaker:
def __init__(self, threshold=5, timeout=30):
self.failures = 0
self.threshold = threshold
self.timeout = timeout
self.state = 'closed'
self.last_failure_time = None
def call(self, func):
if self.state == 'open':
if time() - self.last_failure_time > self.timeout:
self.state = 'half-open'
else:
raise CircuitBreakerOpen()
try:
result = func()
if self.state == 'half-open':
self.state = 'closed'
self.failures = 0
return result
except Exception as e:
self.failures += 1
self.last_failure_time = time()
if self.failures >= self.threshold:
self.state = 'open'
raise4. Retry with Backoff
Transient failures (network timeouts, temporary resource exhaustion) should be retried with exponential backoff. The key is to cap the number of retries and the total time. For example, retry up to 5 times with backoff starting at 100ms, doubling each time, with jitter. If all retries fail, escalate.
import time
import random
def retry_with_backoff(func, max_retries=5, base_delay=0.1):
for attempt in range(max_retries):
try:
return func()
except Exception as e:
if attempt == max_retries - 1:
raise
delay = base_delay * (2 ** attempt) + random.uniform(0, 0.1)
time.sleep(delay)Where to Draw the Line
Not every failure should self-heal. Some require human judgment:
- Resource exhaustion: Disk full, memory leak. Self-healing via restart may work temporarily, but root cause needs investigation. Page after 2 restarts in 1 hour.
- Data corruption: A corrupted database index cannot be healed by restart. Page immediately.
- Security breaches: Any sign of intrusion must page a human.
- Configuration errors: A service that fails due to a config change should not auto-restart; it should stay down and page.
A good rule of thumb: if the recovery action is deterministic and safe (restart, retry, failover), automate it. If it requires analysis or changes state (data repair, config update), page.
Tooling Integration
Combine these patterns into a cohesive system:
- Use systemd for service-level restarts.
- Use Prometheus with alerting rules that have
for:duration to avoid flapping. - Use HAProxy or Envoy for health check-based load balancing.
- Use Consul or etcd for distributed health checks and leader election.
Example Prometheus alert rule with duration:
groups:
- name: self-healing
rules:
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05
for: 2m
annotations:
summary: "High HTTP error rate"The for: 2m ensures that a brief spike doesn't fire the alert. If the spike persists, it's a real problem.
Case Study: Reducing Pager Load by 80%
At a previous company, we ran a fleet of 200 microservices. Pages were frequent (10-15 per week). We implemented:
- systemd restart policies for all stateless services.
- Circuit breakers for all HTTP clients.
- Health check grace periods in the load balancer.
- Prometheus alerting with
for:durations.
Result: pages dropped to 2-3 per week. Those remaining were genuine issues: disk full, bug in new release, config error. The team regained trust in alerts.
Final Thoughts
The no-ops threshold is a design decision, not a fixed number. Start conservative—page more—then tune down as you gain confidence. Document every automated recovery action. Monitor the self-healing itself: if a service restarts 100 times a day, that's a symptom of a deeper problem. The goal is not zero pages, but meaningful pages.
Self-healing infrastructure buys you sleep, but it also buys you time to focus on architecture, not firefighting. Draw your threshold wisely.