systemd RuntimeMaxSec Ignored for Oneshot Units: What Caps Them Instead

Why a oneshot service can outlive its RuntimeMaxSec, and which knobs actually enforce a time limit

by

A common source of confusion in systemd administration is the behavior of RuntimeMaxSec on Type=oneshot units. An operator adds RuntimeMaxSec=300 to a oneshot service, expecting the unit to be killed if it runs longer than five minutes. The service runs for hours. No warning, no error, no kill. The directive is silently ignored.

This is not a bug. It is a consequence of how systemd defines runtime and which unit types have a runtime clock at all. Understanding the distinction prevents a class of production incidents where a stuck oneshot unit blocks a boot sequence, a dependency chain, or a maintenance window.

What RuntimeMaxSec actually measures

RuntimeMaxSec= is documented as the maximum time a service is allowed to run. The key word is run. In systemd's state machine, a unit is considered running only after it has entered the active state. For Type=simple, Type=exec, Type=notify, and Type=forking services, this happens shortly after the main process is spawned. The runtime clock starts then and RuntimeMaxSec can fire.

For Type=oneshot, the unit never enters the active state while the process is executing. It transitions from activating to active only after the main process exits successfully. During execution, the unit is in activating. The runtime clock, which measures time in the active state, never starts. Therefore RuntimeMaxSec has nothing to measure and is ignored.

This is stated, somewhat obliquely, in the systemd.service man page: RuntimeMaxSec= applies to units that are in the running state. Oneshot units are not running; they are activating. The directive is accepted by the parser, stored in the unit structure, and then never consulted for oneshot units. No log message is emitted because from systemd's perspective there is nothing wrong.

What caps oneshot units instead

The directive that actually applies to oneshot units is TimeoutStartSec=. This timeout covers the entire activation phase, including the execution of the oneshot process. If the process does not exit before TimeoutStartSec elapses, systemd considers the start operation failed and terminates the unit.

[Unit]
Description=Long-running maintenance task

[Service]
Type=oneshot
ExecStart=/usr/local/bin/maintenance.sh
TimeoutStartSec=300
RemainAfterExit=no

With this configuration, the maintenance script is killed if it runs longer than five minutes. The unit enters a failed state. RuntimeMaxSec is not needed and would have no effect.

The default for TimeoutStartSec is 90 seconds unless overridden by DefaultTimeoutStartSec= in /etc/systemd/system.conf. For oneshot units that perform real work, this default is often too short, and operators frequently set it explicitly. The common mistake is to set RuntimeMaxSec instead, believing it is the more specific directive for runtime limits. It is not.

Why the confusion persists

The naming is misleading. RuntimeMaxSec sounds like it should apply to any service that runs, and a oneshot unit does run a process. But systemd's internal model distinguishes between the activation phase and the running phase, and the directive names follow that model rather than everyday language.

Documentation contributes to the problem. The man page for systemd.service describes RuntimeMaxSec in a section that does not prominently exclude oneshot units. The exclusion is implied by the definition of the running state, which is described elsewhere. An operator reading only the RuntimeMaxSec entry can reasonably conclude it applies.

A second source of confusion is that TimeoutStartSec is often described as a startup timeout, suggesting it covers only the time to launch a process, not the time the process runs. For Type=simple services, this is effectively true: the process is considered started as soon as it is spawned, so TimeoutStartSec covers little. For Type=oneshot, the process is not considered started until it exits, so TimeoutStartSec covers the entire execution. The same directive has very different scope depending on unit type.

Interaction with RemainAfterExit

RemainAfterExit=yes changes the picture slightly. With this setting, a oneshot unit that exits successfully remains in the active state. The runtime clock starts at that point. RuntimeMaxSec then applies to the time the unit spends in active after the process has exited.

This is rarely useful. The process is already gone; there is nothing to kill. RuntimeMaxSec firing on a RemainAfterExit oneshot unit would simply transition it to failed, which might be desired if the unit is meant to represent a lease or a lock with an expiry. But it does not cap the execution of the oneshot process itself.

[Service]
Type=oneshot
ExecStart=/usr/local/bin/acquire-lock.sh
RemainAfterExit=yes
RuntimeMaxSec=3600

In this configuration, the lock acquisition script runs under TimeoutStartSec (default 90 seconds). Once it exits successfully, the unit remains active. After one hour, RuntimeMaxSec fires and the unit is marked failed. The lock is not released by systemd; the unit's ExecStop would need to handle that, and ExecStop runs on the transition out of active.

What actually terminates a stuck oneshot process

When TimeoutStartSec expires, systemd sends SIGTERM to the main process. If the process does not exit within TimeoutStopSec (which for the start timeout is governed by the same value unless overridden), systemd sends SIGKILL. The unit is then placed in a failed state.

The kill signal is sent to the process group of the main process, not just the main process itself, provided KillMode= is set appropriately. The default KillMode=control-group sends the signal to all processes in the unit's cgroup. This matters for oneshot units that spawn children: a shell script that backgrounds a worker will have that worker killed as well, because it remains in the unit's cgroup.

[Service]
Type=oneshot
ExecStart=/usr/local/bin/orchestrate.sh
TimeoutStartSec=600
TimeoutStopSec=30
KillMode=control-group

With this configuration, the orchestrator has ten minutes to complete. If it exceeds that, systemd sends SIGTERM to the cgroup, waits thirty seconds, then sends SIGKILL. Any child processes are included.

A subtlety: if the oneshot process daemonizes itself and the daemon escapes the cgroup, systemd cannot kill it. This is one reason Type=oneshot is a poor fit for anything that forks into the background. The correct type for a forking service is Type=forking, which has different timeout semantics and does start the runtime clock.

Diagnosing silent ignores

When a timeout directive appears to have no effect, the first step is to check the unit type and the state transitions. systemctl show -p Type -p TimeoutStartUSec -p RuntimeMaxUSec -p ActiveEnterTimestamp -p InactiveEnterTimestamp <unit> reveals what systemd believes.

systemctl show -p Type -p TimeoutStartUSec -p RuntimeMaxUSec myservice.service

If Type=oneshot and RuntimeMaxUSec is non-zero, the directive is set but inert. If TimeoutStartUSec is at the default 90 seconds and the process is expected to run longer, that is the actual cap.

systemd-analyze can also help. systemd-analyze verify <unit> checks for syntax errors but does not flag semantically inert directives. There is no built-in warning for RuntimeMaxSec on a oneshot unit. This is a known gap; patches to add such a warning have been discussed but not merged as of recent systemd releases.

Patterns that avoid the trap

Several patterns reduce the chance of relying on an inert directive.

First, treat TimeoutStartSec as the primary timeout for oneshot units. Set it explicitly and document why. Avoid setting RuntimeMaxSec on oneshot units unless RemainAfterExit=yes is in use and the intent is to expire the active state.

Second, for long-running tasks, consider whether Type=oneshot is the right choice. If the task is a daemon that should run indefinitely, Type=simple or Type=notify is appropriate, and RuntimeMaxSec works as expected. If the task is a batch job that should complete and exit, Type=oneshot is correct, and TimeoutStartSec is the cap.

Third, wrap the work in a script that enforces its own timeout. This provides a second layer independent of systemd's state machine and makes the intent explicit in the unit file.

#!/bin/bash
set -euo pipefail
timeout --signal=TERM --kill-after=30s 600 /usr/local/bin/real-work.sh

With this wrapper, the timeout command enforces a ten-minute limit regardless of systemd's configuration. The unit's TimeoutStartSec should be set slightly higher to avoid a race between the two mechanisms.

Fourth, monitor for units that remain in activating for unusually long periods. A oneshot unit stuck in activating is either running longer than expected or has hit a timeout that is not configured. Alerting on ActiveState=activating with a duration threshold catches both cases.

RuntimeMaxSec is not the only directive whose applicability depends on unit type. WatchdogSec= applies to services that support the watchdog protocol, typically Type=notify or Type=notify-reload. It is ignored for oneshot units because there is no running process to ping the watchdog.

Restart= is ignored for oneshot units. A oneshot unit that fails does not restart automatically; the failure propagates to dependents. This is by design, but it surprises operators who expect restart behavior from all service types.

KillSignal=, KillMode=, and TimeoutStopSec= apply to the stop operation, which for oneshot units is triggered by systemctl stop or by dependency ordering. They do not cap the start operation. TimeoutStartSec is the only directive that caps the activation phase.

The mental model to keep

systemd's service types map to distinct state machine paths. The directives that apply to each path are determined by which states the unit passes through and how long it spends in each.

For oneshot units, the relevant states are activating and active. Time in activating is bounded by TimeoutStartSec. Time in active is bounded by RuntimeMaxSec, but only if the unit remains active, which requires RemainAfterExit=yes. The execution of the oneshot process occurs in activating, so TimeoutStartSec is the cap that matters.

For simple, exec, notify, and forking units, the process runs in active. RuntimeMaxSec bounds that. TimeoutStartSec bounds only the transition into active, which is typically brief.

Once this model is internalized, the behavior of each directive becomes predictable. The silent ignore is no longer mysterious; it is the expected result of applying a running-state directive to a unit that never enters the running state.

Practical checklist

When configuring a oneshot unit with a time limit:

  • Set TimeoutStartSec= to the maximum acceptable execution time.
  • Do not rely on RuntimeMaxSec= unless RemainAfterExit=yes is set and the intent is to expire the active state.
  • Set TimeoutStopSec= to bound the termination grace period.
  • Use KillMode=control-group to ensure child processes are terminated.
  • Consider an external timeout wrapper for defense in depth.
  • Verify with systemctl show that the effective values match expectations.

When a timeout appears to be ignored, check the unit type first. The directive may be valid but inapplicable. systemd does not warn about this, so the operator must know the state machine well enough to recognize the mismatch. That knowledge is the difference between a unit that fails fast under a stuck process and one that hangs indefinitely, blocking everything that depends on it.

#init-systems#linux#oneshot#service-management#systemd#timeouts
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.