Disabled Is Not Stopped: How Orchestrators Resurrect Masked systemd Units

The gap between systemd's mask state and an orchestrator's desired-state reconciliation.

by

Masking a systemd unit is one of the few operations on a Linux host that feels genuinely final. systemctl mask foo.service symlinks the unit file to /dev/null, and systemd will refuse to start it, even as a dependency. Operators reach for it when they want something to stay dead: a legacy daemon that conflicts with a replacement, a vendor agent that keeps re-enabling itself, a service that has been migrated off the box entirely.

Then the unit comes back. Not immediately, and not because systemd changed its mind. It comes back because something outside systemd has an opinion about desired state, and that opinion does not include the word "masked."

What mask actually means

It helps to be precise about the three states people conflate:

  • stopped: the unit is not currently running. Nothing prevents it from starting.
  • disabled: the unit will not be started automatically at boot via its [Install] section. It can still be started manually or as a dependency.
  • masked: the unit name resolves to /dev/null, so systemd cannot load a definition for it at all. Any attempt to start it fails with Unit foo.service is masked.

Masking is the strongest of the three. It is also purely a systemd-level concept. The mask lives in the filesystem as a symlink, typically under /etc/systemd/system/foo.service -> /dev/null. Nothing broadcasts that fact to the rest of the machine.

$ systemctl mask nginx.service
Created symlink /etc/systemd/system/nginx.service -> /dev/null.

$ systemctl start nginx.service
Failed to start nginx.service: Unit nginx.service is masked.

That failure is local to systemd. An orchestrator that shells out to systemctl enable --now will see a non-zero exit code, but many orchestrators do not treat that as a fatal condition. They log it, retry on the next reconciliation loop, and move on. The unit stays masked, but the orchestrator keeps trying, and the operator sees a stream of failures that look like noise rather than a policy conflict.

The reconciliation loop does not know about masks

Container orchestrators, configuration-management agents, and node bootstrappers all share a common design: they hold a desired state and periodically drive the host toward it. The mechanism varies, but the shape is the same.

A typical reconciliation step looks like this:

for each unit in desired_units:
    if not unit_is_active(unit):
        enable_and_start(unit)

There is no branch for unit_is_masked. The agent asks systemd whether the unit is active, gets inactive or failed, and proceeds to start it. The start fails because of the mask, the agent records a failure, and the loop runs again. Depending on the agent's backoff, this can repeat every few seconds indefinitely.

The failure mode is not that the orchestrator overrides the mask. It cannot; systemd will not load a masked unit. The failure mode is that the orchestrator never stops trying, and the operator never gets a clean signal that the mask is the reason.

Where the resurrection actually happens

There are several distinct mechanisms by which a masked unit appears to come back. They are worth separating because the fix differs for each.

1. The agent recreates the unit file

Some configuration agents do not just call systemctl. They write the unit file themselves from a template, then run daemon-reload and enable --now. If the agent writes to /etc/systemd/system/foo.service, it overwrites the mask symlink with a regular file. The mask is gone, and the unit starts.

This is the most common cause of "masked units keep coming back." The mask was never durable because the agent owns the path.

# What the agent does, in effect:
cat > /etc/systemd/system/foo.service <<'EOF'
[Unit]
Description=Foo agent
[Service]
ExecStart=/usr/local/bin/foo
[Install]
WantedBy=multi-user.target
EOF
systemctl daemon-reload
systemctl enable --now foo.service

The cat > clobbers the symlink. There is no error, because writing a file over a symlink is a normal filesystem operation.

2. The agent uses a drop-in or a different unit name

Some agents avoid touching the vendor unit and instead install a drop-in under /etc/systemd/system/foo.service.d/override.conf. A drop-in does not un-mask the unit, but it can change ExecStart, Restart=, and dependencies. If the agent also enables a companion unit that Wants=foo.service, the companion will pull foo in. Because foo is masked, the pull fails, but the companion may be configured with Restart=always, producing a restart loop that looks like the masked unit is running.

3. A watchdog or supervisor restarts the wrapper

A common pattern is a small supervisor unit that runs a script, and the script calls systemctl start foo.service. The supervisor is not masked. It keeps running, keeps failing to start foo, and keeps restarting. From the outside, the host is busy doing something with foo even though foo never starts.

4. The orchestrator schedules a pod that binds the same port

In Kubernetes and similar systems, a DaemonSet or a static pod may bind the port the masked unit used to own. The masked unit is not running, but the port is occupied, and the orchestrator's health checks fail. Operators sometimes misread this as the masked unit having restarted.

Why systemctl is-enabled lies to you

A masked unit reports as masked from systemctl is-enabled, but many agents do not call that. They call systemctl is-active, which returns inactive for a masked unit. The distinction matters:

$ systemctl is-enabled foo.service
masked

$ systemctl is-active foo.service
inactive

An agent that only checks is-active cannot distinguish "stopped and startable" from "masked and unstartable." It will attempt the start every time.

Making the mask durable

The reliable fix is to make the mask part of the same desired-state system that the orchestrator reads. If the orchestrator owns the unit, the mask must be expressed in the orchestrator's own configuration, not as a local filesystem operation.

For configuration-management agents, this usually means a resource that explicitly disables and masks the unit, and a guard that prevents the agent from writing the unit file:

# Illustrative shape, not a specific tool's syntax
- name: ensure foo is masked
  systemd:
    name: foo.service
    masked: true
    enabled: false
    state: stopped

The important part is masked: true as a first-class field. If the agent's schema does not have it, the agent will keep fighting the mask.

For container orchestrators, the unit is usually not the right layer. If a node bootstrapper is installing systemd units, the mask should be applied by the same bootstrapper, and the bootstrapper should be idempotent about it. A common pattern is to have the bootstrapper check for an existing mask and skip unit installation if one is present:

if [ -L /etc/systemd/system/foo.service ] && \
   [ "$(readlink /etc/systemd/system/foo.service)" = "/dev/null" ]; then
    echo "foo.service is masked, skipping install"
    exit 0
fi

This is a guard, not a fix. The real fix is to remove foo from the bootstrapper's desired state entirely.

The Restart= trap

A masked unit cannot start, but a unit that Wants= or Requires= a masked unit can still be started. If that unit has Restart=always, systemd will restart it on failure, and each restart will attempt to pull in the masked dependency. The result is a restart loop that consumes CPU and fills the journal, even though the masked unit never runs.

[Unit]
Description=Wrapper
Requires=foo.service
After=foo.service

[Service]
ExecStart=/usr/local/bin/wrapper
Restart=always
RestartSec=5

With foo.service masked, wrapper.service fails immediately, restarts after five seconds, and repeats. The journal shows a steady stream of Failed to start wrapper.service entries, and the underlying cause is the mask on foo.

Diagnosing the loop

When a masked unit appears to be resurrecting, the useful questions are:

  1. Is the unit actually running, or is something else failing because of it?
  2. Who is writing to /etc/systemd/system/foo.service?
  3. Which units have Wants=, Requires=, or After=foo.service?
  4. Is there a supervisor or watchdog with Restart=always that references foo?

systemctl list-dependencies --reverse foo.service shows the units that pull foo in. journalctl -u foo.service shows the start attempts. auditd or a filesystem watcher on /etc/systemd/system shows who is writing the unit file.

$ systemctl list-dependencies --reverse foo.service
foo.service
● └─wrapper.service

If wrapper.service is the only dependent and it has Restart=always, the fix is to stop and mask wrapper as well, or to remove the dependency.

The general principle

Masking is a local, systemd-level assertion. Orchestrators operate on a different layer and do not read systemd's mask state. Any time a host is managed by an agent that owns unit files, a local systemctl mask is a temporary measure at best. The durable fix is to express the intent in the layer that owns the host's desired state.

This is not a systemd bug. It is a consequence of two systems with different notions of authority over the same resource. The operator's job is to decide which system is authoritative and to make the other one defer. A mask that is not reflected in the orchestrator's configuration is a note to self, not a policy.

A checklist for operators

  • Confirm the unit is masked, not just stopped or disabled.
  • Check whether a configuration agent owns the unit file path.
  • Look for Restart=always units that depend on the masked unit.
  • Check for companion units that Wants= the masked unit.
  • Verify that the orchestrator's desired state does not include the unit.
  • If the orchestrator cannot express a mask, remove the unit from its desired state instead.
  • Watch the journal for repeated start attempts, not just for the unit's own logs.

The pattern repeats across tools and years because the underlying tension is structural: local imperative commands versus remote declarative state. Masking is the imperative side. Reconciliation is the declarative side. When they disagree, the declarative side usually wins, because it runs on a timer.

#configuration-management#infrastructure#linux#orchestration#reconciliation#systemd
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.