TL;DR
- A service that spawns an updater and then asks systemd to restart the service will lose the updater. In my lab, the helper was started with
setsid(Python’sstart_new_session=True), and it still never logged its final line. Its cgroup was/system.slice/demo-app.service, the same as the service it was restarting. setsidchanges the process group and session. It does not change the cgroup, and for a service with the defaultKillMode=control-group, the cgroup is what systemd kills on stop.KillMode=processkept the helper alive, but it stayed in the unit’s cgroup. When it was still running, it showed up as part of the new instance. The systemd documentation calls this setting “not recommended!”.- Starting the helper with
systemd-runput it in its own unit (demo-updater.service), and it finished normally. Having the service simply exit and lettingRestart=bring the new version up also worked. - All of this ran on real systemd 257 inside a throwaway namespace on a Linux box. I did not test
systemd --user, macOSlaunchd, containers, or a non-root service.
The symptom
A daemon can update itself in a way that sounds sensible: download the new version, start a small helper, and let the helper restart the service and then check that the new process is healthy. If it is not healthy, the helper rolls back. The last step is the whole point of having a helper, and it is the one I want to see in the log: “restart finished; verification runs now”.
The failure looks like this: the new process comes up fine, and the helper’s log has its first line (“start”) and nothing after. No error, no crash, no partial verification.
Candidates, in the order I would suspect them:
- The helper crashed. There was no error output, and the shell steps are trivial.
- The helper received SIGHUP when its parent exited. That is what
setsidis for, and it was already in the code. - Something sent it a signal when the service restarted. Plausible, and the next section looks for how.
Evidence
A symptom like this is easy to blame on the wrong thing, so I made the smallest reproduction I could. It needs real systemd as the service manager. I booted systemd as PID 1 inside a throwaway PID, mount and cgroup namespace on a Linux box (as root, with unshare; I do not print that setup because it is specific to my machine, and a disposable VM with systemd is the easier route for you). Then I ran four variants with a one-file service.
#!/usr/bin/env python3
"""A tiny daemon that can update itself. SIGUSR1 = "update requested".
usage: app.py <how> how = popen | systemd-run | exit
"""
import os, signal, subprocess, sys, time
HOW = sys.argv[1]
LOG = "/run/demo/log"
def log(msg):
with open(LOG, "a") as f: f.write(f"{time.strftime('%H:%M:%S')} [app pid={os.getpid()}] {msg}\n")
def on_usr1(*_):
log(f"update requested, starting the updater via {HOW}")
if HOW == "popen": # the obvious way: detach into a new session so nothing can touch it
subprocess.Popen(["/run/demo/updater.sh"], start_new_session=True,
stdin=subprocess.DEVNULL, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
elif HOW == "systemd-run": # hand the updater to the service manager as its own unit
subprocess.run(["systemd-run", "--unit=demo-updater", "--collect", "/run/demo/updater.sh"], check=True)
else: # no helper at all: leave, and let the supervisor (Restart=) start the new version
log("exiting; the supervisor restarts me"); os._exit(0)
signal.signal(signal.SIGUSR1, on_usr1)
signal.signal(signal.SIGTERM, lambda *_: (log("SIGTERM received, exiting"), sys.exit(0)))
log("started")
while True: time.sleep(1)
SIGUSR1 stands for “an update was requested”. The helper restarts the service and then writes the line I care about:
#!/bin/bash
log() { echo "$(date +%H:%M:%S) [updater pid=$$] $*" >> /run/demo/log; }
log "start; cgroup=$(cut -d: -f3 /proc/self/cgroup)"
systemctl restart demo-app.service
sleep "${UPDATER_SLEEP:-1}"
log "restart finished; the post-restart verification runs now"
The unit file is generated per trial, so the same service can run with different options:
#!/bin/bash
# usage: unit.sh <popen|systemd-run|exit> [extra [Service] lines...]
HOW=$1; shift
cat > /run/systemd/system/demo-app.service <<EOT
[Unit]
Description=demo app
[Service]
ExecStart=/usr/bin/python3 /run/demo/app.py $HOW
$(printf '%s\n' "$@")
EOT
systemctl daemon-reload
And the trial script resets, runs, and prints the log and the processes left over:
#!/bin/bash
# usage: trial.sh <label> <popen|systemd-run|exit> [extra [Service] lines]
label=$1; shift
echo "=== $label"
systemctl stop demo-app.service demo-updater.service 2>/dev/null; : > /run/demo/log
/run/demo/unit.sh "$@"
systemctl start demo-app.service; sleep 1
systemctl kill -s SIGUSR1 --kill-whom=main demo-app.service
sleep 4
cat /run/demo/log
echo "--- processes still around:"; ps -eo pid,cgroup:40,cmd | grep -E 'updater.sh|app.py' | grep -v grep
Output of the four trials:
=== A: Popen(start_new_session=True), default KillMode
03:08:37 [app pid=92] started
03:08:38 [app pid=92] update requested, starting the updater via popen
03:08:38 [updater pid=95] start; cgroup=/system.slice/demo-app.service
03:08:38 [app pid=92] SIGTERM received, exiting
03:08:38 [app pid=101] started
--- processes still around:
101 0::/system.slice/demo-app.service /usr/bin/python3 /run/demo/app.py popen
=== B: same, KillMode=process
03:08:42 [app pid=130] started
03:08:43 [app pid=130] update requested, starting the updater via popen
03:08:43 [updater pid=133] start; cgroup=/system.slice/demo-app.service
03:08:43 [app pid=130] SIGTERM received, exiting
03:08:43 [app pid=139] started
03:08:44 [updater pid=133] restart finished; the post-restart verification runs now
=== C: systemd-run --unit=demo-updater
03:08:47 [app pid=170] started
03:08:48 [app pid=170] update requested, starting the updater via systemd-run
03:08:48 [updater pid=175] start; cgroup=/system.slice/demo-updater.service
03:08:48 [app pid=170] SIGTERM received, exiting
03:08:48 [app pid=180] started
03:08:49 [updater pid=175] restart finished; the post-restart verification runs now
=== D: app exits, Restart=always
03:08:52 [app pid=212] started
03:08:53 [app pid=212] update requested, starting the updater via exit
03:08:53 [app pid=212] exiting; the supervisor restarts me
03:08:53 [app pid=217] started
(The script also prints a --- processes still around: listing. After trial A it showed only the new app process, so the updater was gone. For B, C and D it showed only the new app process too, because the helper had already finished within the 4-second wait, so I left those listings out. These are the lines from a second session; the first session, which the conclusions below were first drawn from, gave the same shape.)
Reading it:
- A is the bug. The updater logs its start, and the log line shows its cgroup is the service’s own. The old app exits on
SIGTERM, the new one starts, and the updater never prints its last line and does not appear in the process list. The service restarted successfully, but the verification and any rollback that depended on the updater did not happen. - B kept the updater alive. But the same
KillMode=processthat spared it also means anything the service spawns can outlive its stop. - C gave the updater its own cgroup (
/system.slice/demo-updater.service) and it ran to the end. - D removed the helper altogether: the app logged that it would exit, and
Restart=alwaysstarted the new instance.
To see what B costs, I ran it again with the helper sleeping for 30 seconds, and asked systemd for the status of the service while the helper was still running:
● demo-app.service - demo app
Loaded: loaded (/run/systemd/system/demo-app.service; static)
Active: active (running) since Mon 2026-10-05 03:09:11 JST; 3s ago
Main PID: 302 (python3)
CGroup: /system.slice/demo-app.service
├─296 /bin/bash /run/demo/updater.sh
├─302 /usr/bin/python3 /run/demo/app.py popen
└─303 sleep 30
The updater from the previous instance is sitting inside the new instance’s cgroup. A later systemctl stop would only signal the main process; in my run, the sleep 30 stayed alive after I stopped the service. The helper has become part of the next instance’s lifecycle by accident.
I had not captured the signal that killed the updater in A, so I added one line to updater.sh and ran trial A again:
trap 'log "got SIGTERM"; exit 1' TERM # added after the log() definition
03:09:06 [app pid=250] update requested, starting the updater via popen
03:09:06 [updater pid=253] start; cgroup=/system.slice/demo-app.service
03:09:06 [app pid=250] SIGTERM received, exiting
03:09:06 [updater pid=253] got SIGTERM
03:09:06 [app pid=261] started
The updater receives the same SIGTERM as the main process, at the same moment. That fits the documentation below: with KillMode=control-group, every process in the unit’s cgroup is sent SIGTERM first. The updater here is a shell script that can trap it; without the trap, SIGTERM simply ends it.
Why: the documentation says so
From systemd.kill(5): “If set to control-group, all remaining processes in the control group of this unit will be killed on unit stop”. The default is control-group. For process: “only the main process itself is killed (not recommended!)”, and the note that this “allows processes to escape the service manager’s lifecycle and resource management, and to remain running even while their service is considered stopped” (systemd.kill(5)).
So setsid() is the wrong tool for this job. It protects against terminal hangups and process-group signals. As far as the documentation above goes, systemd decides what belongs to a service by its control group, and that is what the log line in trial A shows the helper still inside.
And systemd-run “may be used to create and start a transient .service or .scope unit and run the specified COMMAND in it” (systemd-run(1)). That is why trial C worked: the helper became a separate unit, and the restart of one unit does not stop the other.
The fix, and its alternatives
| Option | What it does | Cost | I ran it |
|---|---|---|---|
systemd-run --unit=… --collect for the helper | The helper gets its own unit and cgroup | Needs permission to call it (I ran as root); unit names must be unique | yes (C) |
The service exits, and the supervisor restarts it (Restart=) | No helper; the new instance does the post-start checks | The check and rollback must live in the new instance, or elsewhere | yes (D) |
KillMode=process | The helper survives | Leaks into the next instance; the docs advise against it | yes (B) |
| A separate updater unit started by a timer or a path unit, not by the service | The updater has no relation to the service’s cgroup | More units | no |
launchd on macOS | The launchd.plist(5) page says that on a job’s death, launchd kills remaining processes with the same process group ID unless AbandonProcessGroup is true | Different mechanism, different escape rules | no, and I do not have a macOS machine in this setup |
For launchd, the page is launchd.plist(5). Because the rule there is about process groups, a setsid’d child would have its own group, so the behaviour may differ from systemd’s. That is an inference from the text; I did not run it.
My pick would be D when the service can verify itself after it starts, and C when an outside observer is needed to verify and roll back. The helper that decides to roll back should not be inside the thing it may have to stop.
When the fix applies
| Situation | Choice |
|---|---|
| Self-updating daemon under systemd, needs a watcher that survives the restart | systemd-run for the watcher |
| The new version can check its own health at start and exit non-zero if it is broken | Exit and let Restart= (with sensible limits) do the work |
| Updates are done by a package manager or a deploy tool outside the service | Leave the service out of it. Do not make the service restart itself |
A user-level service (systemd --user), a container without systemd, or macOS | Not tested here. Check the equivalent rule on that platform |
Habits that bring the bug back
- Relying on
setsid,nohup,disownor a double fork. Symptom: the helper works when you run the service by hand, and disappears under systemd.setsidwas reproduced above (A); I did not test the others. They change process or terminal relations, not the cgroup. - “Fixing” it with
KillMode=process. Symptom: orphaned children after a stop; a helper that shows up in the next instance’s status. Reproduced (B). - Using a fixed unit name for every update. Symptom: the second update fails to start the helper while the first is still around. I did not test the collision;
--collectremoves a finished unit from systemd’s memory, but verify what your systemd version does when the name is still active. - Not logging outside the dying cgroup. Symptom: you see the first line and no last line. Log to somewhere that survives (the journal, or a file the helper opens itself), and log before the risky step.
- Treating “the service restarted” as “the update succeeded”. Symptom: a broken version is running and nothing noticed. The verification step is the part that was lost in A. Put it where it survives.
- Testing the updater only by hand. Symptom: it works in a terminal and not under systemd. A terminal run has no unit cgroup to clean up. Test the update under the real service manager.
Run the four trials
You need a machine where systemd is PID 1 and you can use root. A throwaway VM is best. The scripts write units under /run/systemd/system, which is volatile.
sudo mkdir -p /run/demo && sudo cp app.py updater.sh unit.sh trial.sh /run/demo/ && sudo chmod +x /run/demo/*.shsudo /run/demo/trial.sh "A" popen. Success looks like trial A: no “restart finished” line, and the cgroup shown in the “start” line isdemo-app.service.sudo /run/demo/trial.sh "B" popen KillMode=process: the last line appears.sudo /run/demo/trial.sh "C" systemd-run: the cgroup isdemo-updater.serviceand the last line appears.sudo /run/demo/trial.sh "D" exit Restart=always: the app restarts by itself.- Add
Environment=UPDATER_SLEEP=30as an extra argument to apopen KillMode=processtrial and runsystemctl status demo-app.servicein the first seconds, to see the leaked helper. - Clean up:
sudo systemctl stop demo-app.service demo-updater.service; sudo rm -rf /run/demo /run/systemd/system/demo-app.service; sudo systemctl daemon-reload.
What the throwaway systemd showed, and what it could not
Verified: trials A, B, C, D, the extended B with a sleeping helper, and A with the trap, as printed above, on a Linux 6.12 box with systemd 257 running as PID 1 of a throwaway PID, mount and cgroup namespace, as root. One run each; the extended B once. The quotes from systemd.kill(5), systemd-run(1) and launchd.plist(5) were read at the original pages while writing.
Not verified: whether the updater could survive a SIGTERM handler that ignores the signal until SIGKILL arrives (I only trapped it and exited); systemd --user; a service running as a non-root user (and the permissions systemd-run then needs); containers; other distributions or systemd versions; macOS launchd and whether AbandonProcessGroup matters for a setsid’d child; unit-name collisions for the helper; what happens when the update itself fails and the rollback runs. The log timestamps are from my clock and the one-second sleep is a lab value.
Same cgroup, same fate
A process in the same cgroup is part of the same service, whatever its session or process group says. If the thing must outlive the service’s restart, give it its own unit, or remove the need for it.