TL;DR

  • Replacing a daemon has four separate jobs: new connections should not be refused, in-flight requests should finish, a bad release should be detected and undone, and state should be handed over safely. Each has its own tool.
  • In a hammer test (about 100 new connections per second for 5 seconds, a restart 2 seconds in; two sets of five runs per mode), systemctl restart of a self-binding server produced errors in every run: 4 to 37 refused connections and 1 to 3 resets per run. Starting v2 beside v1 with SO_REUSEPORT and then stopping v1 gave 0 refused, but a reset in 2 of 5 runs in the first set and 3 of 5 in the second. Socket activation (systemd holds the listening socket) gave no errors in all 10 runs.
  • The price of socket activation: connections queue while the new process starts. The slowest request in those runs was 58 to 285 ms (typically about 250), against 1 to 16 ms otherwise.
  • The resets in the overlap runs are the kernel aborting connections that were waiting in the closing listener’s accept queue. With net.ipv4.tcp_migrate_req=1 (Linux 5.14 and later), 0 of 16 runs showed a reset; with the default 0, 5 of 8 runs did (checked in a throwaway network namespace on this box).
  • Draining: a 3-second request in flight survived a restart under the default stop timeout. With TimeoutStopSec=1 it was killed. Sending EXTEND_TIMEOUT_USEC from the draining process kept it alive past that timeout.
  • A release swap by an atomic symlink change, with a health check that insists on the expected version and an automatic rollback, correctly kept a good release and rolled back both a crashing one and one that started but answered wrongly.
  • Everything ran as root on one Linux box under a real systemd 257 in a throwaway namespace. State handoff is only described, not exercised.

What “replace” has to cover

Restarting a service is one command. Replacing it without anyone noticing means answering four questions in order:

  1. When the old process stops listening and the new one has not started yet, where do arriving connections go?
  2. What happens to the requests the old process is still working on?
  3. How do you know the new version is the one serving, and healthy?
  4. If it is not, how do you get the old one back?

State handoff (what the new version needs to know from the old one) is a fifth question that only some services have.

The server I used is small: it accepts a connection, reads a line, “works” for a configurable time, and replies with its version and pid. It stops accepting on SIGTERM, drains, and exits. It can bind its own socket (optionally with SO_REUSEPORT) or take one from systemd, and it can speak enough of sd_notify to say it is ready and to ask for more stop time.

#!/usr/bin/env python3
"""A small line-based server that can be replaced while it serves.
Env: VERSION (label in replies), WORK (seconds of "work" per request), REUSEPORT=1 (bind with SO_REUSEPORT),
     EXTEND=1 (ask the service manager for more stop time while requests are still running).
Under socket activation (LISTEN_FDS) the listening socket is inherited and never re-created."""
import os, signal, socket, sys, threading, time

VERSION = os.environ.get("VERSION", "v1")
WORK = float(os.environ.get("WORK", "0"))
PORT = int(os.environ.get("PORT", "9100"))

def notify(msg):                                   # minimal sd_notify(3)
    path = os.environ.get("NOTIFY_SOCKET")
    if not path: return
    s = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
    s.connect("\0" + path[1:] if path[0] == "@" else path); s.send(msg.encode()); s.close()

def listener():
    if os.environ.get("LISTEN_PID") == str(os.getpid()) and int(os.environ.get("LISTEN_FDS", "0")) >= 1:
        return socket.socket(fileno=3), "inherited from the service manager"
    s = socket.socket(); s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
    if os.environ.get("REUSEPORT"): s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEPORT, 1)
    s.bind(("127.0.0.1", PORT)); s.listen(128)
    return s, "bound by myself"

lock = threading.Lock(); inflight = 0; stopping = False

def handle(c):
    global inflight
    try:
        c.recv(100); time.sleep(WORK); c.sendall(f"{VERSION} {os.getpid()}\n".encode())
    except OSError: pass
    finally:
        c.close()
        with lock: inflight -= 1

def on_term(*_):
    global stopping; stopping = True
signal.signal(signal.SIGTERM, on_term)

srv, how = listener(); srv.settimeout(0.05)
print(f"{VERSION} pid={os.getpid()} listening ({how})", flush=True)
notify("READY=1")
while not stopping:
    try: c, _ = srv.accept()
    except socket.timeout: continue
    with lock: inflight += 1
    threading.Thread(target=handle, args=(c,), daemon=True).start()
srv.close()                                         # stop accepting, then drain
print(f"{VERSION} pid={os.getpid()} SIGTERM: draining {inflight} request(s)", flush=True)
notify("STOPPING=1")
while inflight:
    if os.environ.get("EXTEND"): notify("EXTEND_TIMEOUT_USEC=2000000")
    time.sleep(0.2)
print(f"{VERSION} pid={os.getpid()} drained, exiting", flush=True)

A load generator opens a fresh connection about 100 times per second from 4 threads for the length of the run and counts what each connection got:

#!/usr/bin/env python3
"""Open a new connection ~100x/s from 4 threads for DURATION seconds; count what happened."""
import collections, socket, sys, threading, time
DURATION = float(sys.argv[1]); PORT = 9100
res = collections.Counter(); first_err = []; slowest = [0.0]; t0 = time.monotonic(); lock = threading.Lock()
def one():
    c = socket.socket(); c.settimeout(8); t = time.monotonic()
    try:
        c.connect(("127.0.0.1", PORT)); c.sendall(b"GET\n")
        data = c.recv(100)
        out = data.decode().split()[0] if data else "EOF-without-reply"
    except Exception as e: out = type(e).__name__
    finally: c.close()
    with lock:
        slowest[0] = max(slowest[0], time.monotonic() - t)
        res[out] += 1
        if out not in ("v1", "v2", "v3"): first_err.append(round(time.monotonic() - t0, 2))
def loop():
    while time.monotonic() - t0 < DURATION: one(); time.sleep(0.03)
ts = [threading.Thread(target=loop) for _ in range(4)]; [t.start() for t in ts]; [t.join() for t in ts]
print("  results:", dict(res), "| errors between", (min(first_err), max(first_err)) if first_err else "-", "s | slowest request", round(slowest[0]*1000), "ms")

Step 1: stop semantics and draining

When systemd stops a service it sends SIGTERM to the main process, waits for it to exit, and if it has not exited within TimeoutStopSec, “it will be forcibly terminated by SIGKILL” (systemd.service(5)). A server that wants to finish in-flight work has to (a) stop accepting, (b) finish what it has, and (c) fit into that time, or ask for more.

The server above does (a) and (b): it closes its listener and waits for in-flight handlers. For (c), the manual says a service of Type=notify that sends EXTEND_TIMEOUT_USEC=... can extend the stop time past TimeoutStopSec; the first message must arrive before the timeout is exceeded, and it must be repeated within the stated interval, or the service must terminate on its own (systemd.service(5), sd_notify(3)). My server sends it every 0.2 seconds while draining, when EXTEND=1.

One 3-second request is in flight, and one second later the service is restarted:

#!/bin/bash
# usage: drain.sh <label> <TimeoutStopSec> <EXTEND 0|1>   : one 3-second request is in flight when the service is restarted
echo "=== $1  (TimeoutStopSec=$2, EXTEND_TIMEOUT_USEC sent: $( [ "$3" = 1 ] && echo yes || echo no ))"
systemctl stop web.socket web.service web-b.service 2>/dev/null
cat > /run/systemd/system/web.service <<EOT
[Unit]
Description=swapdemo drain
[Service]
Type=notify
Environment=VERSION=v1 WORK=3 $( [ "$3" = 1 ] && echo EXTEND=1 )
TimeoutStopSec=$2
ExecStart=/usr/bin/python3 /run/swapdemo/srv.py
EOT
systemctl daemon-reload; systemctl start web.service; sleep 0.5
python3 - <<'PY' &
import socket, time
t = time.monotonic(); c = socket.socket(); c.connect(("127.0.0.1", 9100)); c.sendall(b"GET\n")
try: d = c.recv(100); print(f"  client: {'reply ' + d.decode().strip() if d else 'connection closed with no reply'} after {time.monotonic()-t:.1f}s")
except Exception as e: print(f"  client: {type(e).__name__} after {time.monotonic()-t:.1f}s")
PY
P=$!
sleep 1; t=$(date +%s.%N); systemctl restart web.service; echo "  restart returned after $(python3 -c "import time;print(f'{time.time()-$t:.1f}')")s"
wait $P
journalctl -u web --no-pager -o cat --since "-8s" | grep -E 'SIGTERM|drained|killed|timed out|Failed|SIGKILL|signal' | sed 's/^/  journal: /'
=== default timeout  (TimeoutStopSec=90, EXTEND_TIMEOUT_USEC sent: no)
  client: reply v1 47537 after 3.0s
  restart returned after 2.1s
=== short timeout, no extension  (TimeoutStopSec=1, EXTEND_TIMEOUT_USEC sent: no)
  client: connection closed with no reply after 2.1s
  restart returned after 1.2s
=== short timeout, extended while draining  (TimeoutStopSec=1, EXTEND_TIMEOUT_USEC sent: yes)
  client: reply v1 47611 after 3.0s
  restart returned after 2.1s

(The script’s trailing journal: lines are omitted here: its time window also picked up lines from earlier runs.) A single run each. With the default, the reply took 3.0 s and the restart waited about 2.1 s for the drain. With a 1-second limit and no extension, the client’s connection was closed with no reply after 2.1 s: the process was killed before it finished. With the same 1-second limit and extension, the reply arrived after 3.0 s. This is the knob for services with long requests: a high fixed TimeoutStopSec also delays every normal stop and every stuck process, while the extension is paid only while real work is draining.

Step 2: the gap between old and new

Draining protects requests that already arrived. It does nothing for connections that arrive after the old process closed its listener and before the new one opened its own. I measured three ways to cover that, each with five runs of the hammer, with the replacement fired 2 seconds in:

#!/bin/bash
# plain service, bind itself
cat > /run/systemd/system/web.service <<EOT
[Unit]
Description=swapdemo app
[Service]
Type=notify
Environment=VERSION=v1 REUSEPORT=1
ExecStart=/usr/bin/python3 /run/swapdemo/srv.py
EOT
cat > /run/systemd/system/web-b.service <<EOT
[Unit]
Description=swapdemo app (new version, same port)
[Service]
Type=notify
Environment=VERSION=v2 REUSEPORT=1
ExecStart=/usr/bin/python3 /run/swapdemo/srv.py
EOT
cat > /run/systemd/system/web.socket <<EOT
[Socket]
ListenStream=127.0.0.1:9100
EOT
systemctl daemon-reload
#!/bin/bash
# usage: run.sh <label> '<start command>' '<what to do at t=2s>'
echo "=== $1"
systemctl stop web.socket web.service web-b.service 2>/dev/null; sleep 0.5
eval "$2"; sleep 1
python3 /run/swapdemo/hammer.py 5 &
H=$!
sleep 2
eval "$3"
wait $H
echo "  journal:"; journalctl -u web -u web-b --no-pager -o cat --since "-12s" 2>/dev/null | grep -v -E 'Started|Stopped|Starting|Stopping|Deactivated|Consumed|Succeeded|Finished' | sed 's/^/    /'
#!/bin/bash
# usage: bench.sh plain|overlap|socket [rounds]     (units.sh must have been run once; run.sh and hammer.py are in the same directory)
D=$(dirname "$0")
for i in $(seq 1 "${2:-5}"); do
  case $1 in
    plain)   $D/run.sh "plain restart, run $i"          'systemctl start web.service'  'systemctl restart web.service' ;;
    overlap) $D/run.sh "start v2 beside v1, run $i"     'systemctl start web.service'  'systemctl start web-b.service; sleep 0.3; systemctl stop web.service' ;;
    socket)  $D/run.sh "socket activation, run $i"      'systemctl start web.socket'   'systemctl restart web.service' ;;
  esac
done

Per run, the hammer prints counts by outcome (about 650 requests each time). Condensed from the five runs of each mode, in the order printed:

ModeRefused (per run)Reset (per run)OtherSlowest request
systemctl restart of a self-binding server5, 10, 34, 30, 41, 1, 1, 2, 31 connection closed with no reply (run 1)1 to 16 ms
Start v2 beside v1 (SO_REUSEPORT), 0.3 s later stop v10, 0, 0, 0, 01, 0, 1, 0, 0none1 to 10 ms
Socket activation, systemctl restart the service0, 0, 0, 0, 00, 0, 0, 0, 0none248 to 262 ms

In the plain-restart runs, the errors all fell within about a quarter of a second after the restart began (the hammer reports them between 2.0 and 2.24 s, with the restart at 2 s). Socket activation had no errors in any of its five runs.

A second set of five runs per mode, run later with the same scripts, gave: plain restart refused 33, 4, 37, 9 and 9 connections, with 2, 3, 2, 2 and 2 resets; overlap refused none and reset connections in 3 of 5 runs (0, 2, 1, 0, 1); socket activation had no errors, with a slowest request of 267, 248, 255, 285 and 58 ms. The ordering held. The counts, and which overlap runs had a reset, did not.

Notes on each:

  • Plain restart. In the window between the old process closing its socket and the new one binding, the port had no listener, so the kernel refused connections. This is the expected behaviour of a closed port; the numbers only show how many connections fall in a window of a few tens of milliseconds at this request rate.
  • Overlap with SO_REUSEPORT. Both versions listen on the same port at the same time, so there is never a moment without a listener. No connection was refused. But in some runs one connection was reset, at about the moment v1 closed its listener (2.3 s into the run: the new version starts at 2 s and the old one stops 0.3 s later). The next subsection finds the cause.
  • Socket activation. systemd owns the listening socket and passes it to the service as file descriptor 3, with LISTEN_FDS and LISTEN_PID set (sd_listen_fds(3)). With Accept=no (the default), “all listening sockets themselves are passed to the started service unit, and only one service unit is spawned for all connections” (systemd.socket(5)). When the old process exits, the listener stays open in systemd, so arriving connections queue instead of being refused. That is where the 248 to 262 ms slowest request comes from: those connections waited while the new process started.

Why the overlap lost a connection

The kernel documentation describes exactly this. A connection is tied to one listening socket when its SYN arrives. When that listener closes, connections still in the handshake and established ones waiting in its accept queue are aborted, even if another listener of the same SO_REUSEPORT group could have taken them. net.ipv4.tcp_migrate_req makes the kernel move them to another listener instead; the default is 0, and the setting arrived in Linux 5.14 (ip-sysctl). It has to be enabled before the reuseport group is created.

I did not want to leave that as a plausible story, so I tested it: the same server and hammer, without systemd, in a fresh network namespace (the setting is per namespace, so the host is untouched). The script replaces v1 by v2 eight times per setting; the timing is the one used above.

#!/bin/bash
# usage: migrate.sh <0|1>      Run as root inside a fresh network namespace, e.g.
#   ip netns add mig; ip netns exec mig bash migrate.sh 0
# Sets net.ipv4.tcp_migrate_req (per network namespace), then replaces v1 by v2 eight times under load,
# with both listening on one port via SO_REUSEPORT. srv.py and hammer.py must be in the same directory.
HERE=$(cd "$(dirname "$0")" && pwd)
ip link set lo up
sysctl -qw net.ipv4.tcp_migrate_req=$1
echo "tcp_migrate_req=$(cat /proc/sys/net/ipv4/tcp_migrate_req)"
for i in 1 2 3 4 5 6 7 8; do
  VERSION=v1 REUSEPORT=1 python3 $HERE/srv.py >/dev/null & P1=$!
  sleep 0.7
  python3 $HERE/hammer.py 4 & H=$!
  sleep 2
  VERSION=v2 REUSEPORT=1 python3 $HERE/srv.py >/dev/null & P2=$!
  sleep 0.3
  kill -TERM $P1                       # v1 closes its listener; v2 keeps listening
  wait $H
  kill -TERM $P2; wait $P1 $P2 2>/dev/null
done

Results, counting runs in which the hammer saw at least one reset:

tcp_migrate_reqRuns with a resetTotal resets
0 (the default)5 of 87
10 of 80

An earlier, longer pass of the same experiment, in the same kind of namespace, gave 11 of 16 runs with a reset at 0 and 0 of 16 at 1. So on this kernel (Linux 6.12) the resets in the overlap runs are the aborted accept-queue connections, and the sysctl removes them.

Two cautions. The same kernel page warns that migrating between listeners with different socket settings may crash applications, so do not flip this on a fleet without reading it. And it is a kernel-wide (per-namespace) setting, which a deployment script usually does not own; the other way out is the one in the last section of this post, socket activation, where no listener closes at all.

Step 3: swap the release, verify it, roll it back

Restarting into a new version is the risky half. The routine I use is boring on purpose:

  1. Releases live in separate directories; a symlink named current points to one.
  2. Switching is atomic: create the new symlink under a temporary name, then rename it over the old one (mv -T).
  3. Restart, then ask the service which release it is, with a deadline. “It answers” is not enough.
  4. If the check fails, switch the symlink back and restart again.
#!/bin/bash
# usage: swap.sh <release-name>      Releases live in /run/swapdemo/rel/<name>/srv.py ; /run/swapdemo/current is a symlink.
set -u
new=$1; base=/run/swapdemo; unit=web.service
old=$(basename "$(readlink $base/current)")
echo "swap: $old -> $new"
healthy() {   # healthy = the service answers one request AND says it is the release we expect
  for i in $(seq 1 20); do
    out=$(python3 - <<PY 2>/dev/null
import socket
c = socket.socket(); c.settimeout(1); c.connect(("127.0.0.1", 9100)); c.sendall(b"GET\n"); print(c.recv(100).decode().split()[0])
PY
)
    [ "$out" = "$1" ] && return 0
    sleep 0.25
  done
  return 1
}
ln -sfn "$base/rel/$new" "$base/current.tmp" && mv -T "$base/current.tmp" "$base/current"   # atomic switch
systemctl restart $unit 2>/dev/null
if healthy "$new"; then echo "swap: $new is healthy, keeping it"; exit 0; fi
echo "swap: $new did not become healthy within 5 s, rolling back to $old"
ln -sfn "$base/rel/$old" "$base/current.tmp" && mv -T "$base/current.tmp" "$base/current"
systemctl reset-failed $unit 2>/dev/null; systemctl restart $unit
healthy "$old" && echo "swap: $old is serving again" || echo "swap: ROLLBACK FAILED"
exit 1
#!/bin/bash
b=/run/swapdemo; mkdir -p $b/rel/{v1,v2,v3,v4}
for v in v1 v2 v3 v4; do cp $b/srv.py $b/rel/$v/srv.py; sed -i "s/^VERSION = .*/VERSION = \"$v\"/" $b/rel/$v/srv.py; done
# v3 crashes at start; v4 starts, but answers with the wrong label
sed -i 's/^srv, how = listener().*/import sys; sys.exit("v3 cannot start: bad config")/' $b/rel/v3/srv.py
sed -i 's/^VERSION = .*/VERSION = "v1-but-actually-broken"/' $b/rel/v4/srv.py
ln -sfn $b/rel/v1 $b/current
cat > /run/systemd/system/web.service <<EOT
[Unit]
Description=swapdemo release swap
[Service]
Type=notify
Environment=REUSEPORT=1
ExecStart=/usr/bin/python3 /run/swapdemo/current/srv.py
TimeoutStartSec=3
EOT
systemctl daemon-reload; systemctl stop web.socket web-b.service 2>/dev/null; systemctl restart web.service

Three swaps from the same starting point (v1 serving). v3 crashes at start; v4 starts but answers with the wrong label:

swap: v1 -> v2
swap: v2 is healthy, keeping it
swap: v2 -> v3
swap: v3 did not become healthy within 5 s, rolling back to v2
swap: v2 is serving again
swap: v2 -> v4
swap: v4 did not become healthy within 5 s, rolling back to v2
swap: v2 is serving again

The v4 case is why the check compares the version label and does not stop at “got a reply”. A health check that only tests that the port answers would have kept v4. In this lab the “label” stands for whatever your service can truthfully report about itself (a build id, a schema version).

Step 4: state handoff, as a rule only

I did not exercise this. The rule I would follow: the old version writes its state to a file or database with an atomic replace (write to a temporary name, flush, rename over the target); the new version reads it at start and treats a missing or unreadable file as “start fresh” rather than failing. Anything held only in memory is gone when the process exits, and no restart strategy above changes that.

Which one to use

SituationChoice
Short requests, brief refused connections are fine (a retrying client, internal tool)Plain restart, with a drain and a sane TimeoutStopSec
New connections must never be refused, and a few hundred ms extra latency on one restart is fineSocket activation
You cannot use systemd’s socket units, or you need both versions serving for a whileOverlap with SO_REUSEPORT, accepting that some connections in the closing listener’s queue can be reset (2 of 5 and 3 of 5 runs here) unless net.ipv4.tcp_migrate_req=1 is on (Linux 5.14 and later; read the kernel’s warning first)
Long requestsDrain, and EXTEND_TIMEOUT_USEC instead of a huge TimeoutStopSec
A bad release is costlyThe swap with a version-checked health check and rollback, regardless of the above

Ways a replacement goes wrong

  • Not draining. Symptom: clients see a closed connection with no reply at every deploy. Reproduced (the 1-second timeout run). Fix: stop accepting, finish in-flight work, fit into the timeout or extend it.
  • Extending once. Symptom: killed anyway after the extension elapses. The documentation says the message must be repeated within the interval; my server repeats it every 0.2 s. I did not test stopping the messages.
  • A health check that cannot tell versions apart. Symptom: a broken release stays. Reproduced (v4). Fix: make the check assert the expected identity.
  • Rolling back to something that no longer exists. Symptom: the rollback fails too. The swap keeps old releases in place and flips a symlink; clean old releases only after the new one has proven itself.
  • Counting on SO_REUSEPORT overlap to be lossless. Symptom: rare resets at the moment of handover. Reproduced, and removed with tcp_migrate_req=1 in the namespace test. Fix: enable the migration (after reading the kernel note), or use socket activation.
  • Persistent client connections. A server that drains must also close idle keep-alive connections, or the drain waits until a client sends something. My lab server closes every connection after one reply, so I did not exercise this.
  • Deploying without a test at your real request rate. The counts above come from about 100 connections per second on loopback. Your window will hold more or fewer.

Run it yourself

Needs a machine with systemd as PID 1, root, and Python 3. A disposable VM is best.

  1. sudo mkdir -p /run/swapdemo && sudo cp srv.py hammer.py units.sh run.sh bench.sh drain.sh swap.sh setup_rel.sh /run/swapdemo/ && sudo chmod +x /run/swapdemo/*
  2. sudo /run/swapdemo/units.sh && sudo /run/swapdemo/bench.sh plain 5. Success: lines with ConnectionRefusedError counts in every run.
  3. sudo /run/swapdemo/bench.sh overlap 5 and sudo /run/swapdemo/bench.sh socket 5. Success: no refusals; the socket mode shows a slowest request of a couple of hundred ms and no errors.
  4. sudo /run/swapdemo/drain.sh "a" 90 0, then ... "b" 1 0, then ... "c" 1 1. Success: reply, no reply, reply.
  5. sudo /run/swapdemo/setup_rel.sh, then sudo /run/swapdemo/swap.sh v2, v3, v4.
  6. Optional: ip netns add mig; sudo ip netns exec mig bash migrate.sh 0, then ... migrate.sh 1, with srv.py and hammer.py next to it, to see the reset count change; ip netns del mig afterwards.
  7. Clean up: stop web.service, web-b.service, web.socket, and remove the unit files in /run/systemd/system and /run/swapdemo; then systemctl daemon-reload.

What this box proved, and what it cannot

Verified: all outputs above, on a Linux 6.12 box with systemd 257 as PID 1 of a throwaway PID, mount and cgroup namespace, as root, Python 3.13. Five runs for each of the three modes (a clean sequential set; an earlier set that overlapped with another run was discarded) and a second set of five on a later day, one run for each drain case (reproduced once more with the same results), one for each swap (also reproduced), and the tcp_migrate_req comparison in a separate network namespace without systemd. The statements from the systemd manual pages, the kernel ip-sysctl text and the socket-activation pages were read at the original pages while writing.

Not verified: whether tcp_migrate_req also helps with in-flight handshakes under heavier load, and the BPF-based migration policy the kernel page mentions; persistent connections, TLS, or real clients; Accept=yes socket activation; behaviour when the extension messages stop; state handoff; behaviour on other systemd versions; production load. Counts depend on the hammer’s rate and on how fast this box starts a process, so treat them as an illustration of the ordering (plain > overlap > socket activation in errors), not as numbers to expect.

Where connections wait

Don’t ask “how fast is my restart?” Ask where connections wait while it happens. Give them somewhere to wait that outlives the process, drain what is already in flight, and make the new version prove its identity before you stop worrying.