Socket Activation Keeps Connections Waiting, Not Refused, During a systemd Restart

A runbook for replacing a running daemon without dropping requests, tested on a small TCP server under real systemd. A plain restart refused or reset connections every time; starting the new version beside the old one avoided refusals but still reset a connection in some runs, and a kernel setting made those resets go away; socket activation produced no errors. Also covered: draining in-flight work against TimeoutStopSec, and a release swap that checks the new version and rolls back.

October 3, 2026 | 18 min

systemd Kills the Updater Your Service Started, Even With setsid

A daemon that updates itself by starting a helper and restarting its own unit has a trap: the helper starts inside the service’s cgroup, and systemd kills everything in that cgroup on restart. In a throwaway systemd, a setsid’d helper vanished before its last log line, KillMode=process kept it alive but leaked it into the next instance, and systemd-run gave it its own unit. This post shows the evidence, the fix, and what I did not test.

September 28, 2026 | 12 min

A Durable Object Alarm Retries 6 Times, Then Stops

A Durable Object has one alarm, runs it at least once, and retries a failing handler with backoff up to six times. After that nothing wakes the object again. A walkthrough of a lease-expiry ledger that survives all three: a table as the source of truth, a constructor that re-arms the alarm, and an effect that tolerates being run twice. Run on local workerd, with the documented limits quoted from Cloudflare.

September 24, 2026 | 10 min

APNs Stores One Pending Notification per App, So Treat Push as a Hint

Apple and Google both document what happens to a push when the device is offline, the app is killed or the sender is too chatty: messages are replaced, dropped, reordered or delayed. A table of those documented failure modes, and a small simulation showing that a client which pulls from a cursor converges where a client which applies push payloads does not.

September 19, 2026 | 9 min

After Backgrounding, a WebSocket Reporting OPEN Is Only a Claim

Apple’s and Android’s own documentation say a backgrounded app can be suspended, its network access deferred, and its existing connections closed. So when your app comes back, readyState === OPEN is only a memory of the last event, not a measurement. This post derives a small foreground routine (probe, rebuild, catch up) from those documented rules, runs it against a frozen-process stand-in on Linux, and is explicit about what no device was used to check.

September 12, 2026 | 15 min

Cancel the Stream, Not the Connection, When One HTTP/2 Request Times Out

A request deadline fired on a connection that carries many requests at once. Do you close the connection, or only the request? In a small Node lab, closing the HTTP/2 session failed two innocent requests, while cancelling only the stream let them finish on the same TCP connection. The same lab shows the one case where cancelling is not enough: a lost packet stalls every stream on a TCP connection, and a connection-level PING is the right way to tell.

September 8, 2026 | 19 min

Resuming an Event Stream with a Cursor, a Bounded Log and a Snapshot

How a client catches up after a reconnect without losing state, without duplicates, and without re-reading history. Built in steps (versions 0 to 4), from “read everything again” to a bounded log with a snapshot fallback, with what a real browser’s EventSource does on reconnect and on a non-200 response, a snapshot-ordering bug, and a decision tree.

September 4, 2026 | 17 min

Retries Duplicate Your Writes, and Exactly-Once Won't Save You

When a request times out, the client cannot tell whether the request or only its acknowledgement was lost. This is why exactly-once delivery cannot be built, and why effectively-once processing is at-least-once delivery plus a receiver that deduplicates. A runnable experiment with a flaky network, the bug that still duplicates 89 of 200 requests, an atomic dedupe store, and a jitter simulation.

August 28, 2026 | 16 min

A Dead TCP Peer Goes Unnoticed for 15 Minutes on Writes, Forever on Reads

TCP never tells you that the peer is gone. A reader blocks forever, a writer keeps retransmitting for a quarter of an hour, and SO_KEEPALIVE does nothing while data is unacknowledged. This post builds a one-file Linux lab with no root, climbs a ladder of mechanisms (nothing, keepalive, TCP_USER_TIMEOUT, an application heartbeat), measures how long each takes to notice, and shows where NAT and the RFCs fit in.

August 23, 2026 | 19 min

Liveness Without Pings, and Idle Sleep

A fixed-interval keepalive costs a message in each direction even when the connection is busy. Treating every inbound frame as proof of life, probing only quiet connections, and letting an existing heartbeat tell an idle host when to disconnect removes most of that traffic, at the price of a bounded wake-up delay.

August 18, 2026 | 9 min