<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>reliability on Ikoma's Blog</title><link>https://blog.yusukeikoma.com/tags/reliability/</link><description>Recent content in reliability on Ikoma's Blog</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 03 Oct 2026 18:08:27 +0900</lastBuildDate><atom:link href="https://blog.yusukeikoma.com/tags/reliability/index.xml" rel="self" type="application/rss+xml"/><item><title>Socket Activation Keeps Connections Waiting, Not Refused, During a systemd Restart</title><link>https://blog.yusukeikoma.com/posts/replace-running-daemon-without-downtime/</link><pubDate>Sat, 03 Oct 2026 18:08:27 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/replace-running-daemon-without-downtime/</guid><description>A runbook for replacing a running daemon without dropping requests, tested on a small TCP server under real systemd. A plain restart refused or reset connections every time; starting the new version beside the old one avoided refusals but still reset a connection in some runs, and a kernel setting made those resets go away; socket activation produced no errors. Also covered: draining in-flight work against TimeoutStopSec, and a release swap that checks the new version and rolls back.</description></item><item><title>systemd Kills the Updater Your Service Started, Even With setsid</title><link>https://blog.yusukeikoma.com/posts/systemd-killmode-self-update/</link><pubDate>Mon, 28 Sep 2026 09:48:47 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/systemd-killmode-self-update/</guid><description>A daemon that updates itself by starting a helper and restarting its own unit has a trap: the helper starts inside the service&amp;rsquo;s cgroup, and systemd kills everything in that cgroup on restart. In a throwaway systemd, a setsid&amp;rsquo;d helper vanished before its last log line, KillMode=process kept it alive but leaked it into the next instance, and systemd-run gave it its own unit. This post shows the evidence, the fix, and what I did not test.</description></item><item><title>A Durable Object Alarm Retries 6 Times, Then Stops</title><link>https://blog.yusukeikoma.com/posts/durable-object-alarm-retries-six-times-then-stops/</link><pubDate>Thu, 24 Sep 2026 21:33:45 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/durable-object-alarm-retries-six-times-then-stops/</guid><description>A Durable Object has one alarm, runs it at least once, and retries a failing handler with backoff up to six times. After that nothing wakes the object again. A walkthrough of a lease-expiry ledger that survives all three: a table as the source of truth, a constructor that re-arms the alarm, and an effect that tolerates being run twice. Run on local workerd, with the documented limits quoted from Cloudflare.</description></item><item><title>APNs Stores One Pending Notification per App, So Treat Push as a Hint</title><link>https://blog.yusukeikoma.com/posts/push-notifications-are-hints-apns-stores-one/</link><pubDate>Sat, 19 Sep 2026 21:10:16 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/push-notifications-are-hints-apns-stores-one/</guid><description>Apple and Google both document what happens to a push when the device is offline, the app is killed or the sender is too chatty: messages are replaced, dropped, reordered or delayed. A table of those documented failure modes, and a small simulation showing that a client which pulls from a cursor converges where a client which applies push payloads does not.</description></item><item><title>After Backgrounding, a WebSocket Reporting OPEN Is Only a Claim</title><link>https://blog.yusukeikoma.com/posts/mobile-os-background-sockets/</link><pubDate>Sat, 12 Sep 2026 14:18:24 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/mobile-os-background-sockets/</guid><description>Apple&amp;rsquo;s and Android&amp;rsquo;s own documentation say a backgrounded app can be suspended, its network access deferred, and its existing connections closed. So when your app comes back, readyState === OPEN is only a memory of the last event, not a measurement. This post derives a small foreground routine (probe, rebuild, catch up) from those documented rules, runs it against a frozen-process stand-in on Linux, and is explicit about what no device was used to check.</description></item><item><title>Cancel the Stream, Not the Connection, When One HTTP/2 Request Times Out</title><link>https://blog.yusukeikoma.com/posts/http2-timeout-stream-vs-connection/</link><pubDate>Tue, 08 Sep 2026 12:56:00 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/http2-timeout-stream-vs-connection/</guid><description>A request deadline fired on a connection that carries many requests at once. Do you close the connection, or only the request? In a small Node lab, closing the HTTP/2 session failed two innocent requests, while cancelling only the stream let them finish on the same TCP connection. The same lab shows the one case where cancelling is not enough: a lost packet stalls every stream on a TCP connection, and a connection-level PING is the right way to tell.</description></item><item><title>Resuming an Event Stream with a Cursor, a Bounded Log and a Snapshot</title><link>https://blog.yusukeikoma.com/posts/resumable-streams-cursors-last-event-id-snapshots/</link><pubDate>Fri, 04 Sep 2026 19:26:37 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/resumable-streams-cursors-last-event-id-snapshots/</guid><description>How a client catches up after a reconnect without losing state, without duplicates, and without re-reading history. Built in steps (versions 0 to 4), from &amp;ldquo;read everything again&amp;rdquo; to a bounded log with a snapshot fallback, with what a real browser&amp;rsquo;s EventSource does on reconnect and on a non-200 response, a snapshot-ordering bug, and a decision tree.</description></item><item><title>Retries Duplicate Your Writes, and Exactly-Once Won't Save You</title><link>https://blog.yusukeikoma.com/posts/exactly-once-delivery-idempotency-and-backoff/</link><pubDate>Fri, 28 Aug 2026 10:07:45 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/exactly-once-delivery-idempotency-and-backoff/</guid><description>When a request times out, the client cannot tell whether the request or only its acknowledgement was lost. This is why exactly-once delivery cannot be built, and why effectively-once processing is at-least-once delivery plus a receiver that deduplicates. A runnable experiment with a flaky network, the bug that still duplicates 89 of 200 requests, an atomic dedupe store, and a jitter simulation.</description></item><item><title>A Dead TCP Peer Goes Unnoticed for 15 Minutes on Writes, Forever on Reads</title><link>https://blog.yusukeikoma.com/posts/tcp-half-open-connection-detection/</link><pubDate>Sun, 23 Aug 2026 15:09:06 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/tcp-half-open-connection-detection/</guid><description>TCP never tells you that the peer is gone. A reader blocks forever, a writer keeps retransmitting for a quarter of an hour, and SO_KEEPALIVE does nothing while data is unacknowledged. This post builds a one-file Linux lab with no root, climbs a ladder of mechanisms (nothing, keepalive, TCP_USER_TIMEOUT, an application heartbeat), measures how long each takes to notice, and shows where NAT and the RFCs fit in.</description></item><item><title>Liveness Without Pings, and Idle Sleep</title><link>https://blog.yusukeikoma.com/posts/liveness-without-pings-and-idle-sleep/</link><pubDate>Tue, 18 Aug 2026 22:41:09 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/liveness-without-pings-and-idle-sleep/</guid><description>A fixed-interval keepalive costs a message in each direction even when the connection is busy. Treating every inbound frame as proof of life, probing only quiet connections, and letting an existing heartbeat tell an idle host when to disconnect removes most of that traffic, at the price of a bounded wake-up delay.</description></item><item><title>When the Close Event Never Comes</title><link>https://blog.yusukeikoma.com/posts/when-the-close-event-never-comes/</link><pubDate>Mon, 17 Aug 2026 17:27:04 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/when-the-close-event-never-comes/</guid><description>A client that waits for its own socket&amp;rsquo;s close event before cleaning up can stall silently when that event never arrives. Settle your state when you decide to end the connection, make the cleanup idempotent, and keep the event for closes you did not start.</description></item><item><title>Hardening a Go Daemon and a Python API</title><link>https://blog.yusukeikoma.com/posts/go-daemon-and-python-api-hardening/</link><pubDate>Sat, 01 Aug 2026 14:11:31 +0900</pubDate><guid>https://blog.yusukeikoma.com/posts/go-daemon-and-python-api-hardening/</guid><description>Five techniques from a polyglot codebase: rolling out golangci-lint on existing Go code, running the race detector in CI, bounding subprocess lifetimes, re-registering launchd services without the bootout race, and choosing PostgreSQL row-lock strength around foreign keys.</description></item></channel></rss>