Keeping an Agent Resident Behind NAT
Designing a resident agent behind NAT with outbound connections, separate liveness checks, and on-demand SSH sessions for interactive access.
Designing a resident agent behind NAT with outbound connections, separate liveness checks, and on-demand SSH sessions for interactive access.
Five techniques from a polyglot codebase: rolling out golangci-lint on existing Go code, running the race detector in CI, bounding subprocess lifetimes, re-registering launchd services without the bootout race, and choosing PostgreSQL row-lock strength around foreign keys.
Why I built an AI agent platform in Go instead of using LangChain, and how Clean Architecture made model, memory, tool, and streaming concerns independently replaceable.
Welcome to my blog – a space for sharing notes on CS, ML, and engineering.
A machine with Wi-Fi, a VPN and a container bridge has several addresses, and a hostname lookup can return the wrong one. Connecting a UDP socket to a destination makes the kernel pick the source address it would use, and getsockname() reads it back without a single packet being sent. Checked on Linux with three interfaces: five destinations, five answers, no IPv4 or ARP frames on the wire.
A runbook for replacing a running daemon without dropping requests, tested on a small TCP server under real systemd. A plain restart refused or reset connections every time; starting the new version beside the old one avoided refusals but still reset a connection in some runs, and a kernel setting made those resets go away; socket activation produced no errors. Also covered: draining in-flight work against TimeoutStopSec, and a release swap that checks the new version and rolls back.
An over-the-air update is JavaScript that calls into a native binary you cannot change. Treat the runtime version as the ABI version of that binary. Experiments with @expo/fingerprint 0.20.13 show which edits change the hash (extra, version, build numbers, native modules) and which do not (JavaScript, pure-JS dependencies), and a skip list that silently drops a default.
A daemon that updates itself by starting a helper and restarting its own unit has a trap: the helper starts inside the service’s cgroup, and systemd kills everything in that cgroup on restart. In a throwaway systemd, a setsid’d helper vanished before its last log line, KillMode=process kept it alive but leaked it into the next instance, and systemd-run gave it its own unit. This post shows the evidence, the fix, and what I did not test.
A client library shared between a website and a React Native app usually fails on the boring parts of the Web platform, not on fetch. Reading React Native 0.87.1 and Expo 57 sources shows three layers of runtime, a URL class that returns ‘/b/c/d/../x’ where the web returns ‘/b/x’, and a fetch with no response stream unless Expo replaces it. A small client that checks what it needs, tested against four simulated runtimes.
A Durable Object has one alarm, runs it at least once, and retries a failing handler with backoff up to six times. After that nothing wakes the object again. A walkthrough of a lease-expiry ledger that survives all three: a table as the source of truth, a constructor that re-arms the alarm, and an effect that tolerates being run twice. Run on local workerd, with the documented limits quoted from Cloudflare.
Cloudflare’s docs say hibernation keeps WebSocket connections open and discards in-memory state. This post runs a small Durable Object under local workerd and watches exactly what that means: the socket stays OPEN, the constructor runs again on the next message, class fields reset, attachments survive only if you serialize them, and a pending timer or the standard WebSocket API quietly turns hibernation off. Production behaviour and billing are not tested.
Apple and Google both document what happens to a push when the device is offline, the app is killed or the sender is too chatty: messages are replaced, dropped, reordered or delayed. A table of those documented failure modes, and a small simulation showing that a client which pulls from a cursor converges where a client which applies push payloads does not.
The JWT specification tells verifiers to ignore claims they do not understand. That is harmless for claims that grant things and dangerous for claims that restrict them. Four ways to cap what a client kind may hold, tested with a small issuer and three verifiers, and the one I would pick.
RFC 9470 lets an API tell a client that its access token was obtained with too weak or too old a login. Reading the RFC section by section, then building a toy server and client, shows what the document specifies (a 401 challenge, two parameters) and the three things it leaves to you: concurrent challenges, loops, and caches.
Apple’s and Android’s own documentation say a backgrounded app can be suspended, its network access deferred, and its existing connections closed. So when your app comes back, readyState === OPEN is only a memory of the last event, not a measurement. This post derives a small foreground routine (probe, rebuild, catch up) from those documented rules, runs it against a frozen-process stand-in on Linux, and is explicit about what no device was used to check.
The Origin header on a WebSocket handshake is set by browsers and can be set to anything by every other client. So an allow-list protects a cookie-authenticated browser session from other sites, and it proves nothing about who is connecting. A runnable server, a real headless-browser attack, and the handshake rule that follows.
A request deadline fired on a connection that carries many requests at once. Do you close the connection, or only the request? In a small Node lab, closing the HTTP/2 session failed two innocent requests, while cancelling only the stream let them finish on the same TCP connection. The same lab shows the one case where cancelling is not enough: a lost packet stalls every stream on a TCP connection, and a connection-level PING is the right way to tell.
WebSocket send() never waits, and bufferedAmount only sees the part of the path that is in your own process. In a small lab, a sender that politely waited for bufferedAmount to drain still let a slow consumer queue nearly all 400 messages in its own memory, while a credit window of 8 kept the backlog at 8. This post goes through four common beliefs about send(), bufferedAmount, message size and backpressure, tests each one, and ends with a table for choosing between them.
How a client catches up after a reconnect without losing state, without duplicates, and without re-reading history. Built in steps (versions 0 to 4), from “read everything again” to a bounded log with a snapshot fallback, with what a real browser’s EventSource does on reconnect and on a non-200 response, a snapshot-ordering bug, and a decision tree.
Send a header and a body as two small write() calls, then wait for a reply, and every round trip can stall for about 40 ms on an idle machine with an idle network. Neither Nagle’s algorithm nor delayed ACK is a bug; together they deadlock until a timer fires. This post predicts the stall, reproduces it with a 50-line script, reads it off a packet trace, and compares the four ways out.
Refresh token rotation turns a reused token into a theft alarm. A client that fires three requests with an expired access token and refreshes once per 401 presents the same refresh token three times, and trips that alarm on itself. This post breaks a naive client on purpose, fixes it with single-flight refresh and a stale check, and compares the alternatives: a server grace window, request gating, proactive refresh and sender-constrained tokens.
When a request times out, the client cannot tell whether the request or only its acknowledgement was lost. This is why exactly-once delivery cannot be built, and why effectively-once processing is at-least-once delivery plus a receiver that deduplicates. A runnable experiment with a flaky network, the bug that still duplicates 89 of 200 requests, an atomic dedupe store, and a jitter simulation.
epoll’s edge-triggered mode stalls after a partial read; level-triggered mode cannot. The same distinction decides whether a realtime client survives lost, duplicated and reordered events. A reproducible epoll experiment, a five-way simulation on a lossy channel, and a 40-line invalidator that handles the race everyone misses: an event that arrives while the refetch is running.
TCP never tells you that the peer is gone. A reader blocks forever, a writer keeps retransmitting for a quarter of an hour, and SO_KEEPALIVE does nothing while data is unacknowledged. This post builds a one-file Linux lab with no root, climbs a ladder of mechanisms (nothing, keepalive, TCP_USER_TIMEOUT, an application heartbeat), measures how long each takes to notice, and shows where NAT and the RFCs fit in.
After the 101 response, no request carries credentials: the server remembers a verdict, not the evidence. This article follows one token through a browser handshake, shows what a page can and cannot see, and compares three ways to renew it (reconnect, in-band, out-of-band) with a runnable server, tests, and a headless-Chrome check.
How a mobile client moved from timer-driven polling to shared change subscriptions over a per-message-billed relay. The subtle part is not the stream. It is deciding honestly when the fallback poll may stand down.
When every WebSocket message costs money and every frame has a hard cap, you need batching with independent flush triggers, a single-frame threshold below the cap, bounded chunking, and capability flags that survive mixed-version rollouts. The sharp edges are ordering on shutdown and a zero value that disappears from the wire.
A fixed-interval keepalive costs a message in each direction even when the connection is busy. Treating every inbound frame as proof of life, probing only quiet connections, and letting an existing heartbeat tell an idle host when to disconnect removes most of that traffic, at the price of a bounded wake-up delay.
A client that waits for its own socket’s close event before cleaning up can stall silently when that event never arrives. Settle your state when you decide to end the connection, make the cleanup idempotent, and keep the event for closes you did not start.
Designing collaborative document editing with conflict-free synchronization, saved snapshots, and safeguards against overwriting content before initial sync.
Controlling distributed task starts with one unfinished admission per workspace surface, a fixed destination, and idempotent retries during machine outages.
Designing API authentication and role-based authorization with actors, route permissions, and separate credentials for people and machines.
Designing AI agent memory with retrieved context, durable task records, and evidence-based create or update decisions that tolerate search index lag.
Designing a meeting AI pipeline with separate transcription, summarization, and task extraction, preserving evidence and resolving deadlines from meeting time.
Designing a resident agent behind NAT with outbound connections, separate liveness checks, and on-demand SSH sessions for interactive access.
Five techniques from a polyglot codebase: rolling out golangci-lint on existing Go code, running the race detector in CI, bounding subprocess lifetimes, re-registering launchd services without the bootout race, and choosing PostgreSQL row-lock strength around foreign keys.
Why I built an AI agent platform in Go instead of using LangChain, and how Clean Architecture made model, memory, tool, and streaming concerns independently replaceable.
Welcome to my blog – a space for sharing notes on CS, ML, and engineering.