TL;DR

  • TCP has no “are you still there?” signal. A connection is just state held at each end. If the peer vanishes silently (a NAT drops its mapping, a radio link dies, a machine loses power), the local kernel finds out only when something it sends goes unanswered for long enough.
  • So what you see depends on what your side is doing. In my lab (Linux 6.12, kernel defaults): a socket that only reads was still blocked, with no error, after 60 s of observation. A socket that writes returned from write() happily, then failed with ETIMEDOUT after 938 s (about 15.6 minutes), the retransmission give-up time (the man page says 13–30 minutes for the default of 15 retries).
  • SO_KEEPALIVE detects a dead peer on an idle connection (4.6 s with a tuned 2 s / 1 s / 3 probes), but does nothing while sent data is unacknowledged, which is when you most want it. TCP_USER_TIMEOUT bounds that case (5.5 s with a 5 s setting). An application heartbeat is the only mechanism that also checks the other application, and it noticed after 3.0 s of silence with a 1 s ping interval and a 3 s deadline.
  • The classic “half-open” case, where the peer lost its state but is reachable again, resolves itself with a RST on your next write. The harder case is the silent black hole, which never resolves by itself.
  • Everything here is reproducible with a single 107-line Python file. No root: it uses unprivileged user and network namespaces.

Background: what “the connection is dead” even means

Long-lived TCP connections (a message stream, a database session, a WebSocket) tend to sit behind NAT and, for mobile clients, behind a radio link that comes and goes. When something in the path dies without saying goodbye, the application keeps believing the connection is fine.

RFC 9293 §3.5.1 defines a half-open connection: one end has closed or aborted without the other’s knowledge, or the ends have become desynchronized by a failure or reboot that lost memory. Such a connection “will automatically become reset if an attempt is made to send data in either direction”: the end that still has state sends data, the end that lost it answers with a RST.

That is the textbook case, and it is the benign one. The case that hurts is the silent black hole: packets towards the peer are dropped and nothing comes back, not a FIN, not a RST, not an ICMP error. A NAT that times out a mapping may do either. RFC 5382 leaves it to the implementation: when a NAT abandons a live connection it “MAY either send TCP RST packets to the endpoints or MAY silently abandon the connection”, and for later packets without a mapping, “the decision to either silently drop such packets or to respond with a TCP RST packet is left up to the implementation.” So you cannot assume you will be told.

The central question of this post: how long does your program take to find out, and what can you do to make that time short without false alarms?

Why the kernel cannot know

TCP’s job is reliable delivery. From the sender’s side, a lost packet and a dead peer look the same: no ACK. The only defensible policy is to retransmit with exponential backoff and give up eventually. RFC 9293 §3.8.3 describes two thresholds, R1 (start suspecting the path) and R2 (close the connection), says R2 “SHOULD correspond to at least 100 seconds”, and requires that “an application MUST be able to set the value for R2 for a particular connection”. The retransmission timer itself backs off up to a ceiling; RFC 6298 §2 allows a maximum RTO “provided it is at least 60 seconds”, and Linux defines its cap as 120 s (TCP_RTO_MAX_SEC).

Linux’s R2 is the tcp_retries2 sysctl. tcp(7): default 15, “approximately between 13 to 30 minutes, depending on the retransmission timeout”, and the page adds that the RFC 1122 minimum of 100 seconds “is typically deemed too short”.

If you are only waiting for data, nothing is retransmitted, so there is no timer at all. That is the nastiest case, and the first rung of the ladder.

The ladder

Each rung adds one mechanism. For every rung there is one scenario in the lab script (the full script is in the lab section below), so each time below is a measurement, not a calculation.

Rung 0: do nothing, only read

A client sent a request, and is now blocked waiting for the reply. The peer’s end is then cut off. In the lab:

[    0.00s] peer lost (all packets dropped)
[   60.06s] still blocked after 60 s: no data, no EOF, no error

ss -tno shows the socket as ESTAB with no timer at all. The kernel has nothing to retransmit, nothing to probe, and no reason to wake up. I stopped observing at 60 s; there is no reason it would ever end. This is why a reading client needs its own timeout, whatever else is configured.

Rung 1: do nothing, but write

Now the client writes 100 bytes after the peer is lost. write() succeeds immediately, because it only copies into the send buffer. Then the retransmission timer takes over. ss -tno makes the backoff visible as timer:(on,<time left>,<retransmit count>):

[    0.00s] peer lost (all packets dropped)
[    0.00s] write() returned: the bytes only reached the send buffer
[   30.01s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,22sec,7)
[  120.02s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min31sec,9)
[  240.04s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min33sec,10)
[  360.05s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min34sec,11)
[  480.07s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min35sec,12)
[  600.08s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min35sec,13)
[  720.10s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min36sec,14)
[  840.11s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,1min37sec,15)
[  930.13s] ss: ESTAB 0 100 10.0.0.1:52164 10.0.0.2:9000 timer:(on,7.464sec,15)
[  938.42s] recv() failed: errno=110 Connection timed out

(Excerpt: one ss sample about every 30 s from the full log, whitespace condensed.) On this machine the connection failed with ETIMEDOUT after 938 s (about 15.6 minutes). That matches the 15-retry default in the man page, but the exact figure depends on the RTO, which depends on the path’s round-trip time; on a real network it will differ. The kernel documentation offers a cross-check: with the default tcp_retries2 of 15, a connection that starts at the 200 ms minimum RTO and doubles up to the 120 s ceiling has a hypothetical timeout of 924.6 s (ten doubling waits that add up to 204.6 s, plus six waits capped at 120 s). That figure is a lower bound, because TCP gives up at the first RTO that exceeds it (ip-sysctl). 938 s sits just above it, as it should.

Rung 2: SO_KEEPALIVE on an idle connection

Keepalive was designed for the idle case. RFC 9293 §3.8.4 makes it optional (MAY-5), requires it to default to off, and requires the interval to default to “no less than two hours”. Linux’s defaults are 7200 s idle, 75 s between probes, 9 probes (tcp(7)): an idle dead connection dies “approximately an additional 11 minutes” after the two hours. Those defaults are for reaping, not for fast failure. Per socket you can tune them (TCP_KEEPIDLE, TCP_KEEPINTVL, TCP_KEEPCNT). With 2 s, 1 s, 3 probes in the lab:

[    4.59s] recv() failed: errno=110 Connection timed out

Probes start after 2 s of idleness (counted from the last activity, which was half a second before the peer was lost), three unanswered probes later the kernel gives up. So the arithmetic is roughly idle + probes × interval, minus the idle time already elapsed. Note why three probes: RFC 9293 requires that a keepalive implementation “MUST NOT interpret failure to respond to any specific probe as a dead connection”, because pure ACKs are not retransmitted reliably.

Rung 3: the same keepalive, but with unacknowledged data

Set exactly the same options, but write 100 bytes after the peer is lost. If keepalive were a general liveness check, this would take about 4.6 s again. It does not:

[    0.00s] peer lost (all packets dropped)
[    0.00s] write() returned: the bytes only reached the send buffer
[    1.01s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,660ms,2)
[   30.11s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,22sec,7)
[  120.47s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,1min30sec,9)
[  300.15s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,33sec,10)
[  600.27s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,1min35sec,13)
[  930.48s] ss: ESTAB 0 100 10.0.0.1:52196 10.0.0.2:9000 timer:(on,7.108sec,15)
[  938.42s] recv() failed: errno=110 Connection timed out

The keepalive timer never gets its turn; the retransmission timer owns the connection, and the outcome is 938 s, the same as rung 1. This is specified behavior, not a quirk: “Keep-alive packets MUST only be sent when no sent data is outstanding” (RFC 9293 §3.8.4). Keepalive covers idle connections; it does not cover a writer.

Rung 4: TCP_USER_TIMEOUT

TCP_USER_TIMEOUT (Linux 2.6.37 and later) is the per-connection R2 that RFC 9293 requires: the maximum time in milliseconds that transmitted data may remain unacknowledged, or buffered data may remain untransmitted, before the connection is closed with ETIMEDOUT. With 5000 ms, same dead peer, same write:

[    1.01s] ss: ESTAB 0      100  10.0.0.1:57044  10.0.0.2:9000 timer:(on,660ms,2)
[    2.01s] ss: ESTAB 0      100  10.0.0.1:57044  10.0.0.2:9000 timer:(on,1.304sec,3)
[    3.01s] ss: ESTAB 0      100  10.0.0.1:57044  10.0.0.2:9000 timer:(on,300ms,3)
[    4.02s] ss: ESTAB 0      100  10.0.0.1:57044  10.0.0.2:9000 timer:(on,1.400sec,4)
[    5.02s] ss: ESTAB 0      100  10.0.0.1:57044  10.0.0.2:9000 timer:(on,396ms,4)
[    5.46s] recv() failed: errno=110 Connection timed out

The retransmissions still happen on their normal schedule (“the option has no effect on when TCP retransmits a packet”); only the give-up point moves. The check happens at a retransmission, so it fires a little after 5 s, not at exactly 5.000 s. The man page also says that with keepalive enabled, TCP_USER_TIMEOUT overrides keepalive to decide when to close.

What it does not cover: an idle reader. The option is about transmitted data staying unacknowledged. In the lab, an idle reader with a 5000 ms user timeout was still blocked after 60 s with no error, exactly like rung 0.

It cuts both ways, too. A small user timeout will also kill a connection that would have survived a short outage; the man page notes that longer timeouts “allow a TCP connection to survive extended periods without end-to-end connectivity”.

Rung 5: an application heartbeat

None of the kernel mechanisms checks that the program on the other end is alive and responsive: a wedged process behind a healthy kernel still ACKs everything. And a middlebox that terminates TCP separately on each side can answer for a dead backend. The only complete check lives in the application protocol: send a small message regularly and require some inbound byte within a deadline.

[    3.00s] peer lost
[    6.01s] 3 s without any inbound byte -> close and reconnect

In the lab the client pings every second and treats any inbound data as proof of life. The peer was lost at 3.00 s and the client gave up at 6.01 s, 3.0 s later. In the worst case, detection takes about the silence deadline plus one ping interval (the loop only checks once per interval), and it is independent of the kernel’s retransmission state, because the check is a timer in your program, not an event from the socket.

Bonus rung: the textbook half-open case

If the peer rebooted (or otherwise forgot the connection) and the path is open again, the RST arrives on the first write. The lab reproduces it by dropping the packets, making the server discard the connection, and then restoring the network:

[    0.00s] peer lost (all packets dropped)
[    0.31s] network is back, but the peer has no such connection any more
[    5.32s] client idle 5 s, nothing noticed: True  | ss: ESTAB 0      0     10.0.0.1:37880     10.0.0.2:9000
[    5.32s] write() returned: the bytes only reached the send buffer
[    5.32s] recv() failed: errno=104 Connection reset by peer

An idle client notices nothing for as long as it stays idle, then learns instantly when it sends. That is what RFC 9293 §3.5.1 predicts. It also shows why the rungs above are about silence, not about resets: resets are the easy case.

The options side by side

MechanismDetectsMissesTime to detect (lab)Cost
Nothingnothingeverything silentreader: not within 60 s; writer: 938 s (about 15.6 minutes)none
SO_KEEPALIVE (tuned)dead peer on an idle connectionany connection with unacknowledged data4.6 s (2 s idle, 1 s × 3)a few tiny packets per interval; per-socket options
TCP_USER_TIMEOUTdead peer while data is unacknowledgedidle readers5.5 s (5000 ms)none on the wire; short values kill connections in short outages
Application heartbeatdead peer, wedged peer, black hole, for idle and busy socketsnothing in the path it can measure3.0 s (1 s pings, 3 s silence deadline)one small message per interval, each way

The mechanisms stack: a heartbeat for correctness, TCP_USER_TIMEOUT so that a stuck writer fails fast, and keepalive on the server side to reap clients that disappeared.

Where NAT fits

A NAT keeps a mapping per connection and discards it after an idle period. If you are idle for longer than that, your next packet meets a NAT with no memory of you. RFC 5382 REQ-5 says the “established connection idle-timeout” “MUST NOT be less than 2 hours 4 minutes” when the NAT cannot tell whether the connection is alive. Its justification is the arithmetic of the default keepalive: “applications can send keep-alive packets at the default rate (every 2 hours) such that the NAT can passively determine that the connection is alive. The additional 4 minutes allows time for in-flight packets to cross the NAT.”

That is a requirement for NATs that follow the RFC. Nothing guarantees that every NAT or carrier-grade gateway does, and I have not measured any real network’s timeout; I will not give you a number. What follows for design:

  • Treat the idle timeout as unknown and per-network. Make the heartbeat interval configurable, and keep it well below whatever you expect the shortest timeout on your path to be.
  • The default TCP_KEEPIDLE of two hours is exactly the value that protects nothing against an aggressive NAT. If you rely on keepalive packets to hold a mapping open, lower the idle time.
  • A heartbeat doubles as keep-the-mapping-alive traffic, which is another reason to run it even when the kernel features are available.

The lab

Everything above was produced by this script. It creates a second network namespace, connects it with a veth pair, and “loses” the peer by taking its end of the pair down, so every packet towards it is dropped silently. Needs: Linux, Python 3, iproute2 (ip, ss), util-linux (unshare, nsenter), and unprivileged user namespaces enabled. No root.

#!/usr/bin/env python3
"""deadpeer.py - how long does a TCP client take to notice that its peer is gone?

Linux only. Needs unprivileged user namespaces (no real root). Usage:
    python3 deadpeer.py <scenario>
Scenarios: silent-read | keepalive-read | user-timeout-read | keepalive-write |
           user-timeout-write | default-write | heartbeat | reset-on-write
The peer is "lost" by taking its end of a veth pair down, so every packet to it
is dropped without any ICMP or RST - what a vanished NAT mapping or a dead radio
link looks like from the client.
"""
import os, select, signal, socket, struct, subprocess, sys, threading, time

PEER = ("10.0.0.2", 9000)

def sh(*cmd, ns=None):
    pre = ["nsenter", "-t", str(ns), "-n"] if ns else []
    return subprocess.run(pre + list(cmd), check=True, capture_output=True, text=True).stdout

def setup_lab():
    """Create the second network namespace and the veth pair. Returns the namespace's pid."""
    holder = subprocess.Popen(["unshare", "-n", "sleep", "3600"]); time.sleep(0.3)
    sh("ip", "link", "add", "c0", "type", "veth", "peer", "name", "s0")
    sh("ip", "link", "set", "s0", "netns", str(holder.pid))
    sh("ip", "addr", "add", "10.0.0.1/24", "dev", "c0"); sh("ip", "link", "set", "c0", "up")
    sh("ip", "addr", "add", "10.0.0.2/24", "dev", "s0", ns=holder.pid)
    sh("ip", "link", "set", "s0", "up", ns=holder.pid)
    mac = sh("ip", "-o", "link", "show", "s0", ns=holder.pid).split("link/ether ")[1].split()[0]
    # pin the neighbour entry, so a dead peer is not reported as an ARP failure ("No route to host")
    sh("ip", "neigh", "replace", PEER[0], "lladdr", mac, "dev", "c0", "nud", "permanent")
    return holder

SERVER = r'''
import socket, signal, struct, sys
echo = sys.argv[1] == "echo"
s = socket.socket(); s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
s.bind(("10.0.0.2", 9000)); s.listen(); c, _ = s.accept()
def forget(*_):                      # close with RST ... which the lab will drop
    c.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER, struct.pack("ii", 1, 0)); c.close(); raise SystemExit
signal.signal(signal.SIGUSR1, forget)
while True:
    data = c.recv(65536)
    if not data: break
    if echo: c.sendall(b"PONG\n")
'''

def main(scenario):
    holder = setup_lab()
    srv = subprocess.Popen(["nsenter", "-t", str(holder.pid), "-n", "python3", "-c", SERVER,
                            "echo" if scenario == "heartbeat" else "sink"])
    time.sleep(0.5)
    link = lambda state: sh("ip", "link", "set", "s0", state, ns=holder.pid)
    ss = lambda: " | ".join(l.strip() for l in sh("ss", "-tno", "dst", PEER[0]).splitlines()[1:]) or "(socket gone)"
    t0 = time.monotonic()
    def log(msg): print(f"[{time.monotonic() - t0:8.2f}s] {msg}", flush=True)
    def err(e): return f"errno={e.errno} {os.strerror(e.errno or 0)}"

    c = socket.socket()
    if scenario.startswith("keepalive"):          # probe after 2 s idle, every 1 s, give up after 3 misses
        c.setsockopt(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1)
        c.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, 2)
        c.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, 1)
        c.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, 3)
    if scenario.startswith("user-timeout"):
        c.setsockopt(socket.IPPROTO_TCP, socket.TCP_USER_TIMEOUT, 5000)   # ms
    c.connect(PEER); c.sendall(b"hello"); time.sleep(0.5)
    t0 = time.monotonic()

    if scenario == "heartbeat":                       # application-level liveness (see the article)
        c.setblocking(False); last_rx = time.monotonic(); lost = False
        while True:
            if not lost and time.monotonic() - t0 >= 3: link("down"); lost = True; log("peer lost")
            try: c.send(b"PING\n")
            except BlockingIOError: pass
            if select.select([c], [], [], 1.0)[0]:
                try: data = c.recv(100)
                except OSError: data = b""
                if data: last_rx = time.monotonic()
            if time.monotonic() - last_rx > 3:
                log("3 s without any inbound byte -> close and reconnect"); break
    else:
        link("down"); log("peer lost (all packets dropped)")
        if scenario == "reset-on-write":              # the peer also forgets the connection, then comes back
            srv.send_signal(signal.SIGUSR1); time.sleep(0.3); link("up")
            log("network is back, but the peer has no such connection any more")
            idle = not select.select([c], [], [], 5)[0]
            log(f"client idle 5 s, nothing noticed: {idle}  | ss: {ss()}")
        if scenario.endswith("write") or scenario == "reset-on-write":
            c.sendall(b"x" * 100); log("write() returned: the bytes only reached the send buffer")
        stop = threading.Event()
        def sample():
            while not stop.wait(30 if scenario == "default-write" else 1): log("ss: " + ss())
        if scenario in ("default-write", "keepalive-write", "user-timeout-write"):
            threading.Thread(target=sample, daemon=True).start()
        wait = 60 if scenario in ("silent-read", "user-timeout-read") else 3600
        if select.select([c], [], [], wait)[0]:
            try: c.recv(1); log("recv() returned")
            except OSError as e: log(f"recv() failed: {err(e)}")
        else:
            log(f"still blocked after {wait} s: no data, no EOF, no error")
        stop.set()
    srv.kill(); holder.kill()

if __name__ == "__main__":
    if os.geteuid() != 0:                             # become root inside new user + network namespaces
        os.execvp("unshare", ["unshare", "-Urn", sys.executable] + sys.argv)
    main(sys.argv[1])

Two details worth knowing. The neighbour (ARP) entry is pinned, because otherwise the dead peer is reported as No route to host from a failed ARP lookup instead of as a retransmission timeout, which would measure a different thing. And “peer lost” means dropped packets, which approximates a vanished NAT mapping or a dead radio link; it does not reproduce a real NAT’s behaviour, nor mobile radio states.

When to use what

  • A server holding many idle client connections: turn on SO_KEEPALIVE with an idle time and interval you chose, so dead clients are reaped. RFC 1122 §4.2.3.6 describes exactly this use: server applications that “might otherwise hang indefinitely and consume resources unnecessarily if a client crashes”.
  • A client with a long-lived connection that mostly reads (subscriptions, push channels): an application heartbeat is required. Kernel features cannot cover a reader behind NAT, and keepalive defaults are far too slow.
  • A writer that must fail fast: add TCP_USER_TIMEOUT, and choose a value longer than a plausible transient outage.
  • Short-lived request/response connections: use per-request deadlines; none of this is needed.

Mistakes that bring the long timeouts back

  • Tuning keepalive and assuming writers are covered. Symptom: a stuck writer takes minutes to fail despite aggressive keepalive (rung 3). Fix: TCP_USER_TIMEOUT or an application deadline.
  • Keepalive idle time longer than the NAT’s idle timeout. Symptom: connections that die after a quiet period, then hang. Fix: heartbeat or keepalive idle time below the shortest timeout you expect, configurable per deployment.
  • A single missed heartbeat treated as death. TCP already retransmits lost segments, so give the deadline room for a few round trips, and reset it on any inbound byte, not only on the expected reply.
  • A deadline that is extended by later ticks. Measure from when the probe was sent; a second tick must not push the deadline further out.
  • A heartbeat answered by an intermediary. If a proxy replies on the backend’s behalf, you have proven the proxy is alive, not the service. Put the check in a message that the endpoint itself must answer.
  • Treating ETIMEDOUT as fatal for the user. It means this connection is gone, not that the service is. Reconnect with exponential backoff and jitter so a mass failure does not become a thundering herd.
  • User timeout set too short. It converts every brief outage into a reconnect.

Try it yourself

  1. Save the script as deadpeer.py and run python3 deadpeer.py silent-read. After about a minute you should see the “still blocked” line.
  2. Run keepalive-read and then keepalive-write. Compare: the first ends in a few seconds; the second keeps retransmitting (watch the ss samples, one per second there). It takes minutes, so let it run, or stop it with Ctrl-C.
  3. Run user-timeout-write, then change 5000 and watch the give-up point follow it.
  4. Run heartbeat, then change the 1 s ping interval and the 3 s deadline and predict the new detection time before running it.
  5. Run reset-on-write and confirm that the idle client sees nothing until it writes.
  6. In a second terminal, run ss -tno while a scenario is running and watch the timer column.

What I measured, and what I did not

All times are from one Linux 6.12 sandbox with default sysctls (tcp_retries2 15, tcp_keepalive_* 7200/75/9) and a virtual link with sub-millisecond RTT; numbers will differ on your kernel and your network. The two slow cases (a writer with default settings, and a writer with aggressive keepalive) both failed after 938.4 s in the final script; an earlier version of the script, run separately, gave 939.0 s. The short cases were run two to four times each and varied by up to about 0.2 s (the TCP_USER_TIMEOUT case gave 5.44 to 5.65 s). I read the RFC and man-page passages quoted above, the kernel constants in include/net/tcp.h, and the tcp_retries2 entry of the kernel’s ip-sysctl documentation (the 924.6 s lower bound above). I did not test: real NAT idle timeouts, mobile radio behavior, other operating systems, IPv6, or TLS/WebSocket proxies. The statement that a middlebox can answer a heartbeat on behalf of a dead backend is reasoning from how proxies terminate connections, not something I measured.

How long will you believe a dead connection?

  • Silence is the only evidence TCP has about a dead peer, and the kernel only collects it when it has something outstanding. A reader gets no signal at all.
  • Writers fail only after the retransmission give-up time: 938 s (about 15.6 minutes) here, and 13–30 minutes by the man page for the default. Keepalive cannot shorten it, because keepalive only runs when no data is outstanding.
  • Use TCP_USER_TIMEOUT to bound the writer, keepalive to reap idle clients on servers, and an application heartbeat for everything that has to be right, including NAT idle timeouts you cannot know in advance.
  • The most important question to ask of any long-lived connection is not “how do I detect disconnects?” but “what is the longest time I’m willing to believe a dead connection is alive?” Then pick the mechanism that enforces it.