TL;DR
- On a connection that carries many requests, a request deadline should cancel that request (for HTTP/2,
RST_STREAM). In my lab, closing the whole session on one timeout made two healthy requests fail; cancelling the one stream let both finish on the same TCP connection. - That rule only holds if replies can be matched to requests by an id. With matching by position, a late reply was handed to the next request, which received the previous request’s answer and no error. I reproduced this on a WebSocket-based request/response layer.
- Cancelling a stream does not help when the stall is below the streams. In a Linux network-namespace lab, a 100 ms server-to-client blackout on one TCP connection left every stream sharing it silent for roughly 220 to 320 ms (two sessions of five runs). With one connection per stream, the unlucky stream went silent for about 130 ms and the other kept its normal 10 ms cadence.
- So there are two different questions: “is this request too slow?” (answer per stream) and “is this connection still alive?” (answer with a connection-level probe such as HTTP/2
PING, with your own deadline). Close the connection only for the second. - Everything below runs with Node, plus root,
ip netnsandnftfor the packet-loss lab.
Two things that look alike
A client opens one connection to a server and sends many requests over it at the same time. HTTP/2 does this by design, and many WebSocket or RPC protocols do it by convention. One request is slow and its deadline fires. What now?
There are two honest answers, and they cost very different amounts:
- Close the connection. Simple, and it surely releases everything. Every other request in flight on it dies too, and the next request pays for a new TCP (and TLS) handshake.
- Cancel only that request. The connection stays up and the others carry on. You have to be sure the connection is healthy, because a deadline on one request tells you nothing about the others.
Let me show each branch with something you can run.
The duel: close the session, or cancel the stream
HTTP/2 specifies a way to end one request without touching the rest. RST_STREAM “allows for immediate termination of a stream” and “is sent to request cancellation of a stream” (RFC 9113 §6.4). After sending one, the sender must still be ready for frames from the peer that were already on the way. Those can be ignored, except frames that change connection state such as header compression or flow control (§5.4.2). The connection is meant to survive.
The lab uses Node’s built-in http2 module over plain TCP. A server answers GET /<ms> after that many milliseconds. A client sends one request that needs 3000 ms with a 500 ms deadline, and two requests that need 800 ms with generous deadlines. The two strategies differ in a single line.
// One HTTP/2 connection, several requests, one of them is slow.
// What should the client close when the slow one hits its deadline: the stream, or the whole session?
import http2 from "node:http2";
const { NGHTTP2_CANCEL } = http2.constants;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const server = http2.createServer(); // cleartext HTTP/2 is enough for this
let sessions = 0; let serverLog = [];
server.on("session", (s) => { sessions++; });
server.on("stream", (stream, headers) => {
const delay = Number(headers[":path"].slice(1)); // GET /<milliseconds>
stream.on("close", () => { if (stream.rstCode) serverLog.push(`server saw RST_STREAM(code ${stream.rstCode}) for the /${delay} stream`); });
setTimeout(() => { if (!stream.destroyed) { stream.respond({ ":status": 200 }); stream.end(`done after ${delay} ms`); } }, delay);
});
await new Promise((r) => server.listen(0, r));
const url = `http://localhost:${server.address().port}`;
async function run(strategy) {
sessions = 0; serverLog = [];
const session = http2.connect(url);
session.on("error", () => {});
await new Promise((r) => session.on("connect", r));
const t0 = Date.now();
const request = (ms, deadline) => new Promise((resolve) => {
const req = session.request({ ":path": `/${ms}` });
let body = "";
const timer = setTimeout(() => {
if (strategy === "close-stream") req.close(NGHTTP2_CANCEL); // RST_STREAM(CANCEL) for this stream only
else session.destroy(); // tear down the TCP connection
resolve(`TIMEOUT after ${Date.now() - t0} ms`);
}, deadline);
req.on("data", (d) => (body += d));
req.on("error", () => {});
req.on("close", () => { clearTimeout(timer); resolve(body.startsWith("done") ? `ok (${body})` : "FAILED: closed without a response"); });
});
// one slow request (3 s) with a 500 ms deadline, two others that need 800 ms and have 2 s deadlines
const results = await Promise.all([request(3000, 500), request(800, 2000), request(800, 2000)]);
console.log(` strategy "${strategy}": slow -> ${results[0]} | other #1 -> ${results[1]} | other #2 -> ${results[2]} | TCP connections opened: ${sessions}`);
await sleep(300);
console.log(` ${serverLog.length ? serverLog.sort().join("; ") : "server saw no RST_STREAM"}`);
session.destroy(); await sleep(100);
}
console.log("timeout handling on a shared HTTP/2 connection");
await run("close-session");
await run("close-stream");
server.close();
timeout handling on a shared HTTP/2 connection
strategy "close-session": slow -> TIMEOUT after 505 ms | other #1 -> FAILED: closed without a response | other #2 -> FAILED: closed without a response | TCP connections opened: 1
server saw RST_STREAM(code 8) for the /800 stream
strategy "close-stream": slow -> TIMEOUT after 501 ms | other #1 -> ok (done after 800 ms) | other #2 -> ok (done after 800 ms) | TCP connections opened: 1
server saw RST_STREAM(code 8) for the /3000 stream
(Node 20.19.2.) In the close-stream run the server saw exactly one RST_STREAM, for the slow stream, with error code 8, which is CANCEL, and only the slow request was lost. In the close-session run, the two healthy requests were cut off at about 505 ms, before their 800 ms of work was done, and a real client would now need a new connection. (That run’s server log shows a single RST_STREAM line for one of the two /800 streams. I did not capture frames or work out why, so nothing above rests on it; the result is the client-side outcome.)
That settles the cost side of the duel. It does not settle the safety side, which depends on how your client matches replies to requests.
Cancelling is only safe if replies carry an id
HTTP/2 gives each request a stream id, so you get this for free. Many other protocols running over a single connection (a WebSocket with a JSON protocol, an RPC channel, a line protocol) do not, and people match replies by order of arrival. Here is that failure, reproduced.
The server in this lab can reply “in order”, like a pipelined protocol: the reply to request B cannot go out before the reply to request A, even if B is ready. Client 1 correlates by position. Client 2 puts an id in every request and in every reply. Each fires A (slow, 2000 ms), then B and C (50 ms each), all with a 500 ms deadline, and then a fresh request D.
// Correlating replies on one shared connection: by position (FIFO) or by request id.
// One slow request (A), then two quick ones (B, C), 500 ms deadline each. Run: node mux_lab.mjs
import { WebSocketServer, WebSocket } from "ws";
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// --- server: one handler per connection; "ordered" replies in request order (like a pipelined protocol), "tagged" replies as soon as ready
const wss = new WebSocketServer({ port: 0 });
await new Promise((r) => wss.on("listening", r));
wss.on("connection", (ws) => {
let chain = Promise.resolve();
ws.on("message", (raw) => {
const m = JSON.parse(raw);
const work = sleep(m.delay).then(() => JSON.stringify({ id: m.id, tag: m.tag }));
if (m.ordered) chain = chain.then(() => work).then((out) => ws.send(out)); // must wait for everything before it
else work.then((out) => ws.send(out));
});
});
const url = `ws://127.0.0.1:${wss.address().port}`;
const open = async () => { const c = new WebSocket(url); await new Promise((r) => c.on("open", r)); return c; };
// --- client 1: replies carry no id, the n-th reply belongs to the n-th request
async function positional() {
const c = await open(); const waiting = []; let late = 0;
c.on("message", (raw) => { const w = waiting.shift(); if (w) w.resolve(JSON.parse(raw).tag); else late++; });
const call = (tag, delay, deadline = 500) => new Promise((resolve) => {
const w = { resolve };
waiting.push(w);
c.send(JSON.stringify({ tag, delay, ordered: true }));
setTimeout(() => { const i = waiting.indexOf(w); if (i >= 0) { waiting.splice(i, 1); resolve("TIMEOUT"); } }, deadline);
});
const first = await Promise.all([call("A", 2000), call("B", 50), call("C", 50)]);
console.log("positional, first round :", first.join(" "));
const d = await call("D", 10, 3000); // a new request, while A's reply is still on its way
console.log(`positional, request D : asked for D, got "${d}"`);
await sleep(500);
c.close();
}
// --- client 2: replies carry the request id; the timeout removes exactly that id
async function tagged() {
const c = await open(); const pending = new Map(); let nextId = 1, orphans = 0;
c.on("message", (raw) => { const m = JSON.parse(raw); const w = pending.get(m.id); if (w) { pending.delete(m.id); w(m.tag); } else orphans++; });
const call = (tag, delay, deadline = 500) => new Promise((resolve) => {
const id = nextId++;
pending.set(id, resolve);
c.send(JSON.stringify({ id, tag, delay }));
setTimeout(() => { if (pending.delete(id)) resolve("TIMEOUT"); }, deadline);
});
const first = await Promise.all([call("A", 2000), call("B", 50), call("C", 50)]);
console.log("by id, first round :", first.join(" "));
const d = await call("D", 10, 3000);
console.log(`by id, request D : asked for D, got "${d}"`);
await sleep(2500);
console.log(`by id, afterwards : late replies dropped as orphans: ${orphans}, pending map size: ${pending.size}`);
c.close();
}
await positional(); await tagged(); wss.close();
positional, first round : TIMEOUT TIMEOUT TIMEOUT
positional, request D : asked for D, got "A"
by id, first round : TIMEOUT B C
by id, request D : asked for D, got "D"
by id, afterwards : late replies dropped as orphans: 1, pending map size: 0
Two separate failures hide in the positional row:
- B and C also time out. The server had to finish A first, so B and C were stuck behind it. This is application-layer head-of-line blocking, the thing HTTP/2’s own introduction says pipelined HTTP/1.1 “still suffers from” (RFC 9113 §1). Ids cannot make an ordered server faster. What they allow is a server that answers as soon as a reply is ready, and the server in the id run does that.
- D received A’s answer. After the timeouts, the client removed its waiting entries, then A’s late reply arrived and was handed to the next waiter, D. That is a wrong answer, delivered successfully. No error anywhere.
The id-based client lost only A, and when A’s reply finally arrived it was recognised as an orphan and dropped: the final pending map was empty. The cost of this safety is one small map and a rule: on timeout, delete the id; on a reply with an unknown id, drop it quietly.
If your protocol has no ids and you cannot add them, the only safe response to a timeout is to close the connection. That is an argument for adding ids, not for closing connections more.
When cancelling does not help: the stall is below the streams
Everything so far assumed the connection itself is healthy. HTTP/2 multiplexes all streams onto one TCP byte stream, and TCP delivers bytes in order. If one packet is lost, all the bytes after it wait, whichever stream they belong to. RFC 9113 says it plainly: “TCP head-of-line blocking is not addressed by this protocol” (§1). QUIC avoids this: “only streams with data in that packet are blocked waiting for a retransmission”, though packets that carry data from several streams block all of them (RFC 9000 §13). I did not test HTTP/3.
To see it, I built a small network: two Linux network namespaces joined by a virtual Ethernet pair, a server in one, a client in the other. The server streams two responses (/a and /b), each a 1000-byte chunk every 10 ms. At about 500 ms into the run, nft drops all packets from the server to the client’s connection #1 for 100 ms (DROP=100). The client records the longest silence in each stream. In “single” mode, both streams share connection #1. In “two” mode, each has its own connection, and only stream a is on the connection that gets the loss.
// Two long responses (/a, /b), each a 1000-byte chunk every 10 ms for 1.5 s. Plain-text HTTP/2.
import http2 from "node:http2";
const server = http2.createServer();
server.on("stream", (stream) => {
stream.respond({ ":status": 200 });
const chunk = Buffer.alloc(1000, 1); let n = 0;
const t = setInterval(() => { if (stream.destroyed || ++n > 150) { clearInterval(t); stream.end(); } else stream.write(chunk); }, 10);
});
server.listen(9200, "10.1.0.2", () => console.log("server up"));
// mode "single": both requests share one HTTP/2 connection. mode "two": one connection per request.
// Around t = 500 ms, packets from the server to connection #1 are silently dropped for 30 ms.
import http2 from "node:http2";
import { execFile } from "node:child_process";
import { promisify } from "node:util";
const mode = process.argv[2];
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const run = promisify(execFile);
const nft = (...a) => run("nft", a); // async: do not block the event loop while measuring
const DROP_MS = Number(process.argv[3] ?? 30);
await nft("add", "table", "inet", "hol"); await nft("add", "chain", "inet", "hol", "in", "{ type filter hook input priority 0; }");
const s1 = http2.connect("http://10.1.0.2:9200");
const s2 = mode === "single" ? s1 : http2.connect("http://10.1.0.2:9200");
await Promise.all([s1, s2].map((s) => new Promise((r) => s.once("connect", r))));
const port1 = s1.socket.localPort; // the loss will hit this connection only
const stats = {};
function fetch(session, name) {
return new Promise((resolve) => {
const req = session.request({ ":path": `/${name}` });
const st = (stats[name] = { last: performance.now(), maxGap: 0, at: 0, bytes: 0 });
const t0 = performance.now();
req.on("data", (d) => { const now = performance.now(); const gap = now - st.last; if (gap > st.maxGap) { st.maxGap = gap; st.at = st.last - t0; } st.last = now; st.bytes += d.length; });
req.on("end", resolve); req.resume?.();
});
}
const done = Promise.all([fetch(s1, "a"), fetch(s2, "b")]);
await sleep(500);
await nft("add", "rule", "inet", "hol", "in", "ip", "saddr", "10.1.0.2", "tcp", "sport", "9200", "tcp", "dport", String(port1), "drop");
await sleep(DROP_MS);
await nft("flush", "chain", "inet", "hol", "in");
await done;
for (const [n, st] of Object.entries(stats)) console.log(` stream ${n}: longest silence ${st.maxGap.toFixed(0).padStart(4)} ms (starting ${st.at.toFixed(0)} ms in), ${st.bytes} bytes`);
s1.destroy(); s2.destroy();
#!/bin/bash
# root needed. Two namespaces joined by a veth pair; server in "sv", client in "cl".
NODE=$(command -v node); HERE=$(cd "$(dirname "$0")" && pwd)
ip netns del sv 2>/dev/null; ip netns del cl 2>/dev/null
ip netns add sv; ip netns add cl
ip link add name vsv type veth peer name vcl
ip link set vsv netns sv; ip link set vcl netns cl
ip -n sv addr add 10.1.0.2/24 dev vsv; ip -n sv link set vsv up; ip -n sv link set lo up
ip -n cl addr add 10.1.0.1/24 dev vcl; ip -n cl link set vcl up; ip -n cl link set lo up
ip netns exec sv $NODE $HERE/server.mjs & SRV=$!
sleep 1
for round in 1 2 3 4 5; do
for mode in single two; do
echo " [$mode connection(s)] round $round"
ip netns exec cl $NODE $HERE/client.mjs $mode ${DROP:-30}
ip netns exec cl nft delete table inet hol 2>/dev/null
done
done
kill $SRV; ip netns del sv; ip netns del cl
Five rounds on Node 20.19.2, as printed (each cell is the longest gap, in milliseconds, between two data events of that stream):
| Mode | Stream a (on the lossy connection) | Stream b |
|---|---|---|
| Single shared connection | 315, 224, 274, 258, 260 | 316, 223, 269, 219, 225 |
| One connection per stream | 129, 132, 133, 126, 132 | 12, 13, 13, 12, 13 |
The 10 ms cadence of the server is why the healthy stream shows 12 to 13 ms. With one connection, a 100 ms outage became a silence of 219 to 316 ms for both streams. With two connections, only the stream on the damaged connection suffered. I did not look inside TCP, so I cannot say which retransmission timer fired in which round, or explain the spread from 219 to 316 ms. A re-run of the 100 ms case on another day (five more rounds) gave 218 to 225 ms for both streams on the shared connection, and 126 to 139 ms and 13 to 15 ms with separate connections. The shape held; the spread did not. I would not quote any of these figures as a property of TCP, and the lab is Linux-specific (network namespaces, nft, and the Linux retransmission behaviour). A shorter outage (DROP=30) behaved differently: both streams on one connection went quiet for 40 to 52 ms, about the length of the outage, and the damaged stream in the two-connection run did the same (40 to 51 ms) while its neighbour stayed at 12 to 13 ms. So the cost of one lost stretch of packets is not simply its length, and the sharing effect is in who else has to wait.
Now suppose a request deadline fires during this stall. RST_STREAM for that request frees nothing here: the frame travels over the very connection that is not delivering. And all your other streams are just as quiet. The deadline is a symptom. You need a connection-level question.
Asking the connection: PING with your own deadline
HTTP/2 has a frame for this. PING can be used “for measuring a minimal round-trip time from the sender, as well as determining whether an idle connection is still functional” (RFC 9113 §6.7). The protocol does not set a deadline for the reply. You choose one.
// Is this HTTP/2 connection alive? PING frame + our own deadline. Then the path is black-holed.
import http2 from "node:http2";
import { execFile } from "node:child_process";
import { promisify } from "node:util";
const run = promisify(execFile);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
function alive(session, ms) { // resolves true/false, never rejects
return new Promise((resolve) => {
const timer = setTimeout(() => resolve(false), ms);
session.ping((err, rttMs) => { clearTimeout(timer); resolve(err ? false : rttMs); });
});
}
const session = http2.connect("http://10.1.0.2:9200");
session.on("error", () => {});
await new Promise((r) => session.once("connect", r));
const t0 = performance.now();
console.log("healthy path : alive() ->", await alive(session, 1000) !== false ? "true" : "false", `(PING round trip ${(await alive(session, 1000)).toFixed(2)} ms)`);
await run("nft", ["add", "table", "inet", "bh"]); await run("nft", ["add", "chain", "inet", "bh", "in", "{ type filter hook input priority 0; }"]);
await run("nft", ["add", "rule", "inet", "bh", "in", "ip", "saddr", "10.1.0.2", "drop"]);
const t1 = performance.now();
const ok = await alive(session, 1000);
console.log(`black-holed : alive() -> ${ok !== false} after ${(performance.now() - t1).toFixed(0)} ms (deadline 1000 ms)`);
session.destroy();
Running run2.sh (same network setup, but no streaming; nft drops every packet from the server after the first PING):
healthy path : alive() -> true (PING round trip 0.12 ms)
black-holed : alive() -> false after 1001 ms (deadline 1000 ms)
The second line shows nothing about HTTP/2; it shows that the probe returns when my timer fires, and not earlier. A connection that stays silent will not tell you that it is dead, and that is why you need to ask with a deadline. How long to wait is a trade-off between fast detection and false alarms on a loaded or high-latency path. The right value depends on your network, so measure round trips on yours.
Options at a glance
| Option | What you gain | What it costs |
|---|---|---|
| Close the connection on every timeout | Trivial to implement | Collateral failures (shown above); reconnect cost; reconnection storms if many requests time out together |
Cancel the stream (HTTP/2 RST_STREAM), keep the connection | Only the slow request is lost | Server must stop the work (see below); useless if the stall is below the streams |
| Cancel the stream, then PING when several unrelated requests time out together | Same as above, plus a way to detect a dead path | One more timer to tune |
| More than one connection per origin | A lost packet hurts only the streams on that connection | More handshakes and memory; the server’s per-connection limits |
| HTTP/3 (QUIC) | Loss on one stream does not block the others (RFC 9000 §13) | Not tested here; availability depends on your stack and network |
How timeouts go wrong
| What you see | Usual cause | What to do |
|---|---|---|
| One slow endpoint causes bursts of unrelated failures; connection counts spike | The whole session is closed on every timeout (reproduced above) | Cancel the request; close the connection only on a failed probe, a GOAWAY or a protocol error |
| A request returns another request’s data, with no error | Replies matched by position (reproduced above) | Add ids; drop replies with an unknown id |
| Memory creeps up under load; late replies resolve promises nobody waits on | The pending map is not cleaned up on timeout (in the id run above it ended at 0) | Load-test with a mix that makes some requests time out |
| CPU and database load stay high though clients gave up | The server keeps working after the cancel. RST_STREAM only says the stream is over; the handler has to look (in h2_timeout.mjs it checks stream.destroyed before answering). Not measured at scale | Check for cancellation in every handler that does real work |
| Reconnect loops that make a slow server slower | Several timeouts taken as proof that the connection is dead | PING first |
| More load on a server that is already slow | Instant retries on a degraded path (not measured here; it is the usual rule, listed because it appears with the first two rows) | Back off with jitter and cap the retries |
| Nothing, until it matters | The assumption that your HTTP client cancels the stream when you abort. I did not check any specific library or browser | Look in a server log, or capture traffic, before relying on it |
Reproduce it
- Save
h2_timeout.mjsand runnode h2_timeout.mjs. Success: theclose-sessionline shows two FAILED, theclose-streamline shows twook, andTCP connections openedis 1 in both. npm i ws@8, savemux_lab.mjs, and run it. Success:positionalshowsasked for D, got "A", andby idshowsTIMEOUT B Candgot "D".- On Linux with root,
ipandnft: putserver.mjs,client.mjs,pingtest.mjs,run.shandrun2.shin one directory. Runsudo DROP=100 bash run.shfor the stall table, andsudo bash run2.shfor the PING lab. (Ifnftis not in root’sPATH, add/usr/sbin.) The scripts create and delete two namespaces,svandcl. - Change
DROPfrom 100 to 30 and to 300, and predict the silence before you look. - Change the deadline in
mux_lab.mjsso that A finishes in time, and check that all three replies line up in both clients.
How far these results go
Verified on one machine (Linux 6.12, Node 20.19.2): the outputs above, as printed. The packet-loss lab ran five rounds per mode at each of two drop lengths (100 ms and 30 ms), plus a later five-round re-run of the 100 ms case; the PING lab ran once and was reproduced once. The RFC 9113 and RFC 9000 sections cited were read at rfc-editor.org while writing.
Not verified: any real network or middlebox; TLS; HTTP/3; browsers and HTTP client libraries (what they do on abort); the exact cause of the 219 to 316 ms spread; server behaviour under real load after a cancel. The deadline values (500 ms, 1000 ms) are lab choices, not recommendations. Earlier trials of the loss lab with a blocking call to nft gave misleading numbers because the call blocked the event loop; the scripts here use the asynchronous version.
A deadline belongs to a request
A deadline belongs to a request, and the request has an id. Cancel what timed out. Ask the connection separately, with PING and a deadline of your own, and close it only when it fails to answer.