A timeout answers exactly one question: did a reply arrive before my deadline? It says nothing about whether the request ran. Send a request that adds one item to an order, wait 30 milliseconds, hear nothing, and three different histories fit what you saw.
TL;DR
- A timeout is consistent with three worlds: the request never arrived, the request arrived and was applied but the response was lost, or the request is still being processed and you stopped waiting too early. The sender cannot tell them apart, and no protocol on an unreliable link can make them distinguishable (the Two Generals problem).
- So “exactly-once delivery” is not a property you can add to a link. What you can build is at-least-once delivery plus a receiver that deduplicates, which gives exactly-once effects inside a boundary you control.
- The mechanism is an idempotency key: chosen by the client once per intent, before the first send, reused on every retry, and recorded by the server atomically with the effect.
- In a test with 15% of requests lost and a client timeout shorter than some handler runs: retrying without a key duplicated about half the requests, retrying with a key duplicated none. A key store that writes its record after the effect still duplicated 89 of 200 in the run I made of that variant.
- Retries need backoff with jitter, a cap on attempts, and respect for
Retry-After. Without jitter, 1000 clients that failed together retry together, every round.
Three worlds behind one timeout
| What happened | Server state | What the client saw |
|---|---|---|
| The request was lost on the way | Nothing applied | Timeout |
| The request was applied; the response was lost | Applied once | Timeout |
| The request is slow; the client stopped waiting | Will be applied once | Timeout |
The right reaction differs. In the first world you must resend, or the work is lost. In the second you must not resend, or the work is done twice. In the third, resending races with the original.
This is the Two Generals problem, first published in 1975 and given its name in 1978: two parties that can communicate only over an unreliable channel cannot reach certain agreement that both have acted. Its practical corollary is that acknowledgements cannot end the regress, because the acknowledgement of the acknowledgement can be lost as well.
HTTP’s own specification draws the line in the right place. RFC 9110 §9.2.2 says that PUT, DELETE and safe methods are idempotent, so they can be repeated automatically after a communication failure “even if the original request succeeded”; and “a client SHOULD NOT automatically retry a request with a non-idempotent method unless it has some means to know that the request semantics are actually idempotent, regardless of the method, or some means to detect that the original request was never applied.” The rest of this article is those two “means”.
What “exactly once” really means in systems that claim it
Systems that advertise exactly-once semantics are not defeating the impossibility; they are deduplicating inside a boundary. Apache Kafka’s producer documentation is explicit: the idempotent producer “strengthens Kafka’s delivery semantics from at least once to exactly once delivery”, it works by the broker tracking a producer ID and sequence numbers (delivery semantics), and “the producer can only guarantee idempotence for messages sent within a single session”; applications are told to avoid “application level re-sends since these cannot be de-duplicated” (KafkaProducer javadoc). That is the whole recipe: identity for each message, memory at the receiver, and a stated scope.
The practical statement of the equation is: exactly-once processing = at-least-once delivery + idempotent effects.
Experiment: three clients, one flaky network
The script below starts a local HTTP server whose handler takes up to 60 ms and applies a side effect (increments a per-intent counter). The “network” drops 15% of requests before they arrive, and the client gives up after 30 ms, so a good share of requests are applied after the client has stopped listening. Three clients try to apply 200 intents each:
- never retry: send once.
- retry, no key: up to 6 attempts, full-jitter backoff.
- retry, same key: the same, but with an
Idempotency-Keyfixed per intent; the server records the key and replays the stored result for a duplicate.
once.mjs
// Three clients, one flaky network, one side effect that must not happen twice.
// Requires Node 18+ (global fetch). Run: node once.mjs (counts vary slightly between runs: real timers)
import http from "node:http";
const INTENTS = 200; // distinct things the "user" wants to happen exactly once
const CONCURRENCY = 10;
const applied = new Map(); // intent id -> how many times the side effect really ran
const seen = new Map(); // idempotency key -> { fingerprint, promise } (the dedupe store)
const server = http.createServer(async (req, res) => {
let body = ""; for await (const c of req) body += c;
const key = req.headers["idempotency-key"];
const send = (code, obj) => { res.writeHead(code, { "content-type": "application/json" }); res.end(JSON.stringify(obj)); };
const run = async () => { // the side effect
await new Promise((r) => setTimeout(r, Math.random() * 60)); // slow enough that clients sometimes give up first
const { intent } = JSON.parse(body);
applied.set(intent, (applied.get(intent) ?? 0) + 1);
return { status: 201, intent };
};
if (!key) return send(201, await run()); // no key: every request is a new request
const fingerprint = body;
const prev = seen.get(key);
if (prev && prev.fingerprint !== fingerprint) return send(422, { error: "key reused with a different request" });
if (prev) { const r = await prev.promise; return send(r.status, r); } // duplicate or concurrent duplicate: share the one result
const promise = run();
seen.set(key, { fingerprint, promise }); // record BEFORE awaiting, so concurrent duplicates find it
const r = await promise;
send(r.status, r);
});
// A flaky network: drops 15% of requests before they arrive; the client gives up after 30 ms.
const call = async (intent, key) => {
if (Math.random() < 0.15) throw new Error("request lost");
const headers = { "content-type": "application/json", ...(key && { "idempotency-key": key }) };
const res = await fetch(url, { method: "POST", headers, body: JSON.stringify({ intent }), signal: AbortSignal.timeout(30) });
if (res.status >= 500) throw new Error("server error");
return res.status;
};
const jitter = (n, base = 5, cap = 100) => Math.random() * Math.min(cap, base * 2 ** n); // full jitter
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const clients = {
"never retry": async (i) => { try { await call(i); } catch {} },
"retry, no key": async (i) => { for (let n = 0; n < 6; n++) { try { return await call(i); } catch { await sleep(jitter(n)); } } },
"retry, same key": async (i) => { const key = `intent-${i}`; for (let n = 0; n < 6; n++) { try { return await call(i, key); } catch { await sleep(jitter(n)); } } },
};
await new Promise((r) => server.listen(0, "127.0.0.1", r));
const url = `http://127.0.0.1:${server.address().port}/`;
for (const [name, client] of Object.entries(clients)) {
applied.clear(); seen.clear();
let next = 0; // a small worker pool, so that we measure the protocol and not connection storms
await Promise.all(Array.from({ length: CONCURRENCY }, async () => { while (next < INTENTS) await client(`${name}-${next++}`); }));
await sleep(200); // let slow in-flight requests finish
const counts = [...applied.values()];
console.log(`${name.padEnd(16)} lost=${INTENTS - counts.length} applied once=${counts.filter((c) => c === 1).length} duplicated=${counts.filter((c) => c > 1).length}`);
}
server.close();
One run (Node.js 20.19, Linux 6.12; ten concurrent workers; real timers, so your numbers will differ run to run):
never retry lost=32 applied once=168 duplicated=0
retry, no key lost=0 applied once=95 duplicated=105
retry, same key lost=0 applied once=200 duplicated=0
Across repeated runs the picture was the same: with no key about half the intents were applied twice or more, with a key none were, and never retrying lost roughly 10–20% of the work. (“Lost” in the first line is only the requests dropped before arrival. The client also timed out on many that were applied anyway; it just did not know, which is the point.)
The bug that looks like a fix
Look at the order in the server handler:
const promise = run();
seen.set(key, { fingerprint, promise }); // record BEFORE awaiting
const r = await promise;
The key is recorded before the effect finishes, and a duplicate that arrives while the first is still running awaits the same promise. Move the seen.set after await run() and the code still “has a dedupe store”, but concurrent duplicates both miss it. I made that change and ran it:
retry, same key lost=0 applied once=111 duplicated=89
Eighty-nine of 200 intents were duplicated by a client that was sending the key correctly. (The number moves with timing: three later runs of the same variant gave 75, 96 and 90.) The window is exactly the case the client’s timeout creates: the first request is still running when the retry arrives.
The idempotency-key contract
A key is a promise between client and server. The server’s side is a small state machine:
| Request | Server does | Status |
|---|---|---|
| Key never seen | Record the key and run the operation | 2xx |
| Key seen, same request, finished | Replay the stored response; do not run again | The original status |
| Key seen, same request, still running | Wait for it, or tell the client to come back | 200/2xx after waiting, or 409 |
| Key seen, different request body | Refuse: the client is misusing the key | 422 |
These status choices follow the (expired) IETF draft for an Idempotency-Key header, which proposed 409 for a request that is still in progress and 422 for key reuse with a different payload. As of 2026-10-04 the Datatracker lists draft-ietf-httpapi-idempotency-key-header-07 as an expired Internet-Draft (last updated 2026-04-18), not an RFC. Treat it as a convention to compare against, not as a standard.
The client’s side matters just as much:
- Generate the key once per intent, before the first send. Not per attempt. A key generated inside the retry loop is a different key every time, which is no key.
- Persist the key with the pending operation. If the app restarts between attempts, the retry after restart must reuse the key. Store it next to the queued request.
- Never reuse a key for a different intent. The server’s 422 is the safety net, not the plan.
- Send byte-identical retries. Or at least semantically identical ones, if the server fingerprints the payload.
The key and the effect must commit together
An in-memory Map is enough for the experiment, but a real server can crash between “apply” and “record”. If the effect is committed and the key is not, the retry runs the effect again; if the key is committed and the effect is not, the retry replays a response for something that never happened. The remedy is to make them one atomic step. With a relational database that means the same transaction:
atomic_dedupe.py (SQLite, Python standard library only)
# The dedupe record and the side effect must commit together. Run: python3 atomic_dedupe.py
import sqlite3
db = sqlite3.connect(":memory:", isolation_level=None) # we manage transactions ourselves
db.executescript("""
CREATE TABLE balance (id INTEGER PRIMARY KEY, amount INTEGER);
INSERT INTO balance VALUES (1, 0);
CREATE TABLE idempotency (key TEXT PRIMARY KEY, response TEXT);
""")
def charge(key, amount, crash_before_commit=False):
db.execute("BEGIN IMMEDIATE")
row = db.execute("SELECT response FROM idempotency WHERE key = ?", (key,)).fetchone()
if row: # duplicate: replay the stored answer, do nothing else
db.execute("COMMIT")
return row[0]
db.execute("UPDATE balance SET amount = amount + ? WHERE id = 1", (amount,)) # the side effect
if crash_before_commit:
db.execute("ROLLBACK") # the process dies here: neither the effect nor the key survives
raise RuntimeError("crashed before commit")
db.execute("INSERT INTO idempotency VALUES (?, ?)", (key, f"charged {amount}"))
db.execute("COMMIT")
return f"charged {amount}"
def balance():
return db.execute("SELECT amount FROM balance").fetchone()[0]
try:
charge("k1", 100, crash_before_commit=True)
except RuntimeError as e:
print("first attempt:", e, "| balance =", balance())
print("retry :", charge("k1", 100), "| balance =", balance())
print("duplicate :", charge("k1", 100), "| balance =", balance())
first attempt: crashed before commit | balance = 0
retry : charged 100 | balance = 100
duplicate : charged 100 | balance = 100
The “crash” rolls back both the update and the key record, so the retry applies the effect once; the duplicate after that finds the key and replays. (This uses SQLite’s BEGIN IMMEDIATE, which serializes writers. On other databases the usual pattern is a unique constraint on the key, with the insert as part of the same transaction; I did not run that variant here.)
This is also where the scope of the guarantee ends. If the effect is a call to a third party (an email, a payment provider), your transaction cannot roll it back. The only defence is to pass the idempotency identity along: forward the same key, or a key derived from it, to the next hop. Idempotency is an end-to-end property, and a hop that does not honor the key breaks it for everyone upstream.
Retries: backoff, jitter, and budgets
Retrying immediately turns one slow server into a retry storm. Exponential backoff spaces the attempts out; jitter makes clients that failed at the same moment retry at different moments. The widely cited analysis is Exponential Backoff And Jitter, which calls the following variant “Full Jitter”:
$$ \text{sleep}_n = \mathrm{random}\bigl(0,\ \min(\text{cap},\ \text{base}\cdot 2^{n})\bigr) $$where \(n\) is the number of failed attempts so far. The effect is easy to see with a deterministic model: 1000 clients fail at the same instant and every retry fails again, so we only observe scheduling.
jitter.mjs
// 1000 clients fail at the same instant (say, the server restarts). When do their retries arrive?
// Deterministic: seeded PRNG. Run: node jitter.mjs
const rng = (a) => () => { a = (a + 0x6d2b79f5) | 0; let t = Math.imul(a ^ (a >>> 15), 1 | a); t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t; return ((t ^ (t >>> 14)) >>> 0) / 2 ** 32; };
const rand = rng(42);
const BASE = 100, CAP = 10_000, CLIENTS = 1000, ATTEMPTS = 5; // ms
const strategies = {
"no jitter": (n) => Math.min(CAP, BASE * 2 ** n),
"full jitter": (n) => rand() * Math.min(CAP, BASE * 2 ** n),
"equal jitter": (n) => { const d = Math.min(CAP, BASE * 2 ** n); return d / 2 + rand() * (d / 2); },
"decorrelated": (n, prev) => Math.min(CAP, BASE + rand() * ((prev ?? BASE) * 3 - BASE)),
};
for (const [name, delay] of Object.entries(strategies)) {
const buckets = new Map(); // 10 ms bucket -> number of retry requests
for (let c = 0; c < CLIENTS; c++) {
let t = 0, prev;
for (let n = 0; n < ATTEMPTS; n++) {
prev = delay(n, prev); t += prev; // every retry fails again, so we measure pure scheduling
const b = Math.floor(t / 10); buckets.set(b, (buckets.get(b) ?? 0) + 1);
}
}
const peak = Math.max(...buckets.values());
console.log(`${name.padEnd(13)} peak retries in one 10 ms bucket: ${String(peak).padStart(4)} (total ${CLIENTS * ATTEMPTS})`);
}
no jitter peak retries in one 10 ms bucket: 1000 (total 5000)
full jitter peak retries in one 10 ms bucket: 157 (total 5000)
equal jitter peak retries in one 10 ms bucket: 214 (total 5000)
decorrelated peak retries in one 10 ms bucket: 73 (total 5000)
Without jitter, all 1000 clients hit the server in the same 10 ms bucket on every round. With any jitter, the peak drops sharply. Do not rank the three jitter variants from this output: their first-round windows are different widths, so their peaks differ for that reason alone; the linked analysis compares them on completion time and total work.
Backoff is only half of a retry policy. The other half is knowing when to stop and when not to start:
- Retry only what is retryable. Network errors and timeouts,
503(withRetry-After, which you should honor), and429(RFC 6585 §4). Do not retry a4xxthat says the request is wrong. - Cap attempts and total time. A retry loop with no deadline is a leak.
- Retry at one layer. If three layers each make up to three attempts, the layer at the bottom can see up to 27 requests for one user action.
- Do not store transient failures as the final result. If the server saved a
503against the key, every retry would replay that failure forever. Store terminal outcomes.
Where idempotency still breaks
| Failure | Result | Defence |
|---|---|---|
| Key recorded after the effect | Concurrent duplicates both run (89 of 200 in my run) | Record first, or lock the key; let the duplicate wait for the first. |
| New key per attempt | No deduplication at all | Generate once per intent, before the first send. |
| Key not persisted across app restart | A duplicate after a crash | Store the key with the queued operation. |
| Same key, different payload | A wrong result is replayed | Fingerprint the request; answer 422. |
| Effect and key in separate commits | Crash window runs the effect twice or replays a phantom result | One transaction (or unique constraint inside the transaction). |
| Key retention shorter than the client’s retry window | An old retry runs again | Keep keys longer than the longest time a client may still retry. |
| Downstream hop ignores the key | Duplicates reappear below your boundary | Forward the key, or give the downstream its own idempotency identity. |
| Retries without jitter | Synchronized waves of load | Full jitter; cap attempts; honor Retry-After. |
| Retries at every layer | Multiplicative load (3 x 3 x 3 = 27) | Retry at one layer; pass deadlines down. |
| Transient failure stored as result | Permanent replayed error | Store only terminal outcomes. |
Who needs a key, and who does not
Use idempotency keys for any operation with a side effect that is not naturally idempotent and that a client may retry: creating things, charging, sending, enqueuing.
You may not need them when:
- The operation is already idempotent in the sense of RFC 9110 §9.2.2:
PUTthat replaces a resource,DELETEby identifier, reads. - The operation has a natural unique identity, in which case a unique constraint on that identity is your dedupe store.
- At-most-once is acceptable (best-effort telemetry): send once and move on.
- The effect has no retry path at all, because the user would rather see an error than wait.
Run it, then change one thing
# Node.js 18+ (I ran Node.js 20). No dependencies.
node once.mjs # takes a few seconds; the numbers differ per run
node jitter.mjs # deterministic
python3 atomic_dedupe.py
Then change one thing at a time. Move seen.set below the await in once.mjs and watch the duplicates return. Generate the key inside the retry loop. Raise the network’s loss rate to 50%. Set the client timeout to 200 ms so the lost-response case disappears and see which clients still differ.
The bet behind every retry
Every retry is a bet that the first attempt did not work. The idempotency key turns the bet into a no-op when you are wrong: the second attempt finds the first and returns its result. What it costs is memory at the receiver, one atomic commit, and the discipline to choose the key before the first send. What it buys is that “exactly once” is no longer a property of the network, but of a boundary you can name.
What the experiments verified, and what they did not
Everything ran on Linux 6.12 with Node.js 20.19 and Python 3.13. The experiment uses a single process, a local network and an in-memory dedupe store; the loss rate and timings are parameters I picked, not measurements of any real network. Counts in once.mjs vary between runs because it uses real timers. The jitter output is a scheduling model, not a load test. I did not test a PostgreSQL or MySQL version of the atomic dedupe, key expiry, or a third-party downstream. Statements about specifications and documentation are from the linked pages as of 2026-10-04.