TL;DR
- With refresh token rotation, every refresh returns a new refresh token and invalidates the old one. If an invalidated token is presented again, the server cannot tell the thief from the legitimate client, so per RFC 9700 §4.14.2 it revokes the active token and the user must authenticate again.
- A client that sends several requests in parallel, gets several
401s, and refreshes once per401presents the same refresh token several times. Only the first wins. The rest look like a replay. In the toy server below, the naive client ends up with a revoked token family every time I ran it. - Fix on the client: single-flight refresh. All callers wait on one shared refresh. Add a stale check so a late
401does not start a second refresh, and retry each request at most once. - Fix on the server (a trade-off, not a free lunch): a short grace window in which the previous token is still accepted.
- The demo is 93 lines of dependency-free Node. It prints the failing case and the three fixes side by side.
Background: what rotation protects, and what it costs
Public clients such as mobile apps and single-page apps cannot keep a secret, so the refresh token sitting on the device is a valuable thing to steal. RFC 9700 §4.14.2 says authorization servers “MUST utilize one of these methods to detect refresh token replay by malicious actors for public clients”: sender-constrained refresh tokens (RFC 8705 or RFC 9449), or refresh token rotation.
Rotation, as the RFC describes it: the server issues a new refresh token with every refresh response; the previous one is invalidated, but the server remembers the relationship. If a token is stolen and then used by both the attacker and the legitimate client, one of them will present an invalidated token. The RFC is explicit about the cost: “The authorization server cannot determine which party submitted the invalid refresh token, but it will revoke the active refresh token. This stops the attack at the cost of forcing the legitimate client to obtain a fresh authorization grant.”
Now picture an ordinary client. The access token expired while the app was in the background. The user opens a screen that fires three API calls at once. All three go out with the expired access token. All three come back 401. If each 401 handler thinks “refresh, then retry”, three refresh requests carry the same refresh token. The server sees one valid use and two replays, and it is right to be alarmed: from where it sits, this is indistinguishable from theft.
Break it first
Here is the “obvious” handler (it appears in the demo below as strategy naive):
let r = await send(accessToken);
if (r.status === 401) {
await refresh(); // every concurrent 401 gets here
r = await send(accessToken);
}
Against a toy server with rotation and reuse detection, three parallel calls end like this (output is real; which request wins varies from run to run):
naive: every 401 refreshes | burst: ERR,200,ERR | POST /token: 3 | after next expiry: ERR (refresh failed: 400)
Read it left to right. One call succeeds, two fail. Three refresh requests were sent. And the damage shows up later: when the new access token expires, the client’s refresh token has been revoked along with its family, the refresh returns 400, and the user is signed out. The failure is delayed, intermittent and depends on timing, which is exactly why it survives testing on a fast machine with one request at a time.
Options considered
| Approach | What it does | Trade-offs |
|---|---|---|
| Single-flight refresh (client) | Concurrent callers share one in-flight refresh. | Fully client-side; the right default. Covers one process only (see failure modes). |
| Gate requests during a refresh (client) | While a refresh is running, new requests wait instead of sending the old token. | Avoids a wave of doomed 401s. Adds a queue and a place to hang. Complements single-flight rather than replacing it. |
| Proactive refresh before expiry (client) | Refresh a little early based on the token’s lifetime. | Fewer 401s, but not zero: timers are not guaranteed to fire in a suspended or throttled process, and clocks drift. You still need the 401 path. |
| Grace window (server) | The just-rotated token stays valid for a short period. | Absorbs concurrency and lost responses. Widens the replay window by exactly that period. Some providers make the length configurable. That is a threat-model decision, not a bug fix. |
| Sender-constrained tokens (server and client) | Bind the token to a key, so replay by another party fails by construction (RFC 9449, RFC 8705). | The other branch of the RFC 9700 requirement. A bigger change; this post stays with rotation. |
If you operate only the client, single-flight plus a one-time retry is the answer. If you also operate the server, decide the grace window deliberately and write down why.
How single-flight works
Single-flight is a small idea: store the Promise, not the result.
let inflight = null;
const refreshOnce = () =>
(inflight ??= doRefresh().finally(() => { inflight = null; }));
The first caller finds inflight empty, starts the refresh and stores its promise. Every caller that arrives before it settles gets the same promise and waits on it. finally clears the slot on success and on failure, so a failed refresh does not poison later attempts. JavaScript’s single thread makes the check-and-store race-free; in a multi-threaded language you need a mutex or an equivalent primitive.
That alone leaves one gap. A request that was sent with the old access token but whose 401 arrives after the shared refresh finished finds inflight empty and starts a second refresh. It uses the newest refresh token, so it is not a replay, but it is a wasted rotation. The stale check closes it: remember which access token a request used, and when its 401 arrives, compare that with the current one. If they differ, someone already renewed, so skip the refresh and just retry.
if (state.access !== tokenThatFailed) return; // already renewed by someone else
Two more rules complete the picture. Retry once: if the retried request is still 401, surface the error instead of looping. And treat the pair “new access token and new refresh token” as one unit that is stored together, since a response can carry both.
The demo: a toy server and four strategies
The script has two halves. The first is a toy server with the behavior RFC 9700 describes (rotation, reuse detection that revokes the family, and an optional grace window). The second is the client under test with three reactions to a 401. The third request in each burst is deliberately slow so that its 401 arrives after the first refresh has finished. Save it as refresh_demo.mjs and run node refresh_demo.mjs (Node 18 or newer, no packages):
// Refresh-token rotation vs. concurrent 401s. Run: node refresh_demo.mjs (Node >= 18, no dependencies)
import http from 'node:http';
// ---- A toy authorization/resource server: rotation + reuse detection (+ optional grace window) ----
function makeServer({ graceMs = 0 } = {}) {
const s = { seq: 0, access: new Set(), refresh: new Map(), used: new Map(), revoked: new Set(), tokenCalls: 0 };
const issue = (family) => {
const n = ++s.seq;
s.access.add(`a${n}`);
s.refresh.set(`r${n}`, family);
return { access: `a${n}`, refresh: `r${n}` };
};
s.first = () => { const t = issue('family-1'); s.access.delete(t.access); return t; }; // access token starts out expired
const srv = http.createServer((req, res) => {
let body = '';
req.on('data', (d) => (body += d));
req.on('end', async () => {
await new Promise((r) => setTimeout(r, Number(req.headers['x-delay'] ?? 20))); // network + processing
if (req.url === '/api') {
const token = (req.headers.authorization ?? '').replace('Bearer ', '');
res.statusCode = s.access.has(token) ? 200 : 401;
return res.end('{}');
}
s.tokenCalls++; // POST /token
const { refresh } = JSON.parse(body);
const family = s.refresh.get(refresh);
if (family && !s.revoked.has(family)) { // normal rotation
s.refresh.delete(refresh);
s.used.set(refresh, { family, at: Date.now() });
return res.end(JSON.stringify(issue(family)));
}
const old = s.used.get(refresh);
if (old && Date.now() - old.at < graceMs && !s.revoked.has(old.family)) { // grace window
return res.end(JSON.stringify(issue(old.family)));
}
if (old) s.revoked.add(old.family); // reuse detected: revoke everything
res.statusCode = 400;
res.end('{"error":"invalid_grant"}');
});
});
s.listen = () => new Promise((ok) => srv.listen(0, '127.0.0.1', () => { s.url = `http://127.0.0.1:${srv.address().port}`; ok(); }));
s.close = () => srv.close();
return s;
}
// ---- The client under test: three ways to react to a 401 ----
function makeClient(server, strategy, tokens) {
const state = { ...tokens }; // { access, refresh }
let inflight = null; // the one refresh currently running, if any
async function doRefresh() {
const r = await fetch(server.url + '/token', { method: 'POST', body: JSON.stringify({ refresh: state.refresh }) });
if (!r.ok) throw new Error('refresh failed: ' + r.status);
Object.assign(state, await r.json()); // access AND refresh token are replaced together
}
function refresh(tokenThatFailed) {
if (strategy === 'naive') return doRefresh(); // every 401 refreshes
if (strategy === 'single-flight+stale-check' && state.access !== tokenThatFailed)
return Promise.resolve(); // someone already renewed it
return (inflight ??= doRefresh().finally(() => { inflight = null; })); // join the refresh in progress
}
return {
state,
async call(delay = 20) {
const send = (token) => fetch(server.url + '/api', { headers: { authorization: 'Bearer ' + token, 'x-delay': String(delay) } });
const used = state.access;
let r = await send(used);
if (r.status !== 401) return r.status;
await refresh(used); // may throw
return (await send(state.access)).status; // retry exactly once
},
};
}
async function scenario(title, strategy, opts) {
const server = makeServer(opts); await server.listen();
const client = makeClient(server, strategy, server.first());
// three requests start with the expired token; the third is slow, so its 401 arrives after the refresh finished
const results = await Promise.allSettled([client.call(), client.call(), client.call(150)]);
const burst = results.map((r) => (r.status === 'fulfilled' ? r.value : 'ERR')).join(',');
const tokenCalls = server.tokenCalls;
server.access.clear(); // later, the new access token expires too
const next = await client.call().then(String, (e) => `ERR (${e.message})`);
console.log(`${title.padEnd(32)} | burst: ${burst.padEnd(11)} | POST /token: ${tokenCalls} | after next expiry: ${next}`);
server.close();
}
await scenario('naive: every 401 refreshes', 'naive');
await scenario('naive + 2 s server grace window', 'naive', { graceMs: 2000 });
await scenario('single-flight', 'single-flight');
await scenario('single-flight + stale check', 'single-flight+stale-check');
One run:
naive: every 401 refreshes | burst: ERR,200,ERR | POST /token: 3 | after next expiry: ERR (refresh failed: 400)
naive + 2 s server grace window | burst: 200,200,200 | POST /token: 3 | after next expiry: 200
single-flight | burst: 200,200,200 | POST /token: 2 | after next expiry: 200
single-flight + stale check | burst: 200,200,200 | POST /token: 1 | after next expiry: 200
How to read the four rows:
- naive: three refresh requests, two replays. The family is revoked, so the next expiry ends in
400. (In other runs the winning request was a different one.) - naive + grace window: still three refresh requests, but the server tolerates the replays for 2 s, so everyone succeeds. The cost sits on the server side, in the widened replay window.
- single-flight: the first two
401s share one refresh. The slow third401arrives later and starts a second one, which is valid but unnecessary. - single-flight + stale check: exactly one refresh. This is the shape to ship.
The toy server is mine, not a real provider’s. It demonstrates the mechanism, not any vendor’s exact behavior.
Where single-flight applies
- Use single-flight whenever the server rotates refresh tokens and your app can have more than one request in flight. That is nearly every app.
- Without rotation the failure above cannot happen, so single-flight is an optimization (fewer calls) rather than a correctness fix. It is still cheap.
- It is not enough when several processes or contexts share one stored refresh token. See the next section.
- Do not add a grace window just to hide a client bug. Fix the client first; keep the window for the failure you cannot fix on the client (lost responses).
What still goes wrong after the fix
- More than one tab, worker or process. Each has its own
inflight, so they replay each other’s token. Symptom: random sign-outs only when the app is open twice or background work runs. Fix: one cross-context lock around “read token, refresh, write token”. In browsers, the Web Locks API lets scripts in multiple tabs or workers of one origin coordinate, e.g.navigator.locks.request('token-refresh', async () => { /* re-read stored token; refresh only if still stale */ }). I did not run that snippet in this sandbox (it needs a browser). Elsewhere, let one component own the token store. - The refresh response is lost. The server rotated, but the client never received the new token (timeout, process killed). The next refresh presents the old token and looks like a replay. Symptom: sign-outs on flaky networks. Fix: this is inherent to rotation; a grace window is how providers cope. Also persist the new tokens before you use them.
- Signing the user out on any refresh error. A timeout or a
5xxis not a revoked grant. RFC 6749 §5.2 definesinvalid_grantfor a refresh token that is “invalid, expired, revoked”. Fix: sign out oninvalid_grant; retry with backoff on network errors. - Retrying a request whose body cannot be replayed. A fetch
Requestbody is one-shot;clone()throws if the body was already used. Fix: build the request from data on each attempt, not from a consumed object. - Unbounded retries. A permanently rejected token loops through refresh forever. Fix: retry once per request.
- Stale check against the wrong value. Compare the token the failed request used, not the token that happens to be current when you read the code. The demo captures it in
const used = state.accessbefore sending.
Break the demo on purpose
- Run
node refresh_demo.mjsa few times. The naive row should fail every time; only the identity of the winning request changes. - In
makeClient, delete the stale-check line and confirm the single-flight row’sPOST /tokencount goes from 1 to 2. - Change the grace window in the second scenario (
{ graceMs: 2000 }) to0and it fails again. In my runs1passed in only some runs, while5and50passed in the three runs I tried. Real replays are spread out by network latency, so a tight window protects less than it appears to; choose the window from the failure you need to absorb, not from a local test. - Change the slow call
client.call(150)toclient.call(20). The third401now arrives before the refresh finishes, so plain single-flight needs only one refresh (I saw 1 in three runs) and the stale check makes no difference. The stale check only matters for late401s.
What I ran, and what I did not run
I ran the script five times; the naive row failed in all of them, and the single-flight and stale-check rows succeeded in all of them with 2 and 1 refresh requests. The toy server follows the RFC 9700 §4.14.2 description. I did not test any real identity provider, mobile OS, or browser, and I did not run the Web Locks snippet.
Rules to ship with
- Rotation makes reuse of a refresh token an alarm, and parallel
401s make a legitimate client trigger it. - Share one refresh among concurrent callers (store the promise), compare tokens to skip late refreshes, and retry once.
- If you control the server, a grace window is a deliberate, bounded concession for lost responses, not a fix for client concurrency.
- Single-flight is per process. For several tabs or processes, add a cross-context lock or a single owner for the token store.