Say a client goes offline for forty seconds and reconnects, and the server published three hundred events in that time. There are two easy answers, and both fail in ways you notice late. Send nothing, and the client is silently behind. Send the whole history, and the cost of every reconnect grows with the age of the stream.

This article starts from the second answer (version 0) and fixes its problems one at a time, up to a bounded log with a snapshot fallback (version 3) and the rule that makes the snapshot safe (version 4). Along the way it checks what a browser’s EventSource really does, because the browser already ships half of the solution.

TL;DR

  • Resuming needs three things: a position the client can remember (a cursor), a log on the server that can replay from that position, and a fallback for when the position is too old to replay (a snapshot).
  • Server-sent events give you the client half for free: the browser remembers the last id and sends it back as Last-Event-ID on reconnect. I confirmed this in Chrome 154. What it does not give you is the server half, or any way to say “I can’t resume you”.
  • A non-200 response ends a native EventSource for good, with no status or body visible to the page (confirmed in Chrome 154). So for browsers, signal “too old, start over” in-band: answer 200 and send a reset event that carries a snapshot. A 410 Gone is fine for clients you control.
  • A snapshot is a pair (state, cursor). Read the cursor first, then the state, and make replay idempotent. Read them the other way round and an update can vanish: in my toy model, 699 of 1000 trials.
  • A cursor is only safe if sequence numbers become visible in order. If writer A takes number 5 and commits after writer B has published 6, a reader whose cursor is 6 never sees 5.
  • The end-to-end test in this article (random connection kills, plus an outage longer than the log retention) ends with client and server in the same state, with zero gaps and zero duplicates.

Version 0: reconnect and read everything

The simplest correct design: on every connect, send the whole state, then the live events. There is no cursor, no log, nothing to get wrong. Its cost is proportional to the size of the history on every reconnect, which is fine for small data and terrible for large, because reconnects cluster exactly when the network is bad. Keep this version in your pocket: it is the fallback for every later version.

Version 1: a cursor

Give every event a sequence number that increases by one in the order the server publishes them. The client remembers the last number it applied, and on reconnect asks for everything after it: GET /events?after=42. The server answers from a log.

Three rules decide whether this works:

  1. The sequence is monotonic and gapless per stream, so a client can detect a missing event by looking at numbers.
  2. The cursor is opaque to the client. It stores and returns it, nothing else. That leaves you free to change what it encodes.
  3. Numbers become visible in order. This is the one people miss. If sequence numbers are allocated when a write starts but become visible when it commits, then writer A can take 5, writer B can take 6 and commit first, and a reader that has seen 6 moves its cursor past 5 for good:
visibility_gap.mjs
// A cursor is only safe if sequence numbers become visible in order.
// Writer A takes number 5 and is slow to publish; writer B takes 6 and publishes first. Run: node visibility_gap.mjs
const published = [];                                   // what readers can see, in publication order
const publish = (seq, data) => published.push({ seq, data });
const readAfter = (cursor) => published.filter((e) => e.seq > cursor);

publish(4, "x");                                        // everything up to 4 is visible
// writer A allocated seq 5 but has not committed yet
publish(6, "y");                                        // writer B allocated 6 and committed first

let cursor = 4;
const first = readAfter(cursor);                        // the reader polls: sees only 6
cursor = Math.max(cursor, ...first.map((e) => e.seq));  // cursor advances to 6
publish(5, "z");                                        // writer A commits late
const second = readAfter(cursor);                       // the reader polls again with cursor 6

console.log("first read :", first.map((e) => e.seq));
console.log("second read:", second.map((e) => e.seq), "<- event 5 is never delivered");
first read : [ 6 ]
second read: [] <- event 5 is never delivered

The model is deliberately abstract: it shows the shape of the bug, not any particular database. The remedies are all ways of making visibility follow allocation order: allocate the number at the moment of publication (a single writer, or a lock), or have readers stop at the lowest number that might still be in flight.

Version 2: let the browser carry the cursor

Server-sent events standardize exactly this client behavior. The HTML Standard says: an id: field sets the connection’s last event ID; the value persists until the server sets another; on reconnect the browser sends a Last-Event-ID request header with it; an id containing a NUL character is ignored; and retry: sets the reconnection delay. The EventSource client is a cursor store with a built-in reconnect loop.

I did not want to take that on trust, so I checked it in a real browser: a page opens an EventSource, the server sends events 1–3 and then drops the connection, and the server records the header on the reconnect.

eventsource_check.mjs
// What does a real browser EventSource do on reconnect, and on a non-200 response?
// Run: node eventsource_check.mjs   (needs a Chrome/Chromium binary; set CHROME=/path if needed)
import http from "node:http";
import { spawn } from "node:child_process";

const seenHeaders = [];   // Last-Event-ID values the server received on /sse
let attempts410 = 0;

const page = `<script>
const out = { messages: [], sseErrors: 0, gone: { errors: 0, readyState: null } };
const es = new EventSource("/sse");
es.onmessage = (e) => { out.messages.push(e.lastEventId); };
es.onerror = () => { out.sseErrors++; };
const gone = new EventSource("/gone");
gone.onerror = () => { out.gone.errors++; out.gone.readyState = gone.readyState; };
setTimeout(() => { fetch("/report", { method: "POST", body: JSON.stringify(out) }); }, 2500);
</script>`;

const server = http.createServer(async (req, res) => {
  if (req.url === "/") { res.writeHead(200, { "content-type": "text/html" }); return res.end(page); }
  if (req.url === "/sse") {
    const last = req.headers["last-event-id"]; seenHeaders.push(last ?? null);
    res.writeHead(200, { "content-type": "text/event-stream" });
    res.write("retry: 100\n\n");
    const start = last ? Number(last) + 1 : 1;
    for (let id = start; id < start + 3; id++) res.write(`id: ${id}\ndata: hello\n\n`);
    if (!last) setTimeout(() => res.destroy(), 100);          // first connection: drop it after 3 events
    return;
  }
  if (req.url === "/gone") { attempts410++; res.writeHead(410, { "content-type": "application/json" }); return res.end('{"reset":true}'); }
  if (req.url === "/report") {
    let b = ""; for await (const c of req) b += c;
    console.log("page reported:", b);
    console.log("Last-Event-ID headers seen by /sse:", JSON.stringify(seenHeaders));
    console.log("connection attempts to the 410 endpoint:", attempts410);
    res.end("ok"); clearTimeout(guard); chrome.kill(); server.close(); server.closeAllConnections();
  }
});
await new Promise((r) => server.listen(0, "127.0.0.1", r));
const chrome = spawn(process.env.CHROME ?? "google-chrome", [
  "--headless=new", "--no-sandbox", "--disable-gpu", "--user-data-dir=./es-check-profile",
  `http://127.0.0.1:${server.address().port}/`,
], { stdio: "ignore" });
const guard = setTimeout(() => { console.log("timeout"); chrome.kill(); process.exit(1); }, 20000);
page reported: {"messages":["1","2","3","4","5","6"],"sseErrors":1,"gone":{"errors":1,"readyState":2}}
Last-Event-ID headers seen by /sse: [null,"3"]
connection attempts to the 410 endpoint: 1

The first connection carried no header; the reconnect carried 3; the page then received 4, 5, 6. That is half of the resume protocol with no client code at all. (Headless Chrome 154 on Linux; other browsers were not tested.)

The second half of the output matters most for design. The /gone endpoint answered 410. The browser tried once, fired error, and readyState became 2 (CLOSED): it did not retry. The spec says so: if the response status is not 200, or the content type is not text/event-stream, “fail the connection”, and “Once the user agent has failed the connection, it does not attempt to reconnect.” The page never sees the status code or the body.

So in a browser you cannot use 410 to say “your cursor is too old”. You have to say it inside a 200 response, in the stream.

Version 3: bounded log, in-band reset

A log cannot grow forever. Keep the last N events (or the last N minutes) and a client that stays away longer than that cannot be resumed from the log. The server must notice and say so. The design that works for both native EventSource and fetch-based clients:

  • Client sends its cursor (Last-Event-ID).
  • If cursor + 1 >= oldest retained, replay events after the cursor, then tail the live log.
  • Otherwise, send event: reset with the current snapshot and its cursor, then continue from that cursor.
  • A connection with no cursor takes the same path as “too old”.

Here is a complete implementation with a client that follows the same rules as EventSource (remember the last id, send it on reconnect). Retention is 50 events so that the test can overrun it.

resume.mjs: server, client and an end-to-end test with induced failures
// A resumable event stream in ~100 lines: cursor + bounded log + snapshot fallback.
// Requires Node 18+. Run: node resume.mjs
import http from "node:http";

const RETAIN = 50;                                   // the server keeps only the last 50 events
const log = [];                                      // [{ seq, key, value }]
const state = {};                                    // key -> { value, seq }   (the "current" truth)
let seq = 0;

const write = (key, value) => {                      // every state change appends to the log
  const ev = { seq: ++seq, key, value };
  state[key] = { value, seq };
  log.push(ev);
  if (log.length > RETAIN) log.shift();
  return ev;
};
const snapshot = () => ({ seq, state: structuredClone(state) });   // (cursor, state) read together: no await in between

const sse = (id, event, data) => `id: ${id}\nevent: ${event}\ndata: ${JSON.stringify(data)}\n\n`;

export const server = http.createServer((req, res) => {
  res.writeHead(200, { "content-type": "text/event-stream", "cache-control": "no-store" });
  const header = req.headers["last-event-id"];
  let cursor = header === undefined ? null : Number(header);

  const oldest = log.length ? log[0].seq : seq + 1;  // first sequence number still available
  if (cursor === null || !Number.isInteger(cursor) || cursor + 1 < oldest || cursor > seq) {
    const snap = snapshot();                          // cannot resume: tell the client in-band, then continue from the snapshot
    res.write(sse(snap.seq, "reset", snap));
    cursor = snap.seq;
  }
  for (const ev of log) if (ev.seq > cursor) { res.write(sse(ev.seq, "put", ev)); cursor = ev.seq; }

  const timer = setInterval(() => {                   // tail the log
    for (const ev of log) if (ev.seq > cursor) { res.write(sse(ev.seq, "put", ev)); cursor = ev.seq; }
  }, 5);
  req.on("close", () => clearInterval(timer));
});

// ---- a minimal client that behaves like EventSource: remember the last id, send it back on reconnect ----
export async function follow(url, onEvent, { signal }) {
  let lastId = null;
  while (!signal.aborted) {
    try {
      const res = await fetch(url, { headers: lastId === null ? {} : { "last-event-id": String(lastId) }, signal });
      let buf = "";
      for await (const chunk of res.body.pipeThrough(new TextDecoderStream())) {
        buf += chunk;
        for (let i; (i = buf.indexOf("\n\n")) >= 0; ) {
          const raw = buf.slice(0, i); buf = buf.slice(i + 2);
          const ev = Object.fromEntries(raw.split("\n").map((l) => [l.slice(0, l.indexOf(":")), l.slice(l.indexOf(":") + 2)]));
          lastId = Number(ev.id);
          onEvent(ev.event, JSON.parse(ev.data));
        }
      }
    } catch (e) { if (signal.aborted) return; }
    await new Promise((r) => setTimeout(r, 20));      // reconnect delay (EventSource's `retry`)
  }
}

if (import.meta.url === `file://${process.argv[1]}`) {
  await new Promise((r) => server.listen(0, "127.0.0.1", r));
  const url = `http://127.0.0.1:${server.address().port}/`;

  const mine = {}; let expected = null, gaps = 0, dups = 0, resets = 0, applied = 0;
  const ac = new AbortController();
  const done = follow(url, (type, d) => {
    if (type === "reset") { resets++; for (const k in mine) delete mine[k]; Object.assign(mine, d.state); expected = d.seq + 1; return; }
    if (expected !== null && d.seq < expected) { dups++; return; }
    if (expected !== null && d.seq > expected) gaps++;
    mine[d.key] = { value: d.value, seq: d.seq }; expected = d.seq + 1; applied++;
  }, { signal: ac.signal });

  // Producer: 400 writes, sometimes in bursts. Chaos: kill every open connection now and then,
  // and once keep the client away for long enough that the log moves past its cursor.
  for (let i = 0; i < 400; i++) {
    write(`k${i % 7}`, i);
    if (i % 60 === 59) server.closeAllConnections();                  // short outage: resume from the cursor
    if (i === 200) { server.closeAllConnections(); for (let j = 0; j < 120; j++) write(`k${j % 7}`, 1000 + j); } // long outage: > RETAIN events missed
    if (i % 5 === 0) await new Promise((r) => setTimeout(r, 3));
  }
  await new Promise((r) => setTimeout(r, 300));
  ac.abort(); await done; server.closeAllConnections(); server.close();

  const same = JSON.stringify(Object.keys(state).sort().map((k) => [k, state[k].value])) ===
               JSON.stringify(Object.keys(mine).sort().map((k) => [k, mine[k].value]));
  console.log(`server seq=${seq}  applied=${applied}  resets=${resets}  gaps=${gaps}  duplicates=${dups}  client state equals server state: ${same}`);
}

The test at the bottom produces 400 writes, kills every open connection every 60 writes, and once keeps the client away while 120 more events are written (more than the retention of 50):

server seq=520  applied=354  resets=2  gaps=0  duplicates=0  client state equals server state: true

resets=2 is the first connection (no cursor) plus the long outage; every short outage resumed from the cursor. gaps=0 and duplicates=0 are checks the client makes on sequence numbers as it applies events. In three runs the final state matched each time (applied was between 349 and 354 because the timing of the kills varies).

Is the retention check really doing anything? I removed the clause cursor + 1 < oldest from the server and ran it again: gaps=1. The client silently skipped events. In that run the final state still matched, because later writes happened to overwrite the lost ones, which is exactly why this bug survives testing. Count gaps in the client, not just final state.

The reset event is also where the client throws away what it has and installs the snapshot:

if (type === "reset") { for (const k in mine) delete mine[k]; Object.assign(mine, d.state); expected = d.seq + 1; return; }

Version 4: the snapshot is a pair

A snapshot is not “the state”. It is “the state as of cursor C”, and the pairing is the whole point. If the cursor does not describe the state, then the replay that starts at the cursor will leave a hole or a repeat. The rule is the same one you would use for any lock-free read of two things:

Read the cursor first, then the state. Make replay idempotent.

Reading the cursor first means the state is at least as new as the cursor, so replaying from the cursor can only repeat events. If replay is idempotent (set key to value, ignoring anything older than what you hold for that key), repeats are harmless. Reading the state first and the cursor second means the cursor can be newer than the state, and the events in between are never replayed.

snapshot_order.mjs
// A snapshot is a pair (state, cursor). Read them in the wrong order and an update can vanish.
// Run: node snapshot_order.mjs
const rng = (a) => () => { a = (a + 0x6d2b79f5) | 0; let t = Math.imul(a ^ (a >>> 15), 1 | a); t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t; return ((t ^ (t >>> 14)) >>> 0) / 2 ** 32; };

function trial(order, rand) {
  const log = [], state = {}; let seq = 0;
  const write = () => { const key = "k" + Math.floor(rand() * 3); log.push({ seq: ++seq, key, value: seq }); state[key] = { value: seq, seq }; };
  const writes = (n) => { for (let i = 0; i < n; i++) write(); };
  writes(3);

  // Reading the state and reading the cursor are two steps. Writes can land between them (0 to 3 here).
  let snapState, cursor;
  if (order === "state, then cursor") { snapState = structuredClone(state); writes(Math.floor(rand() * 4)); cursor = seq; }
  else                                { cursor = seq; writes(Math.floor(rand() * 4)); snapState = structuredClone(state); }
  writes(Math.floor(rand() * 2));                                   // and a few more after the snapshot

  // The client installs the snapshot, then replays every log entry after the cursor,
  // skipping entries that are not newer than what it already holds for that key.
  const mine = structuredClone(snapState);
  for (const ev of log) if (ev.seq > cursor && ev.seq > (mine[ev.key]?.seq ?? 0)) mine[ev.key] = { value: ev.value, seq: ev.seq };
  return JSON.stringify(Object.entries(mine).sort()) === JSON.stringify(Object.entries(state).sort());
}

for (const order of ["state, then cursor", "cursor, then state"]) {
  const rand = rng(7); let ok = 0;
  for (let i = 0; i < 1000; i++) ok += trial(order, rand) ? 1 : 0;
  console.log(`${order.padEnd(20)} client matches server in ${ok}/1000 trials`);
}
state, then cursor   client matches server in 301/1000 trials
cursor, then state   client matches server in 1000/1000 trials

The 301 depends on my toy interleaving (0–3 writes land between the two reads, 3 keys); the point is the zero versus the non-zero. resume.mjs avoids the problem a different way: snapshot() reads both in one synchronous step with no await in between, which is the single-process equivalent of reading them in one database transaction or one consistent snapshot.

Choosing a design

And some alternatives I considered:

OptionFitsDoes not fit
Re-read everythingSmall state; the repair pathLarge state; frequent reconnects
Cursor + bounded log + snapshotMutable state shown to many clientsStrict audit where a skipped event is unacceptable
HTTP Range requests (RFC 9110 §14)Byte-addressable resources: files, append-only blobsEvent streams where “position” is not a byte offset
Client keeps everything in memoryShort sessionsAnything where the app restarts
Level-triggered refetch (see the article on invalidation)State that can be re-read cheaply by keyLarge histories and ordered events

Ways a resume path breaks

FailureSymptomDefence
Cursor older than the logA silent gapDetect in the server; send reset with a snapshot (or 410 to clients that can read it).
Native EventSource gets a non-200It stops forever, with no status visibleAnswer 200 and signal in-band.
Snapshot state read before the cursorA lost updateCursor first, then state, or one consistent read; idempotent replay.
Sequence numbers visible out of orderA skipped eventMake visibility follow allocation order.
Replay not idempotentDuplicates on overlapCarry a version per key and skip anything not newer.
Client tests only the final stateA bug that survives because later events overwrite lost onesCount sequence gaps in the client.
Unbounded logMemory or disk growthRetain N events or T minutes; make the retention part of the contract.
Reconnect storm after an outageAll clients ask for snapshots at onceJittered reconnect delay. The retry field sets one fixed delay for every client, so add jitter in your own client.
Cursor treated as meaningful by clientsCannot change the encodingKeep it opaque.
id field containing NULThe browser ignores that id field, so the cursor does not advanceUse plain integers or URL-safe tokens. The spec describes the value as any UTF-8 string without NUL, LF or CR.

When to use this, and when not

Use cursor + log + snapshot when many clients watch mutable state, reconnects are frequent, and the cost of re-reading the whole state each time is noticeable.

Do not use it when:

  • The history is small. Version 0 has no failure modes worth the name.
  • You only need the latest value per key. Refetching on a hint is simpler, and it has one code path.
  • Every event is legally meaningful. A snapshot fallback that hides a gap is the wrong tool; keep a durable log and make “too old” an error the client must handle.

Try it, then break it

# Node.js 18+ (I ran Node.js 20). No dependencies except for eventsource_check.mjs, which needs Chrome or Chromium.
node resume.mjs             # random kills + an outage longer than the retention
node snapshot_order.mjs     # why the order of the two reads matters
node visibility_gap.mjs     # why sequence numbers must become visible in order
node eventsource_check.mjs  # what a real browser does (set CHROME=/path/to/chrome if needed)

Then break it: change RETAIN to 5, or delete the cursor + 1 < oldest clause and add an assertion that the client’s key set equals the server’s after every kill, not just at the end.

Three commitments

Resumability comes down to three commitments, and before you ship a resume path you should be able to answer all three:

  1. Position. What does the client carry, and is it opaque? SSE standardizes this part.
  2. Retention. How long does the log honor that position? That is a number you choose and write down.
  3. Reset. What does the server do when the position is too old, and can the client not miss it? Most designs leave this one out, and it is the difference between a stream that recovers and one that silently drifts.

Where this was checked, and where it was not

Everything ran on Linux 6.12 with Node.js 20.19 and headless Google Chrome 154. Browser behaviour (Last-Event-ID on reconnect, no reconnect after a non-200) was observed in that one browser and matches the HTML Standard as of 2026-10-04. The server is a single process with an in-memory log; snapshot_order.mjs and visibility_gap.mjs are models, not database experiments. I did not test multi-node servers, a persistent log, proxies that buffer event streams, or other browsers.