TL;DR

  • send() returns as soon as the bytes are queued. In my runs it accepted 64 MiB in roughly 0.15 to 0.25 s from a receiver that was not reading at all, in both Node (ws) and headless Chrome.
  • bufferedAmount counts only what has not yet been handed to the operating system. With the receiver stalled it stayed 0 for the first 2.56 MiB (Node) and 2.75 MiB (Chrome) on a Linux loopback, because the kernel’s send and receive buffers took that much first. bufferedAmount === 0 does not mean the peer has the data.
  • Waiting for bufferedAmount to fall before each send() protects the sender’s memory. It does not protect a consumer that reads from the socket quickly and queues the work in its own memory: in five runs of the lab, such a consumer still held 387 to 393 of 400 messages (about 24 MiB) at its peak. A credit window of 8 messages held the backlog at 8.
  • A message larger than the receiver’s limit does not fail alone. The receiver closes the whole connection with code 1009, and a small message sent right after it is lost too.
  • Reproduce everything with five short scripts below. They need Node 20, the ws package, and (for one of them) Chrome.

Four beliefs, tested one by one

Most WebSocket code I read looks like this:

socket.send(payload);

It is one line, it does not throw, and nothing in it says where the bytes wait if the other side is slow. Here are four things people (including me) tend to assume about that line, and what a small lab says about each.

The blue box is the only place bufferedAmount can see. TCP’s own flow control protects the box before the receiver’s kernel buffer fills up. Nothing protects the orange box unless you build something.

Belief 1: “send() returned, so it was sent”

The lab: a server that completes the WebSocket handshake and then stops reading from the socket, and a client that sends 1024 messages of 64 KiB (64 MiB) in one synchronous loop.

// A WebSocket server that accepts connections and then stops reading from the socket.
import { WebSocketServer } from "ws";
const wss = new WebSocketServer({ port: 8081 });
wss.on("connection", (ws) => {
  ws._socket.pause();          // stop reading: the kernel receive buffer fills, then the window closes
  console.log("client connected; not reading");
});
console.log("listening on ws://127.0.0.1:8081");
// Send 64 KiB messages as fast as possible to a server that is not reading.
import WebSocket from "ws";
const ws = new WebSocket("ws://127.0.0.1:8081");
const CHUNK = Buffer.alloc(64 * 1024, 1);
ws.on("open", () => {
  let sent = 0, firstBuffered = null;
  const t0 = process.hrtime.bigint();
  for (let i = 1; i <= 1024; i++) {            // 1024 x 64 KiB = 64 MiB in one synchronous loop
    ws.send(CHUNK);
    sent += CHUNK.length;
    if (firstBuffered === null && ws.bufferedAmount > 0) firstBuffered = sent;
  }
  const ms = Number(process.hrtime.bigint() - t0) / 1e6;
  console.log(`send() x1024 returned after ${ms.toFixed(1)} ms`);
  console.log(`bufferedAmount stayed 0 until ${(firstBuffered / 1048576).toFixed(2)} MiB had been "sent"`);
  console.log(`bufferedAmount now: ${(ws.bufferedAmount / 1048576).toFixed(2)} MiB of ${(sent / 1048576)} MiB`);
  console.log(`rss: ${(process.memoryUsage().rss / 1048576).toFixed(0)} MiB`);
  setTimeout(() => process.exit(0), 200);
});

Run node stall_server.mjs in one terminal and node probe_node.mjs in another. Three runs on my machine:

send() x1024 returned after 193.4 ms   (the other two runs: 152.3 and 152.8 ms)
bufferedAmount stayed 0 until 2.56 MiB had been "sent"
bufferedAmount now: 61.51 MiB of 64 MiB
rss: 129 MiB

The loop finished in a fraction of a second, with no exception and no pause, and the receiver had read nothing. About 61.5 MiB was sitting in the sender’s process. The same shape in a real browser, using headless Chrome against the same stalled server:

// Same probe from a real browser WebSocket (headless Chrome). Needs: npm i puppeteer-core, google-chrome.
const puppeteer = require("puppeteer-core");
(async () => {
  const b = await puppeteer.launch({ executablePath: "/usr/bin/google-chrome", headless: "new", args: ["--no-sandbox"] });
  const p = await b.newPage();
  const http = require("http"); const srv = http.createServer((q,r)=>r.end("<!doctype html><title>x</title>")).listen(8082);
  await p.goto("http://127.0.0.1:8082/");
  const r = await p.evaluate(async () => {
    const ws = new WebSocket("ws://127.0.0.1:8081");
    await new Promise((res) => (ws.onopen = res));
    const chunk = new Uint8Array(64 * 1024);
    const t0 = performance.now();
    for (let i = 0; i < 1024; i++) ws.send(chunk);
    const sameTask = ws.bufferedAmount;
    const ms = performance.now() - t0;
    const samples = [];
    for (let k = 0; k < 5; k++) { await new Promise((r) => setTimeout(r, 400)); samples.push(ws.bufferedAmount); }
    return { ms, sameTask, samples, wss: typeof WebSocketStream, readyState: ws.readyState };
  });
  console.log(JSON.stringify(r));
  console.log("sameTask MiB", (r.sameTask / 1048576).toFixed(2), "later MiB", r.samples.map((x) => (x / 1048576).toFixed(2)).join(" "));
  await b.close(); srv.close();
})();
sameTask MiB 64.00 later MiB 61.25 61.25 61.25 61.25 61.25

Right after the loop, in the same task, Chrome reports all 64 MiB. That matches the WebSockets Standard: the getter returns the bytes queued with send() that “as of the last time the event loop reached step 1” had not been transmitted, and that includes everything sent during the current task. Later it settles at 61.25 MiB.

Three runs printed identical byte counts. The 1024 send() calls took roughly 0.15 to 0.25 s in the runs where I printed the timing (the Node figures are above; a re-run on another day gave 190 to 240 ms). The byte counts depend on the socket buffers, the timings on the machine.

The standard also says what happens at the limit: if the data cannot be sent “because it would need to be buffered but the buffer is full”, the browser flags the WebSocket as full and closes the connection (WHATWG WebSockets, send()). It does not say how large that buffer is, and I did not reach it with 64 MiB.

So: send() is “enqueue”. Whether the peer ever receives the data is a separate question that send() cannot answer.

Belief 2: “bufferedAmount is 0, so the peer has it”

The same runs answer this one. bufferedAmount stayed 0 until 2.56 MiB had been sent (Node, ws 8.22.0) and the equivalent figure in Chrome was 64 − 61.25 = 2.75 MiB. The receiver read none of it. The data was in the kernel: on this machine tcp_wmem is 4096 16384 4194304 and tcp_rmem is 4096 131072 6291456 (minimum, default, maximum, in bytes, autotuned), so a few MiB fit in the two socket buffers on a loopback connection. Your numbers will differ.

The specification is explicit that this is by design: bufferedAmount “does not include framing overhead incurred by the protocol, or buffering done by the operating system or network hardware” (WHATWG WebSockets). MDN adds that it does not reset to zero when the connection is closed: if you keep calling send(), it keeps climbing. I did not test that in a browser.

Two small differences for the Node ws library, from its documentation: the value is 0 if the data was sent immediately, and it includes framing bytes, where the standard says it does not.

Use bufferedAmount === 0 as “my own queue is empty”. Never as “delivered”. For delivery you need an acknowledgement from the other application.

Belief 3: “if I wait for bufferedAmount to drain, a slow receiver can’t hurt me”

This is the one that surprised me. The standard’s own example waits for bufferedAmount == 0 before sending the next update, and the MDN WebSocketStream page describes the API as one that “can take advantage of stream backpressure automatically”. It is natural to conclude that gating on bufferedAmount gives you backpressure. It gives you backpressure from the network and the receiver’s kernel. What if the receiver’s kernel is fine, and it is the receiver’s application that is slow?

The lab: a receiver that reads every message immediately, pushes it onto a queue, and works through the queue at 5 ms per message. A sender that sends 400 messages of 64 KiB, in three ways. The script prints how many messages sat in the receiver’s queue at the worst moment.

// Does gating on bufferedAmount protect a slow consumer? Run: node slow_lab.mjs
import { WebSocketServer, WebSocket } from "ws";

const MSGS = 400, SIZE = 64 * 1024, WORK_MS = 5;           // the consumer needs 5 ms per message
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));

async function run(mode) {
  const wss = new WebSocketServer({ port: 0 });
  await new Promise((r) => wss.on("listening", r));
  const port = wss.address().port;
  let peak = 0, done;                                         // peak = most messages ever waiting in the consumer's JS memory
  const finished = new Promise((r) => (done = r));
  wss.on("connection", (ws) => {
    const queue = []; let handled = 0, working = false;
    const pump = async () => {
      if (working) return; working = true;
      while (queue.length) {
        await sleep(WORK_MS); queue.shift(); handled++;
        if (mode === "credit") ws.send("ack");                // credit: one ack per message that is really finished
        if (handled === MSGS) done();
      }
      working = false;
    };
    ws.on("message", () => { queue.push(1); peak = Math.max(peak, queue.length); pump(); });
  });

  const c = new WebSocket(`ws://127.0.0.1:${port}`);
  await new Promise((r) => c.on("open", r));
  const payload = Buffer.alloc(SIZE, 7);
  let inFlight = 0, wake = null, peakBuffered = 0;
  c.on("message", () => { inFlight--; wake?.(); });
  const t0 = Date.now();
  for (let i = 0; i < MSGS; i++) {
    if (mode === "bufferedAmount") while (c.bufferedAmount > 256 * 1024) await sleep(5);
    if (mode === "credit") while (inFlight >= 8) await new Promise((r) => (wake = r));   // window of 8 messages
    c.send(payload); inFlight++;
    peakBuffered = Math.max(peakBuffered, c.bufferedAmount);
  }
  const sendMs = Date.now() - t0;
  await finished;
  console.log(`${mode.padEnd(15)} sender done sending after ${String(sendMs).padStart(5)} ms | sender peak bufferedAmount ${(peakBuffered/1048576).toFixed(2)} MiB | consumer peak backlog ${String(peak).padStart(3)} messages (${(peak*SIZE/1048576).toFixed(1)} MiB)`);
  c.close(); wss.close();
}
for (const m of ["naive", "bufferedAmount", "credit"]) await run(m);

Five runs of node slow_lab.mjs (Node 20.19.2, ws 8.22.0, loopback):

SenderTime to hand over all 400 messagesSender’s peak bufferedAmountReceiver’s peak backlog
naive: send in a loop62 to 78 ms22.57 MiB388 to 390 messages (24.3 to 24.4 MiB)
bufferedAmount: wait until ≤ 256 KiB108 to 216 ms0.25 MiB387 to 393 messages (24.2 to 24.6 MiB)
credit: at most 8 unanswered messagesabout 2050 ms0 MiB8 messages (0.5 MiB)

(Each cell is the range over five runs, except that the naive sender’s peak bufferedAmount was 22.57 MiB every time.)

The middle row is the lesson. The sender’s own queue is perfectly bounded: 0.25 MiB. But the data did not disappear. The receiver’s code read from the socket as fast as it arrived, so TCP’s window never closed, so bufferedAmount never rose, so the sender never waited, and the whole 24 MiB ended up in a JavaScript array on the other side. The problem moved; it did not go away.

The third row is a credit window: the receiver sends back one small "ack" for every message it has finished processing, and the sender never has more than 8 unacknowledged. The backlog is capped by construction, and it costs about two seconds because that is how long the receiver really needs (400 × 5 ms). That is the point: the sender is now paced by the consumer.

bufferedAmount is still useful. It is the right tool to cap the sender’s own memory when the network is slow. The two compose:

  1. credit window against the receiver’s application,
  2. bufferedAmount ceiling against the network and the receiver’s kernel.

Belief 4: “message size is only a performance knob”

The RFC allows a WebSocket message to be fragmented across frames, so it feels as if any size is fine (RFC 6455 §5.4). Receivers still choose a maximum. The same RFC defines close code 1009: “an endpoint is terminating the connection because it has received a message that is too big for it to process” (§7.4.1). Platforms publish such limits; for example, Cloudflare’s Durable Objects documentation lists 32 MiB for a received WebSocket message (limits). The lab uses ws with maxPayload set to 1 MiB.

// What happens to a message that is larger than the receiver allows? Run: node limit_lab.mjs
import { WebSocketServer, WebSocket } from "ws";
const wss = new WebSocketServer({ port: 0, maxPayload: 1024 * 1024 });   // receiver accepts at most 1 MiB per message
await new Promise((r) => wss.on("listening", r));
wss.on("connection", (ws) => {
  ws.on("message", (m) => console.log(`server got a message of ${m.length} bytes`));
  ws.on("error", (e) => console.log(`server error: ${e.message}`));
});
const c = new WebSocket(`ws://127.0.0.1:${wss.address().port}`);
await new Promise((r) => c.on("open", r));
c.send(Buffer.alloc(512 * 1024));                       // fits
c.send(Buffer.alloc(2 * 1024 * 1024));                  // does not fit
c.send(Buffer.alloc(16));                               // a perfectly small message, sent afterwards
await new Promise((r) => c.on("close", (code, reason) => { console.log(`client: closed with code ${code} ${reason}`); r(); }));
wss.close();
server got a message of 524288 bytes
server error: Max payload size exceeded
client: closed with code 1009 

The 512 KiB message arrived. The 2 MiB message closed the connection, and the 16-byte message sent right after it never arrived: the server never printed a second “got a message” line. One message that is too big has a blast radius of the whole connection, including innocent messages in flight behind it.

The practical rule is to pick a maximum message size lower than the smallest limit among all receivers you talk to, split anything bigger into chunks yourself, and send the chunks through the same gate as every other message (credit first, bufferedAmount second). Chunking gives you bounded memory and a clean place to wait. It does not give you retransmission; TCP already does that, and a lost chunk means a lost connection.

What about WebSocketStream?

MDN describes WebSocketStream as a promise-based API on top of streams that “can take advantage of stream backpressure automatically”. It is marked experimental and non-standard, and “not currently a part of any specification” (MDN). In the headless Chrome 154 I used, typeof WebSocketStream was "function". I did not test it, and I did not check other browsers or React Native. Note that stream backpressure on the writing side still only reflects the transport. The credit window is about the other application, and no stream API gives you that for free.

Which one do I need?

SituationWhat to use
Small messages, low rate, receiver does trivial workNothing. send() is fine.
A burst that could exceed your own memory (large payloads, fast producer) on a slow networkGate on bufferedAmount (poll, or use a callback such as send(data, cb) in ws), and stop when readyState is no longer OPEN
Receiver’s processing can be slower than the sender’s productionCredit window with application acknowledgements, sized in bytes if message sizes vary
Payloads that can exceed any receiver’s limitCap message size below the smallest limit, chunk, send chunks through the gates above
Strictly need to know the peer processed somethingApplication-level ack for that message. Nothing below the application can tell you.

Mistakes I would check for first

  • A while (bufferedAmount > high) loop with no exit. Symptom: a tab or process spinning after the connection drops. Fix: also check readyState and reject when it is not OPEN (and remember that MDN says the value does not reset to zero on close).
  • Credit never returned on an error path. Symptom: the sender stalls forever after one failed message. Fix: send the acknowledgement in a finally, or send a negative acknowledgement.
  • Credit counted in messages when sizes vary. Symptom: a window of 8 is fine until one message is 30 MiB. Fix: count bytes.
  • Credit counter not reset on reconnect. Symptom: after reconnecting, the sender believes 8 messages are in flight that will never be acknowledged and sends nothing. Fix: reset the window when a new connection opens.
  • Polling at a fixed interval. Symptom: throughput capped by the sleep, or wasted wakeups. A 5 ms sleep was fine here; measure yours.
  • Treating the first oversized-message close as a transient error. Symptom: reconnect, resend the same too-big message, close again, forever. Fix: treat 1009 as a permanent failure for that payload.

Try it

  1. mkdir lab && cd lab && npm init -y && npm i ws@8 and save the scripts above next to package.json.
  2. Terminal 1: node stall_server.mjs. Terminal 2: node probe_node.mjs. You should see bufferedAmount stay 0 for a couple of MiB and then climb to roughly the total you sent.
  3. node slow_lab.mjs. Success looks like the table above: a receiver backlog near 400 for the first two senders and exactly 8 for credit.
  4. node limit_lab.mjs. Success is a 1009 close and no “got a message of 16 bytes” line.
  5. For the browser version, install puppeteer-core, point executablePath at your Chrome, and run node stall_server.mjs and node probe_chrome.cjs.
  6. Change WORK_MS or the window size (8) in slow_lab.mjs and predict the new backlog and duration before running.

How far these measurements go

Verified, on one machine (Linux 6.12, loopback, Node 20.19.2, ws 8.22.0, headless Chrome 154): the numbers in the tables above, which come from the scripts as printed (five runs of slow_lab.mjs, three runs each of the probes, one run of limit_lab.mjs). The quotes from the WebSockets Standard, RFC 6455, MDN, the ws documentation and the Cloudflare limits page were read at the original pages while writing this.

Not verified: any real network (the kernel absorption figures of 2.56 and 2.75 MiB depend on socket buffer sizes and will differ across systems), other browsers, React Native, Safari, what a browser does when its send buffer is actually full, bufferedAmount after close, WebSocketStream behaviour, permessage-deflate (RFC 7692), and Cloudflare’s runtime itself (I only cite its documented limit). The 5 ms per message and the window of 8 are lab choices, not recommendations.

Make the other side say how much it can take

send() means “queued”. bufferedAmount means “not yet in the kernel”. Neither means “the other side is keeping up”. If the other side can be slower than you, make it say so, with a number it controls: credits.