TL;DR

  • A client that sends a message as two small write() calls and then waits for a reply can see each round trip take about 40 ms on Linux, with idle CPUs and no packet loss. In my sandbox the median was 44 ms, against 0.02 ms when the same bytes went out in one write.
  • Cause: the sender’s Nagle algorithm holds the second small segment until the first is acknowledged. The receiver’s delayed ACK holds that acknowledgment, hoping to piggyback it on a reply. The reply cannot be sent until the second segment arrives. Each side waits for the other until the delayed-ACK timer fires (40 ms minimum on Linux).
  • Fix, in order of preference: send the message in one write (build it in a buffer, or use writev/sendmsg); otherwise set TCP_NODELAY on the writing side; as a last resort, TCP_QUICKACK on the receiving side (Linux-specific, and it must be re-armed).
  • Everything below is reproducible: two short scripts, one measures, the other prints a packet-level timeline.

The puzzle

Imagine a tiny request protocol. A request is a 4-byte header followed by an 8-byte body. The server waits until it has all 12 bytes, then answers with 2 bytes. The client code is the obvious thing:

sock.sendall(HEADER)   # 4 bytes
sock.sendall(BODY)     # 8 bytes
reply = sock.recv(2)

Both ends are on the loopback interface of one machine. Nothing is lossy, nothing is congested, and the server does no work. How long should one round trip take?

Make a guess before reading on. If you said “tens of microseconds”, you are in good company, and you are wrong by three orders of magnitude on a default Linux socket. The script later in this post gives (Linux 6.12, Python 3.13, one sandbox, medians of 50 round trips; your numbers will differ):

two writes, Nagle on (default)     median    44.00 ms
one write,  Nagle on (default)     median     0.02 ms

The bytes are the same. Only the number of write() calls changed. A second detail makes this bug survive code review and one-shot tests: the first round trip on a fresh connection is fast (about 0.3 ms in my trace below). The stall starts on the second request.

Four ways out, and what each costs

OptionWhat it doesCost and catches
Coalesce into one write (build the message in a buffer, or writev / sendmsg with a list of buffers)The whole request leaves as one segment. Nagle has nothing to hold back.You need to control the code that writes. No socket option, portable. This is the one I would reach for first.
TCP_NODELAY on the writerDisables Nagle: small segments go out immediately.More, smaller packets if you keep writing in tiny pieces: each segment carries at least 40 bytes of IPv4 + TCP headers (without options). Fine for request/response traffic, wasteful for bulk.
TCP_QUICKACK on the reader (Linux)Acknowledge immediately instead of delaying.Linux-only, and not permanent: the kernel can drop back to delayed ACKs, so you re-set it around reads. Useful when you cannot change the writer.
TCP_CORK on the writer (Linux)Hold partial frames until you uncork, then send everything at once.Linux-specific; the man page notes a 200 ms ceiling on how long output is corked. Good for “headers, then sendfile”, clumsy for plain request/response.

Linux also has tcp_autocorking (default on since 3.14), which coalesces small writes when an earlier packet is still waiting in a qdisc or device queue (tcp(7)). That is a different mechanism from Nagle: /proc/sys/net/ipv4/tcp_autocorking was 1 on the sandbox kernel and the stall still occurred, so it does not rescue this case.

Many runtimes already pick an answer for you. Go’s net package enables no-delay by default. libcurl sets TCP_NODELAY by default since 7.50.2. Node.js’s http server defaults noDelay to true since v18, while a plain net.Socket starts with Nagle enabled. The Python socket in my script, as the numbers show, leaves it on. Which default you inherit depends on the library you are standing on.

How the two algorithms work

Nagle: at most one small segment in flight

RFC 896 introduced the rule to stop networks from drowning in one-byte packets. RFC 9293 §3.7.4 states it compactly:

If there is unacknowledged data (i.e., SND.NXT > SND.UNA), then the sending TCP endpoint buffers all user data (regardless of the PSH bit) until the outstanding data has been acknowledged or until the TCP endpoint can send a full-sized segment (Eff.snd.MSS bytes).

So with nothing in flight, a small write goes out immediately. With a small segment in flight, the next small write waits. Implementations “SHOULD” do this, but there “MUST be a way for an application to disable the Nagle algorithm on an individual connection” (RFC 1122 §4.2.3.4, RFC 9293 §3.7.4). That is TCP_NODELAY.

Delayed ACK: wait a little, maybe there is data to ride on

A receiver may postpone an acknowledgment so it can piggyback on outgoing data or cover two segments with one ACK. RFC 1122 §4.2.3.2 says the delay “MUST be less than 0.5 seconds”, and RFC 5681 §4.2 repeats that an ACK must go out within 500 ms of the first unacknowledged packet and at least every second full-sized segment. The 500 ms is a ceiling. Linux uses a much shorter timer: in the kernel’s include/net/tcp.h, TCP_DELACK_MIN is HZ/25 (40 ms) and TCP_DELACK_MAX is HZ/5 (200 ms). The “about 40 ms” in the title is that implementation constant, not a number from the RFCs. ss -i shows the per-connection value as ato:40.

Why together they stall

The standards know about this. RFC 9293 notes that “there can be problematic interactions between the Nagle algorithm and delayed acknowledgments” and its Appendix A.3 describes a modification (specified in an IETF draft) that some operating systems implement. With a default socket the sequence is:

  1. Client writes the 4-byte header. Nothing is in flight, so it is sent at once.
  2. Client writes the 8-byte body. The header is still unacknowledged and the body is smaller than a full segment, so Nagle holds it.
  3. Server receives the header. It has no reply to send yet (it needs 12 bytes), so there is nothing to piggyback an ACK on, and it delays the ACK.
  4. After the delayed-ACK timer (about 40 ms) the ACK goes out. The client’s Nagle logic releases the body. The server now has 12 bytes and replies.

Neither side is misbehaving. Remove either algorithm and the stall disappears.

Looking inside: a measurement and a packet trace

The measurement

This script times the round trip for four variants, plus the receiver-side TCP_QUICKACK variant. The server answers only after receiving a complete 12-byte message.

#!/usr/bin/env python3
"""Write-write-read over TCP: why does the reply take ~40 ms?

The server answers only after it has received a full 12-byte message
(4-byte header + 8-byte body). The client sends the message either as two
small writes or as one. Run it on any Linux box and compare the medians.
"""
import socket, threading, time, statistics, sys

HEADER, BODY = b"HEAD", b"BODYBODY"
N = 50

def server(listener, quickack):
    conn, _ = listener.accept()
    while True:
        got = b""
        while len(got) < len(HEADER) + len(BODY):
            if quickack:  # not permanent (tcp(7)): re-arm before every wait
                conn.setsockopt(socket.IPPROTO_TCP, socket.TCP_QUICKACK, 1)
            chunk = conn.recv(4096)
            if not chunk:
                return
            got += chunk
        conn.sendall(b"ok")

def run(label, *, nodelay, send, quickack=False):
    listener = socket.socket()
    listener.bind(("127.0.0.1", 0))
    listener.listen()
    threading.Thread(target=server, args=(listener, quickack), daemon=True).start()
    c = socket.create_connection(listener.getsockname())
    if nodelay:
        c.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
    samples = []
    for _ in range(N):
        t0 = time.perf_counter()
        send(c)
        c.recv(2)
        samples.append((time.perf_counter() - t0) * 1000)
    c.close()
    print(f"{label:<34} median {statistics.median(samples):8.2f} ms   max {max(samples):8.2f} ms")

two_writes = lambda c: (c.sendall(HEADER), c.sendall(BODY))
one_write  = lambda c: c.sendall(HEADER + BODY)
gather     = lambda c: c.sendmsg([HEADER, BODY])      # writev-style: one syscall, one segment

print(sys.platform, "-", "n =", N, "round trips per row")
run("two writes, Nagle on (default)",  nodelay=False, send=two_writes)
run("two writes, TCP_NODELAY",         nodelay=True,  send=two_writes)
run("one write,  Nagle on (default)",  nodelay=False, send=one_write)
run("sendmsg([hdr, body]), Nagle on",  nodelay=False, send=gather)
run("two writes, receiver TCP_QUICKACK", nodelay=False, send=two_writes, quickack=True)

Output from one run in my sandbox (Linux 6.12, Python 3.13, loopback). Absolute numbers depend on kernel, timer configuration and load; what matters is the gap and where it disappears:

linux - n = 50 round trips per row
two writes, Nagle on (default)     median    44.00 ms   max    48.04 ms
two writes, TCP_NODELAY            median     0.02 ms   max     0.08 ms
one write,  Nagle on (default)     median     0.02 ms   max     0.07 ms
sendmsg([hdr, body]), Nagle on     median     0.02 ms   max     0.05 ms
two writes, receiver TCP_QUICKACK  median     0.02 ms   max     0.07 ms

Three things to read off this table. Setting TCP_NODELAY on the writer removes the stall. Sending the same bytes with one sendall or one sendmsg removes it without touching a socket option. And the receiver-side TCP_QUICKACK removes it from the other end. (In an earlier version of my script I set TCP_NODELAY on the server socket instead; it changed nothing, because the held-back segment is the client’s. The option belongs on the side that writes the small pieces.)

The packet trace

A table of latencies tells you that it stalls. A trace tells you why. Inside a throwaway user and network namespace (unshare -Urn, so no real root is needed), a raw AF_PACKET socket on lo can timestamp every outgoing TCP segment:

#!/usr/bin/env python3
# Run: unshare -Urn python3 nagle_trace.py   (needs CAP_NET_RAW inside a throwaway network namespace)
import socket, struct, subprocess, threading, time, sys
subprocess.run(["ip","link","set","lo","up"],check=True)
sniff = socket.socket(socket.AF_PACKET, socket.SOCK_RAW, socket.htons(3)); sniff.bind(("lo",0))
events=[]; t0=None
def run_sniff():
    while True:
        raw,addr=sniff.recvfrom(65535)
        if addr[2]!=4: continue          # lo shows every packet twice; keep PACKET_OUTGOING
        p=raw[14:]
        if len(p)<40 or p[9]!=6: continue
        ihl=(p[0]&15)*4; sp,dp,seq,ack,off,fl=struct.unpack("!HHIIBB",p[ihl:ihl+14]); plen=len(p)-ihl-(off>>4)*4
        if fl&2: continue
        events.append((time.perf_counter(),sp,dp,fl,plen))
threading.Thread(target=run_sniff,daemon=True).start()
H,B=b"HEAD",b"BODYBODY"
def server(l):
    c,_=l.accept()
    while True:
        got=b""
        while len(got)<12:
            d=c.recv(4096)
            if not d: return
            got+=d
        c.sendall(b"ok")
l=socket.socket(); l.bind(("127.0.0.1",0)); l.listen(); sport=l.getsockname()[1]
threading.Thread(target=server,args=(l,),daemon=True).start()
c=socket.create_connection(l.getsockname())
for i in range(4):                       # a few rounds: the first is in quick-ack mode
    time.sleep(0.2); events.clear()
    t=time.perf_counter(); c.sendall(H); c.sendall(B); c.recv(2); dt=(time.perf_counter()-t)*1000
    time.sleep(0.1)
    print(f"--- round {i+1}: reply after {dt:.1f} ms")
    for (ts,sp,dp,fl,pl) in list(events):
        who = "client->server" if dp==sport else "server->client"
        kind = f"{pl} B data" if pl else "pure ACK"
        print(f"  +{(ts-t)*1000:7.2f} ms  {who}  {kind}")

Four rounds on one connection. This is the output for the first two:

--- round 1: reply after 0.3 ms
  +   0.19 ms  client->server  4 B data
  +   0.29 ms  server->client  pure ACK
  +   0.34 ms  client->server  8 B data
  +   0.34 ms  server->client  pure ACK
  +   0.35 ms  server->client  2 B data
  +   0.36 ms  client->server  pure ACK
--- round 2: reply after 44.1 ms
  +   0.20 ms  client->server  4 B data
  +  43.91 ms  server->client  pure ACK
  +  43.94 ms  client->server  8 B data
  +  43.95 ms  server->client  pure ACK
  +  44.03 ms  server->client  2 B data
  +  44.04 ms  client->server  pure ACK

Round 1 is the quick one: the receiver acknowledges at once, and the body follows 0.15 ms after the header. In round 2 the header is sent at +0.20 ms, nothing happens for about 44 ms, then a pure ACK from the server arrives and only after that does the client send the body. The client’s second write() returned immediately; the kernel simply held the data. Rounds 3 and 4 in the same run look like round 2.

Why is round 1 fast? A receiver that sees a connection’s first data starts in a quick-ACK mode and acknowledges the first segments immediately; the kernel initializes it when the first data packet arrives (tcp_event_data_recv in net/ipv4/tcp_input.c). Once the exchange looks interactive, delayed ACKs take over. I observed the effect and read that code path, but I did not trace which exact kernel transition ends quick-ACK mode, so treat the explanation as “consistent with”, not “proved”.

Which fix, when

  • Request/response protocol with small messages: assemble each message in one buffer and write it once. Add TCP_NODELAY anyway if the code path is latency-sensitive and you cannot rule out a second write elsewhere.
  • A framework or library does the writing, and you cannot change it: check whether it already sets no-delay (see the defaults above). If not, set TCP_NODELAY on the socket it uses. If you control only the reading side, TCP_QUICKACK is a stopgap.
  • Bulk, one-directional transfer of large writes: leave Nagle alone. Full-sized segments are never held back, and there is no small-write pattern to fix.
  • Do not reach for TCP_NODELAY to fix slowness you have not diagnosed. If the latency does not cluster near 40 ms (or up to 200 ms), this is not your problem.

How the stall comes back

  • Many tiny writes with TCP_NODELAY on. The stall is gone, but each write() can become its own segment, 40 or more bytes of headers for a few bytes of payload. Symptom: a high packet count for little data. Fix: put a buffered writer in front of the socket and flush once per message.
  • Another layer writes in two pieces. You write one buffer; a library underneath emits a header, then the body. Symptom: the stall is back despite your coalescing. Fix: check the system-call sequence (for example with strace -e trace=write,sendto,sendmsg; I did not use it here) and configure that layer.
  • Option on the wrong side. TCP_NODELAY helps on the endpoint that writes the held-back data. Setting it on the reader does nothing for this stall (I tried).
  • TCP_QUICKACK set once. The kernel can leave quick-ACK mode again, so a single setsockopt after accept() is not a fix. Re-arm it around every read, and expect it to be Linux-only.
  • Benchmarks that hide it. A single request, or a warm-up that is mostly the first round trip, looks fast. Measure many consecutive round trips on one connection.
  • Mistaking it for server slowness. Symptom: latency stuck near one value with idle CPU. Fix: look at the timeline, not at application profiles.

Reproduce it

  1. Save the first script as nagle_demo.py and run python3 nagle_demo.py on a Linux machine. You should see one row roughly 40 ms higher than the others, and that row should be the two-write one.
  2. Run the second script as unshare -Urn python3 nagle_trace.py (needs unprivileged user namespaces). Find the gap between the 4-byte segment and the 8-byte segment.
  3. Raise N to 500 so the first row runs for about 20 s, and run ss -tin dst 127.0.0.1 in another terminal meanwhile. Look for ato:40 on the two sockets of the demo connection.
  4. Move the server to another host and compare. I did not test that; loopback and real links share the mechanism, but not necessarily the numbers.

What I verified and what I did not

I ran both scripts on one Linux 6.12 sandbox and printed the numbers above; three runs of the measurement gave medians of 44.00, 44.00 and 43.99 ms (the sub-millisecond rows moved between 0.02 and 0.04 ms from run to run) for the two-write row. I read the RFC passages, tcp(7) and the kernel constants linked above. I did not test other kernels, other operating systems, real network links, or TLS.

One message, one write

Nagle holds a small write while earlier data is unacknowledged, delayed ACK holds the acknowledgment, and a request/response exchange built from two small writes makes them wait on each other until a timer fires. On Linux that timer is at least 40 ms. The first request on a connection can be fast, so one-shot tests lie. The robust fix is structural: one message, one write. TCP_NODELAY on the writer is the practical safety net, and TCP_QUICKACK on the reader is the last resort. If you keep one diagnostic from this post, keep this: latency that clusters around a fixed value on an idle machine is a timer, not a workload.