TCP for Application Developers: Handshake, Sliding Window, CUBIC, Nagle and Keepalive Tuning
Key takeaways
TCP connections, sliding windows, Reno/CUBIC congestion control, Nagle, TCP_NODELAY, keepalive—reliable transport and production socket tuning.
Why TCP behavior still matters
TCP (Transmission Control Protocol) is the default reliable transport on the internet—HTTP/HTTPS, SSH, most database protocols, and many microservice RPCs run over it. Connection state, retransmissions, and ordering “just work” because of kernel TCP and the socket API.
Yet latency, throughput, and concurrent connections per server depend heavily on settings and traffic patterns. This article maps RFCs and kernel behavior to socket options and common incidents: the handshake and teardown (and why TIME_WAIT and CLOSE_WAIT show up in ss output), the two separate windows that limit how fast data flows, the Nagle/delayed-ACK interaction behind mysterious 40 ms latencies, and what keepalive can and cannot detect.
TCP from Cerf and Kahn to modern congestion control
RFC 793 and later extensions
TCP evolved from 1974 work by Vint Cerf and Bob Kahn; RFC 793 (1981) was the classic reference. Later additions include SACK, window scaling, timestamps, and congestion-control updates, consolidated in RFC 9293 (2022). Implementations differ: BSD Reno, Linux default CUBIC, etc.
Layer 4: a reliable byte stream between ports
TCP is layer 4 (transport). IP (layer 3) delivers packets between hosts; TCP provides a reliable byte stream between ports (processes). Segmentation, retransmission, and ordering are TCP’s job.
Connection state, ordering, flow and congestion control
| Property | Description |
|---|---|
| Connection-oriented | Logical connection and state machine before bulk data. |
| Reliable | Retransmits on loss; sequence numbers fix reordering. |
| Ordered delivery | Application reads bytes in send order (no message boundaries). |
| Flow control | rwnd matches sender rate to receiver buffer. |
| Congestion control | Reacts to network congestion—cwnd and algorithm. |
| Full duplex | Both directions on one connection (implemented with buffers and ACK interplay). |
The “no message boundaries” row causes more application bugs than any other TCP property. TCP delivers a stream of bytes, not the messages you wrote: two send calls of 100 bytes can arrive as one recv of 200 bytes, or as reads of 37 and 163 bytes. Code that assumes one recv equals one message works on localhost, where writes tend to arrive intact, and fails intermittently over real networks. Every protocol on top of TCP therefore defines its own framing: a length prefix, a delimiter such as \r\n, or a self-describing format like HTTP’s headers plus Content-Length.
Handshake, teardown, windows, and congestion control
3-way handshake
SYN → SYN-ACK → ACK. The client sends an ISN, the server responds with its ISN and ACK, the client ACKs—then ESTABLISHED.
sequenceDiagram participant C as Client participant S as Server C->>S: SYN, seq=x S->>C: SYN-ACK, seq=y, ack=x+1 C->>S: ACK, seq=x+1, ack=y+1 Note over C,S: ESTABLISHED
The handshake costs one full round trip before any application data moves, which is why connection reuse matters so much for latency: on a 100 ms RTT path, a fresh TCP connection plus a TLS 1.3 handshake spends about 200 ms before the first request byte. Initial sequence numbers are randomized so an off-path attacker can’t guess them and inject data into the connection.
On the server, a connection that has completed the handshake waits in the listen backlog until the application calls accept. If the application accepts too slowly, the queue fills (bounded by the listen() backlog argument and net.core.somaxconn on Linux), and new SYNs are dropped. Clients then see connection attempts that take one or three seconds longer than normal, because they are waiting for SYN retransmission timers. nstat or netstat -s counters for listen queue overflows are the place to confirm it. A flood of SYNs that never complete (a SYN flood) is handled by SYN cookies, which let the server avoid keeping state until the final ACK arrives.
4-way termination
Each direction closes with FIN/ACK. The side that sends the last ACK enters TIME_WAIT briefly so late duplicates do not collide with a new connection on the same quad.
sequenceDiagram participant A as Host A participant B as Host B A->>B: FIN B->>A: ACK B->>A: FIN A->>B: ACK Note over A: TIME_WAIT (2MSL)
TIME_WAIT lands on whichever side closes first (the active closer), not necessarily the client. It lasts twice the maximum segment lifetime; Linux hardcodes 60 seconds. During that time the four-tuple (source IP and port, destination IP and port) can’t be reused for a new connection by default, so a client that opens and closes thousands of short connections per second to the same server can run out of ephemeral ports, and connect fails with EADDRNOTAVAIL (“Cannot assign requested address”).
The mirror-image state, CLOSE_WAIT, is almost always an application bug. It means the peer has sent FIN and the kernel has ACKed it, but your process has not called close() on its socket. The kernel can’t finish the connection on its own, so CLOSE_WAIT sockets accumulate until the process leaks file descriptors. When ss -tan state close-wait shows a growing number, look for a code path that stops reading after an error or on EOF (a recv returning 0) and never closes the socket.
Flow control (sliding window)
rwnd advertises how much more data the receiver can buffer. The sender limits unacknowledged bytes to stay within rwnd—protects the receiver. If the receiving application stops reading, its buffer fills, the advertised window drops to zero, and the sender stops. That is TCP’s built-in backpressure: a slow consumer eventually makes the producer’s send block (or return EAGAIN on a non-blocking socket). A zero window in a packet capture points at the receiving application, not the network.
Congestion control
Congestion control reacts to router queues and drops to protect the whole network. Reno includes slow start, congestion avoidance, fast retransmit/recovery. Linux CUBIC adjusts window growth on high-BDP links. Exact parameters depend on kernel version and sysctl.
The sender may have at most min(rwnd, cwnd) bytes in flight. cwnd starts small (Linux uses an initial window of 10 segments, about 14 KB) and roughly doubles each RTT during slow start until loss occurs. That is why a new connection is slow for its first few round trips even on a fast link, and why a 50 KB response on a fresh connection takes several RTTs while the same response on a warm, reused connection takes one. Reno and CUBIC both treat packet loss as the congestion signal and cut cwnd sharply when it happens; on links with random, non-congestion loss (some wireless networks), that makes throughput collapse. BBR, available in Linux, instead models bandwidth and RTT directly and is less sensitive to such loss, which is why some content providers use it on their servers.
flowchart TB
subgraph fc [Flow control]
R[Receive buffer rwnd]
end
subgraph cc [Congestion control]
CWND[cwnd / algorithm]
end
SEND[Data actually sent]
R --> SEND
CWND --> SEND
TCP sockets in C++, Python, and Node.js
Educational minimal examples—production needs async I/O, TLS, pools, and logging.
C++ (Berkeley sockets)
// g++ -std=c++17 -O2 tcp_client.cpp -o tcp_client
#include <arpa/inet.h>
#include <cstring>
#include <iostream>
#include <string>
#include <sys/socket.h>
#include <unistd.h>
int main(int argc, char* argv[]) {
const char* host = argc > 1 ? argv[1] : "127.0.0.1";
const uint16_t port = argc > 2 ? static_cast<uint16_t>(std::stoi(argv[2])) : 8080;
int fd = ::socket(AF_INET, SOCK_STREAM, 0);
if (fd < 0) { perror("socket"); return 1; }
sockaddr_in addr{};
addr.sin_family = AF_INET;
addr.sin_port = htons(port);
if (inet_pton(AF_INET, host, &addr.sin_addr) != 1) {
std::cerr << "inet_pton failed\n";
return 1;
}
if (connect(fd, reinterpret_cast<sockaddr*>(&addr), sizeof(addr)) < 0) {
perror("connect");
close(fd);
return 1;
}
const std::string msg = "ping\n";
ssize_t n = send(fd, msg.data(), msg.size(), 0);
if (n < 0) { perror("send"); close(fd); return 1; }
char buf[4096];
n = recv(fd, buf, sizeof(buf) - 1, 0);
if (n < 0) { perror("recv"); close(fd); return 1; }
if (n == 0) std::cout << "peer closed\n";
else { buf[n] = '\0'; std::cout << buf; }
close(fd);
return 0;
}
- Check every
socket/connect/send/recvreturn. send/recvmay be partial—loop for full buffers in real code.
The single recv in this example is the framing mistake described earlier: it happens to work because the reply is short and the server sends it in one write. A real client keeps calling recv and appending to a buffer until it has a complete message by the protocol’s framing rule. recv returning 0 means the peer closed its sending side (an orderly FIN), which is different from an error. Two error cases are worth recognizing: ECONNRESET (“Connection reset by peer”) means the other side sent RST, and writing to a connection the peer has already reset raises SIGPIPE, which by default kills the process silently. Servers either ignore SIGPIPE or pass MSG_NOSIGNAL to send and handle the EPIPE error instead.
Python 3
#!/usr/bin/env python3
import socket
import sys
def main() -> None:
host = sys.argv[1] if len(sys.argv) > 1 else "127.0.0.1"
port = int(sys.argv[2]) if len(sys.argv) > 2 else 8080
with socket.create_connection((host, port), timeout=10.0) as sock:
sock.sendall(b"ping\n")
data = sock.recv(4096)
if not data:
print("peer closed")
else:
print(data.decode("utf-8", errors="replace"))
if __name__ == "__main__":
main()
socket.create_connection resolves the host name, tries each returned address (IPv6 and IPv4) in turn, and applies the timeout to the connect as well as to later operations. sendall loops internally until every byte is written, which plain send does not. recv(4096) still returns whatever has arrived, up to 4096 bytes, so the framing caveat applies here too.
JavaScript (Node.js)
// node tcp_client.mjs
import net from "node:net";
const host = process.argv[2] ?? "127.0.0.1";
const port = Number(process.argv[3] ?? 8080);
const socket = net.createConnection({ host, port });
socket.setTimeout(10_000);
socket.on("timeout", () => socket.destroy(new Error("idle timeout")));
socket.on("connect", () => {
socket.write("ping\n");
});
socket.on("data", (chunk) => {
process.stdout.write(chunk);
});
socket.on("error", (err) => {
console.error(err.message);
process.exitCode = 1;
});
In Node.js, each data event delivers whatever chunk the kernel had ready, so a message can be split across events or several messages can arrive in one. socket.write never blocks: if the kernel buffer is full, Node queues the data in user space and write returns false. Ignoring that return value and writing in a tight loop is how a Node service’s memory grows without bound when a client reads slowly; wait for the drain event before writing more, or use stream.pipeline, which handles backpressure for you. setTimeout here is an idle timeout that only emits an event; the handler must destroy the socket itself.
Timeouts and errors
- Connect timeout: limit long connects with OS options or async wrappers.
- Read timeout:
SO_RCVTIMEO,socket.setTimeout,sock.settimeout, … to avoid infinite recv. - Retries: safest for idempotent application requests only.
Latency, throughput, and header overhead
Latency
RTT caps throughput on many workloads. Request–response with small messages often hits RTT per exchange—sensitive to packet count, ACK timing, and Nagle.
Throughput
If window sizes (receive + congestion) fall short of BDP (bandwidth × delay), you underfill the pipe. Window scaling, buffer tuning, and application I/O matter on fast long-haul links.
A concrete example: a 1 Gbps link with 80 ms RTT has a BDP of 1 Gbit/s × 0.08 s = 80 Mbit, or 10 MB. To keep that link full, 10 MB must be in flight, so both the receiver’s buffer and the sender’s congestion window must reach that size. The original TCP header’s window field maxes out at 64 KB, which would cap this connection at about 6.5 Mbit/s; the window scaling option negotiated in the handshake is what lifts that limit. Modern Linux autotunes socket buffers up to the limits in net.ipv4.tcp_rmem and tcp_wmem, so long-haul throughput problems are usually solved by raising those maximums rather than by application changes.
Overhead
Each segment adds IP + TCP headers (20+ bytes + options). TLS adds handshakes and crypto. Small payloads waste header ratio.
Benchmarks
Loopback can reach Gbps; cross-region may sit at Mbps to tens of Mbps due to RTT and loss. Measure with iperf3 in target environments before committing to architecture.
Web, databases, and file transfer over TCP
Web and APIs
HTTP/1.1 and HTTP/2 ride TCP (HTTP/3 uses QUIC over UDP). Reverse proxies and load balancers tune reuse, timeouts, and buffers—directly impacting latency.
Databases
PostgreSQL, MySQL, … default to TCP. Connection pools amortize TCP + auth cost.
File transfer
FTP control is TCP, SFTP, rsync over SSH—order and reliability matter. Large transfers: profile disk vs network vs window limits.
Nagle, TCP_NODELAY, keepalive, and buffer sizes
Nagle’s algorithm
Nagle batches small writes to reduce packet count but can add delay. Interactive workloads often set TCP_NODELAY.
The rule is: while any sent data is unacknowledged, hold back further small segments until an ACK arrives or a full segment’s worth of data accumulates. On its own that costs at most one RTT. The real problem is its interaction with delayed ACKs on the receiver, which wait briefly (up to about 40 ms on Linux, 200 ms on Windows by default) hoping to piggyback the ACK on a response. Consider a client that sends a request as two writes, a header and then a body: the first segment goes out, Nagle holds the second until the first is ACKed, and the server delays that ACK because it is waiting for the complete request before responding. Both sides wait for the timer. The symptom is request latency that sits suspiciously close to 40 ms (or 200 ms) regardless of payload size.
The first time I ran into this, a small internal RPC measured about 40 ms per call on a LAN with sub-millisecond RTT; one setsockopt(TCP_NODELAY) brought it under a millisecond. The better fix, though, is to write each message with a single send (or writev for several buffers), which avoids the stall and doesn’t give up Nagle’s batching elsewhere.
TCP_NODELAY
Disables Nagle—small writes go out immediately at the cost of more packets and sometimes lower goodput. Many RPC frameworks and database drivers already set it for you.
SO_KEEPALIVE
Probes idle connections. Often overlaps with app heartbeats—coordinate with proxy and firewall idle timeouts.
On Linux the defaults are 7200 s idle before the first probe, then probes every 75 s, and the connection is declared dead after 9 unanswered probes, so over two hours pass before a dead peer is detected. The per-socket options TCP_KEEPIDLE, TCP_KEEPINTVL and TCP_KEEPCNT override these without changing system-wide sysctls. Keepalive only detects peers that are unreachable while the connection is idle; it does nothing for a peer that is alive but whose application is hung, and it doesn’t cover the case where data is waiting to be sent (that is governed by the retransmission timeout, or TCP_USER_TIMEOUT). Application-level heartbeats detect both, which is why protocols such as WebSocket, gRPC and many database protocols have their own.
Buffers and windows
SO_SNDBUF / SO_RCVBUF and global sysctl limits trade throughput vs memory. Change after measurement. On Linux, setting SO_RCVBUF or SO_SNDBUF explicitly also turns off autotuning for that socket, so a value copied from an old tuning guide can make a connection slower than the default. Also, the kernel doubles the value you set to account for bookkeeping overhead, so getsockopt returns twice what you passed.
TIME_WAIT pileups, resets, and slow transfers
Too many TIME_WAITs
Bursty short connections pressure local ports and kernel tables. Mitigate with HTTP keep-alive, pools, and careful server-side tuning (e.g. reuse options—verify kernel and security guidance).
Thousands of TIME_WAIT sockets on a server are usually harmless: each costs a small amount of kernel memory and the server’s listening port isn’t consumed. The problem case is a client (including a proxy or service calling another service) making many short connections to one destination, which exhausts its ephemeral port range (net.ipv4.ip_local_port_range, about 28,000 ports by default). Connection pooling fixes the cause. Of the kernel options, net.ipv4.tcp_tw_reuse lets outgoing connections reuse TIME_WAIT ports safely using TCP timestamps. The once-popular tcp_tw_recycle broke clients behind NAT and was removed from Linux in 4.12; advice that recommends it is outdated.
Resets and disconnects
RST reasons include closed port, middleboxes, timeouts. Use tcpdump to see who sent FIN/RST. Some resets come from the application itself: closing a socket that still has unread data in its receive buffer makes the kernel send RST instead of FIN, so the peer gets “connection reset” even though the close was deliberate. Servers that reply with an error and immediately close, without draining the rest of the request, trigger this, and the client may never see the error response. Setting SO_LINGER with a zero timeout also turns close into an abortive RST, which some code does intentionally to avoid TIME_WAIT.
Slow transfers
Mix of send buffer stall, slow recv, cwnd limits, disk. Separate iperf on one connection from app profiling. On Linux, ss -ti shows per-connection internals: cwnd, rtt, retransmission counts, and whether the connection is limited by the receiver window, the send buffer, or the application (rwnd_limited, sndbuf_limited, app_limited in recent kernels), which answers “whose fault is it” faster than guessing.
TCP recap and when to pick it
Summary
- TCP provides reliability, ordering, flow control, and congestion control—the transport backbone of the internet.
- Handshake, teardown, and TIME_WAIT tie directly to ops incidents.
- Nagle, windows, buffers, and keepalive are the latency vs throughput levers.
When to choose TCP
- Files, APIs, databases, remote shells when integrity dominates. For low-latency realtime, compare the UDP guide and WebRTC guide.
Frequently Asked Questions (FAQ)
Q. Why do idle TCP connections behind a load balancer or NAT get dropped silently?
A. Middleboxes such as NAT gateways, firewalls and load balancers track connections and discard the state after their own idle timeout, and neither endpoint is notified. The next write then fails or hangs until it times out. OS-level SO_KEEPALIVE often does not help with default settings, because Linux waits two hours (tcp_keepalive_time = 7200 seconds) before the first probe. Either lower the keepalive interval below the middlebox timeout or send application-level heartbeats, and make sure clients reconnect cleanly when a connection turns out to be dead.