fix(server): set TCP_NODELAY on every accepted socket - #377
Merged
Merged
Conversation
The server never set TCP_NODELAY on the sockets it accepts. With Nagle's algorithm on, the kernel holds a small write while an earlier small write on the connection is unacknowledged, and the client's delayed ACK leaves with its next request on that connection or when its timer fires (about 40 ms on Linux). Once one response is written while the previous one is unacknowledged, each response waits for the client's next request, and the reply to that request is held behind it in turn. The response latency then equals the gap between the client's requests on the connection, for every gap under the delayed-ACK timer. Measured for PutItem and UpdateItem over HTTP/2 on TLS with 32 connections: p50 = p99 = 32 ms at 1000 rps, 16 ms at 2000, 8 ms at 4000, for a service time under 2 ms; at 500 rps, a 64 ms gap, p50 is 1.7 ms. The same server stack reproduces the signature on loopback against a paced client, and TCP_NODELAY on the accepted socket removes it. Set the option on the raw TCP socket in both serve paths: in the acceptor that runs before the TLS handshake (axum-server) and, on the plaintext path, through axum's ListenerExt::tap_io. A failure to set it is logged at debug and the connection is served with the kernel default. Two tests bind a loopback listener, connect, run the real acceptor or listener on the accepted socket, and read TCP_NODELAY back from it. Both fail without the change. Signed-off-by: Anandh Somasundaram <yesyayen@gmail.com>
yesyayen
marked this pull request as ready for review
September 30, 2026 17:17
yesyayen
requested review from
LeeroyHannigan,
amrith,
c33howard,
jcshepherd,
pdf-amzn and
robinnsc
as code owners
September 30, 2026 17:18
github-merge-queue
Bot
removed this pull request from the merge queue due to no response for status checks
Sep 30, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to no response for status checks
Sep 30, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to no response for status checks
Sep 30, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 1, 2026
yesyayen
enabled auto-merge
October 1, 2026 17:16
github-merge-queue
Bot
removed this pull request from the merge queue due to no response for status checks
Oct 1, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to no response for status checks
Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The server now sets
TCP_NODELAYon every TCP socket it accepts, in both serve paths ofcrates/server/src/lib.rs:axum_server::from_tcp_rustls):HttpsRedirectAcceptorsets it first, on the raw socket before the TLS handshake. axum-server'sNoDelayAcceptorcannot be stacked under this acceptor, so the acceptor sets the option itself.dev-mode,axum::serve): the listener is wrapped with axum'sListenerExt::tap_io.If the call fails, the server logs at
debugand serves the connection with the kernel default. No config, wire, or dependency change.Why
Without
TCP_NODELAY, Nagle's algorithm and the client's delayed ACK hold each small response until the client sends its next request on that connection. An HTTP/2 client with concurrent requests on a connection then sees latency equal to its request gap on that connection (up to about 40 ms), not the service time.Measured on 0.1.12 on EC2: ExtendDB and PostgreSQL 16 on one c7g.4xlarge, the load generator on a c7g.8xlarge, HTTP/2 over TLS, 32 connections, fixed request rate, 1 KB items. Both runs use the same settings; only the image differs.
With the change, both write rates are within 8% of PostgreSQL alone on the same host (8,500 and 8,250 rps). Latency at 500 rps and peak throughput under overload do not change. Clients that send one request at a time per connection are not affected: botocore at 1000 rps measured p50 2.79 ms before and 2.82 ms after.
Closes # n/a
Testing done
crates/server/src/lib.rsaccept a loopback connection through the real code path and readTCP_NODELAYback from the accepted socket:tls_acceptor_sets_tcp_nodelay_before_the_handshakeandplaintext_listener_sets_tcp_nodelay_on_accepted_sockets. Both fail without the change.cargo test --workspace,cargo fmt --check, andcargo clippy --all-targets -- -D warningspass.Checklist
cargo test --workspace)cargo fmt --check)cargo clippy -- -W clippy::pedantic)Storagetrait, auth model, on-disk format, or public CLI surface, an RFC has been accepted or is linked below. Otherwise, an ADR captures the decision (link below).ADR / RFC: n/a (a socket option on accepted connections; no wire, trait, config, or on-disk change)
Breaking changes
n/a
By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0 and I agree to the Developer Certificate of Origin (DCO). See CONTRIBUTING.md for details.