Skip to content

fix(server): set TCP_NODELAY on every accepted socket - #377

Merged
yesyayen merged 3 commits into
ExtendDB:mainfrom
yesyayen:fix/server-tcp-nodelay
Oct 2, 2026
Merged

yesyayen merged 3 commits into
ExtendDB:mainfrom
yesyayen:fix/server-tcp-nodelay

Conversation

@yesyayen

Copy link
Copy Markdown
Member

What

The server now sets TCP_NODELAY on every TCP socket it accepts, in both serve paths of crates/server/src/lib.rs:

  • TLS (axum_server::from_tcp_rustls): HttpsRedirectAcceptor sets it first, on the raw socket before the TLS handshake. axum-server's NoDelayAcceptor cannot be stacked under this acceptor, so the acceptor sets the option itself.
  • Plaintext (dev-mode, axum::serve): the listener is wrapped with axum's ListenerExt::tap_io.

If the call fails, the server logs at debug and serves the connection with the kernel default. No config, wire, or dependency change.

Why

Without TCP_NODELAY, Nagle's algorithm and the client's delayed ACK hold each small response until the client sends its next request on that connection. An HTTP/2 client with concurrent requests on a connection then sees latency equal to its request gap on that connection (up to about 40 ms), not the service time.

Measured on 0.1.12 on EC2: ExtendDB and PostgreSQL 16 on one c7g.4xlarge, the load generator on a c7g.8xlarge, HTTP/2 over TLS, 32 connections, fixed request rate, 1 KB items. Both runs use the same settings; only the image differs.

0.1.12 0.1.12 with this change
PutItem p50 at 1000 rps 32.2 ms 1.6 ms
UpdateItem p50 at 1000 rps 32.2 ms 1.8 ms
80/20 GetItem/UpdateItem p50 at 4000 rps 8.2 ms 0.5 ms
Highest PutItem rate with p50 ≤ 20 ms and p99 ≤ 100 ms 796 rps 9,000 rps
Highest UpdateItem rate, same limits 796 rps 7,625 rps

With the change, both write rates are within 8% of PostgreSQL alone on the same host (8,500 and 8,250 rps). Latency at 500 rps and peak throughput under overload do not change. Clients that send one request at a time per connection are not affected: botocore at 1000 rps measured p50 2.79 ms before and 2.82 ms after.

Closes # n/a

Testing done

  • Two new unit tests in crates/server/src/lib.rs accept a loopback connection through the real code path and read TCP_NODELAY back from the accepted socket: tls_acceptor_sets_tcp_nodelay_before_the_handshake and plaintext_listener_sets_tcp_nodelay_on_accepted_sockets. Both fail without the change.
  • EC2 before and after runs (table above). A packet capture at 1000 rps confirms the cause: on 0.1.12, 97% of PutItem responses leave about 7 us after the client's next request arrives, 32 ms after their own request. With the change, none do.
  • cargo test --workspace, cargo fmt --check, and cargo clippy --all-targets -- -D warnings pass.

Checklist

  • I have read CONTRIBUTING.md
  • All tests pass (cargo test --workspace)
  • Code is formatted (cargo fmt --check)
  • Clippy is clean (cargo clippy -- -W clippy::pedantic)
  • I have added or updated tests for new functionality
  • I have updated documentation if behavior changed
  • Breaking changes are noted below (if any)
  • If this changes the wire protocol, Storage trait, auth model, on-disk format, or public CLI surface, an RFC has been accepted or is linked below. Otherwise, an ADR captures the decision (link below).

ADR / RFC: n/a (a socket option on accepted connections; no wire, trait, config, or on-disk change)

Breaking changes

n/a


By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0 and I agree to the Developer Certificate of Origin (DCO). See CONTRIBUTING.md for details.

The server never set TCP_NODELAY on the sockets it accepts. With Nagle's algorithm on, the kernel holds a small write while an earlier small write on the connection is unacknowledged, and the client's delayed ACK leaves with its next request on that connection or when its timer fires (about 40 ms on Linux). Once one response is written while the previous one is unacknowledged, each response waits for the client's next request, and the reply to that request is held behind it in turn. The response latency then equals the gap between the client's requests on the connection, for every gap under the delayed-ACK timer. Measured for PutItem and UpdateItem over HTTP/2 on TLS with 32 connections: p50 = p99 = 32 ms at 1000 rps, 16 ms at 2000, 8 ms at 4000, for a service time under 2 ms; at 500 rps, a 64 ms gap, p50 is 1.7 ms. The same server stack reproduces the signature on loopback against a paced client, and TCP_NODELAY on the accepted socket removes it.

Set the option on the raw TCP socket in both serve paths: in the acceptor that runs before the TLS handshake (axum-server) and, on the plaintext path, through axum's ListenerExt::tap_io. A failure to set it is logged at debug and the connection is served with the kernel default.

Two tests bind a loopback listener, connect, run the real acceptor or listener on the accepted socket, and read TCP_NODELAY back from it. Both fail without the change.

Signed-off-by: Anandh Somasundaram <yesyayen@gmail.com>
@yesyayen
yesyayen marked this pull request as ready for review September 30, 2026 17:17

@jcshepherd jcshepherd left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yesyayen
yesyayen added this pull request to the merge queue Sep 30, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks Sep 30, 2026
@yesyayen
yesyayen added this pull request to the merge queue Sep 30, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks Sep 30, 2026
@yesyayen
yesyayen added this pull request to the merge queue Sep 30, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks Sep 30, 2026
@yesyayen
yesyayen added this pull request to the merge queue Oct 1, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 1, 2026
@yesyayen
yesyayen enabled auto-merge October 1, 2026 17:16
@yesyayen
yesyayen added this pull request to the merge queue Oct 1, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks Oct 1, 2026
@yesyayen
yesyayen added this pull request to the merge queue Oct 1, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks Oct 1, 2026
@yesyayen
yesyayen enabled auto-merge October 2, 2026 18:35
@yesyayen
yesyayen added this pull request to the merge queue Oct 2, 2026
Merged via the queue into ExtendDB:main with commit c178814 Oct 2, 2026
24 checks passed
@yesyayen
yesyayen deleted the fix/server-tcp-nodelay branch October 2, 2026 19:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants