Summary
OnTcpAcceptCompletion handles a failed accept by printing accept error: {res} and, because the multishot terminated, re-arming SubmitAcceptMultishot at once. When the failure is -EMFILE or -ENFILE the pending connection stays in the backlog, the re-armed accept fails again immediately, and the reactor spins - accept, CQE(-24), log line, re-arm - as fast as the ring turns.
Repro
Tcp/Raw with PLAYGROUND_REACTORS=2 under prlimit --nofile=80:80 (44 fds are in use at idle), then a client that opens 120 connections and holds them for 3 seconds. The kernel completes the handshakes, so the client sees nothing wrong:
client holds 120 connections
server cpu during the 3s hold: 600 ticks of 10ms (= both reactors at 100%)
stderr lines: 5216092
2770071 [r1] accept error: -24
2446018 [r0] accept error: -24
About 1.7 million log lines per second, for as long as the descriptors stay exhausted.
Impact
Running out of descriptors is a state any server can reach - a connection burst, a leak elsewhere in the process, a low ulimit -n in a container. Here it turns into a CPU-bound spin on every reactor plus a log flood that fills a disk in minutes. Existing connections are still served between re-arms, at a fraction of the rate. Relates to #97 (per-reactor connection caps) and #101 (counters that would make the condition visible).
Fix
On -EMFILE, -ENFILE, -ENOBUFS, -ENOMEM: do not re-arm inline. Re-arm from the 250 ms ticker (it already exists) or from the next close() on that reactor, and rate-limit the log line. The classic alternative - keep a spare descriptor, close it, accept + close the waiting connection, reopen the spare - also ends the storm and tells the peer something instead of leaving it in the backlog.
Summary
OnTcpAcceptCompletionhandles a failed accept by printingaccept error: {res}and, because the multishot terminated, re-armingSubmitAcceptMultishotat once. When the failure is-EMFILEor-ENFILEthe pending connection stays in the backlog, the re-armed accept fails again immediately, and the reactor spins - accept, CQE(-24), log line, re-arm - as fast as the ring turns.Repro
Tcp/RawwithPLAYGROUND_REACTORS=2underprlimit --nofile=80:80(44 fds are in use at idle), then a client that opens 120 connections and holds them for 3 seconds. The kernel completes the handshakes, so the client sees nothing wrong:About 1.7 million log lines per second, for as long as the descriptors stay exhausted.
Impact
Running out of descriptors is a state any server can reach - a connection burst, a leak elsewhere in the process, a low
ulimit -nin a container. Here it turns into a CPU-bound spin on every reactor plus a log flood that fills a disk in minutes. Existing connections are still served between re-arms, at a fraction of the rate. Relates to #97 (per-reactor connection caps) and #101 (counters that would make the condition visible).Fix
On
-EMFILE,-ENFILE,-ENOBUFS,-ENOMEM: do not re-arm inline. Re-arm from the 250 ms ticker (it already exists) or from the nextclose()on that reactor, and rate-limit the log line. The classic alternative - keep a spare descriptor, close it,accept+closethe waiting connection, reopen the spare - also ends the storm and tells the peer something instead of leaving it in the backlog.