Repository navigation
update fib: solve the problem of low CPU utilization and optimize performance. - #1531
Merged
Merged
Conversation
All three entries move to fib a241ac2 and run under its prefork package: a master starts a child for every two CPUs, each with two Ps and a heap of its own, and the children share the ports through SO_REUSEPORT. On 64 CPUs one process spent much of its time in the collector and on the heap lock; prefork took baseline from 2.79M to 3.48M req/s, limited-conn from 1.63M to 2.66M and async from 1.61M to 1.93M, and latency-1m from 34 to 28 cores. fib and fib-tuned serve /static through fib's FileCache, which keeps the files and their pre-compressed twins in memory and follows the disk through an inotify watch (static-tls 0.74M -> 1.19M). The pg pool is split between the children. fib-tuned runs with GOGC=200, encodes JSON with sonic, and drops the settings prefork makes redundant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…RPC; update to fddcb68 fib and fib-tuned route their requests through fib's Router (chi's API with fib's handlers), reading path parameters with Context.Param, and answer JSON through Context.JSON, encoding/json's encoding in fib and sonic's in fib-tuned through fib's JSONEncoder, so both now declare the routing and response completeness axes. Context.JSON reuses its buffer on HTTP/1: json-tls went from 877k to 983k requests a second on 64 CPUs in fib and from 1.57M to 1.77M in fib-tuned. Both subscribe to unary-grpc and unary-grpc-tls: BenchmarkService, generated by protoc-gen-go and protoc-gen-go-grpc with its imports pointed at fib's grpc package and google.golang.org/protobuf messages, is served by fib's grpc.Server through the router, so gRPC shares 8080 (h2c) and 8443 (h2 over TLS) with HTTP. On 64 CPUs unary-grpc served 7.89M calls a second (fib-tuned 8.14M) and unary-grpc-tls 5.56M. validate.sh: 116 passed, 0 failed, 1 skipped (HTTP/3, no h2load-h3). All three entries move to fib fddcb68. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
lesismal
requested review from
Kaliumhexacyanoferrat and
MDA2AV
as code owners
October 7, 2026 18:33
Contributor
Author
|
After the last submission, the CPU utilization of fib on the HTTPArena 3995WX server remained low, but it performed well on my local 8-core and 16-core devices. I didn't reproduce the low CPU utilization issue on hardware with lower core counts. Recently, we reproduced the issue on a 5995wx server, resolving the low CPU utilization problem and making further optimizations. This PR aims to show the latest results. Thank you! |
Owner
|
/benchmark --save |
Contributor
|
👋 Benchmark request received. A collaborator will review and approve the run. |
Contributor
Benchmark ResultsFramework:
Full log |
Contributor
Author
|
/benchmark-multiple -f fib,fib-tuned,fib-websocket --save |
Contributor
|
👋 Benchmark request received. A collaborator will review and approve the run. |
Contributor
Benchmark ResultsFrameworks: 3 | Test: ✅
|
| Test | Conn | RPS | Rate | CPU/req | p99 | p99.9 | CPU | Mem | Δ RPS | Δ Mem |
|---|---|---|---|---|---|---|---|---|---|---|
| baseline | 4096 | 3,164,504 | — | — | — | — | 6139.4% | 293MiB | +191.9% | +35.6% |
| pipelined | 4096 | 20,397,414 | — | — | — | — | 6555.1% | 285MiB | +342.1% | -4.7% |
| limited-conn | 4096 | 2,548,722 | — | — | — | — | 6146.9% | 298MiB | +468.3% | +292.1% |
| json-comp | 4096 | 297,385 | — | — | — | — | 6340.3% | 516MiB | +73.9% | +26.2% |
| json-comp | 16384 | 296,064 | — | — | — | — | 6380.6% | 1.1GiB | +67.0% | +49.6% |
| json-tls | 4096 | 774,182 | — | — | — | — | 6376.7% | 375MiB | +118.3% | +10.9% |
| 8gbit | 512 | 49,363 | 0.9873 | 54.68us | 120.0us | 1836.0us | 235.8% | 256MiB | +0.9% | +172.3% |
| static-tls | 1024 | 1,229,892 | — | — | — | — | 5751.7% | 350MiB | +356.9% | +75.0% |
| async-db | 1024 | 197,176 | — | — | — | — | 4489.7% | 389MiB | +140.8% | +149.4% |
| baseline-h2 | 256 | 5,634,126 | — | — | — | — | 2505.0% | 504MiB | +727.8% | +154.5% |
| baseline-h2 | 1024 | 5,490,643 | — | — | — | — | 2464.0% | 804MiB | +586.6% | +88.3% |
| static-h2 | 256 | 930,877 | — | — | — | — | 6273.2% | 778MiB | +819.0% | +157.6% |
| static-h2 | 1024 | 945,065 | — | — | — | — | 6226.0% | 1.1GiB | +581.3% | +24.5% |
| baseline-h2c | 256 | 12,891,491 | — | — | — | — | 4912.5% | 421MiB | +1821.1% | +137.9% |
| baseline-h2c | 1024 | 13,903,429 | — | — | — | — | 5963.3% | 642MiB | +1634.9% | +106.4% |
| baseline-h2c | 4096 | 11,484,657 | — | — | — | — | 6315.3% | 1.7GiB | +1026.9% | +196.1% |
| json-h2c | 1024 | 1,102,054 | — | — | — | — | 6182.7% | 882MiB | +286.5% | +149.9% |
| json-h2c | 4096 | 1,114,439 | — | — | — | — | 6253.7% | 2.3GiB | +230.3% | +259.6% |
| baseline-h3 | 64 | 6,067,295 | — | — | — | — | 3554.6% | 402MiB | +744.1% | +243.6% |
| static-h3 | 64 | 748,540 | — | — | — | — | 4744.4% | 872MiB | +1921.2% | +188.7% |
| unary-grpc | 256 | 6,198,578 | — | — | — | — | 5449.7% | 603MiB | NEW | NEW |
| unary-grpc | 1024 | 6,199,041 | — | — | — | — | 6198.4% | 1.1GiB | NEW | NEW |
| unary-grpc-tls | 256 | 4,246,784 | — | — | — | — | 3672.7% | 562MiB | NEW | NEW |
| unary-grpc-tls | 1024 | 4,160,362 | — | — | — | — | 3662.4% | 1022MiB | NEW | NEW |
| latency-1m | 1024 | 998,149 | 0.9981 | 30.64us | 204.0us | 4041.0us | 3436.9% | 260MiB | +4.8% | +89.8% |
| latency-10k | 1024 | 9,981 | 0.9981 | 45.58us | 103.0us | 229.0us | 43.5% | 161MiB | ~0% | +235.4% |
| latency-500k-8cpu | 1024 | 444,092 | 0.8882 |
17.79us | 2644992.0us | 2810880.0us | 789.5% | 50MiB | +31.1% | -10.7% |
| async | 32000 | 1,798,497 | — | — | — | — | 5580.7% | 1.5GiB | +73.3% | +176.8% |
✅ fib-tuned
| Test | Conn | RPS | Rate | CPU/req | p99 | p99.9 | CPU | Mem | Δ RPS | Δ Mem |
|---|---|---|---|---|---|---|---|---|---|---|
| baseline | 4096 | 3,136,110 | — | — | — | — | 6341.3% | 467MiB | +380.2% | +148.4% |
| pipelined | 4096 | 20,380,428 | — | — | — | — | 6635.6% | 477MiB | +318.8% | +49.5% |
| limited-conn | 4096 | 2,554,016 | — | — | — | — | 6003.1% | 489MiB | +747.3% | +232.7% |
| json-comp | 4096 | 509,221 | — | — | — | — | 6424.1% | 842MiB | +174.7% | +84.2% |
| json-comp | 16384 | 506,000 | — | — | — | — | 6156.8% | 5.7GiB | +145.1% | +505.5% |
| json-tls | 4096 | 1,505,728 | — | — | — | — | 6225.1% | 658MiB | +329.3% | +83.8% |
| 8gbit | 512 | 49,346 | 0.9869 | 54.25us | 117.0us | 3141.0us | 232.9% | 474MiB | +0.1% | +369.3% |
| static-tls | 1024 | 1,225,298 | — | — | — | — | 5673.7% | 567MiB | +456.0% | +229.7% |
| async-db | 1024 | 230,913 | — | — | — | — | 4259.3% | 674MiB | +193.7% | +326.6% |
| baseline-h2 | 256 | 5,636,472 | — | — | — | — | 2296.4% | 811MiB | +615.9% | +270.3% |
| baseline-h2 | 1024 | 5,511,179 | — | — | — | — | 2323.3% | 1.2GiB | +460.1% | +208.7% |
| static-h2 | 256 | 946,028 | — | — | — | — | 6133.8% | 1.1GiB | +821.6% | +208.6% |
| static-h2 | 1024 | 945,950 | — | — | — | — | 6018.8% | 2.0GiB | +658.2% | +187.2% |
| baseline-h2c | 256 | 13,245,546 | — | — | — | — | 4889.2% | 669MiB | +1699.0% | +310.4% |
| baseline-h2c | 1024 | 14,246,107 | — | — | — | — | 5904.5% | 1.0GiB | +1460.3% | +173.8% |
| baseline-h2c | 4096 | 12,405,361 | — | — | — | — | 5345.7% | 1.8GiB | +894.0% | +251.1% |
| json-h2c | 1024 | 2,337,570 | — | — | — | — | 6114.7% | 1.6GiB | +975.3% | +463.0% |
| json-h2c | 4096 | 2,254,913 | — | — | — | — | 6295.1% | 4.3GiB | +684.7% | +684.9% |
| baseline-h3 | 64 | 6,167,355 | — | — | — | — | 3546.0% | 766MiB | +787.4% | +549.2% |
| static-h3 | 64 | 770,725 | — | — | — | — | 4863.4% | 1.8GiB | +1819.4% | +572.7% |
| unary-grpc | 256 | 6,594,862 | — | — | — | — | 5046.2% | 873MiB | NEW | NEW |
| unary-grpc | 1024 | 6,699,375 | — | — | — | — | 5429.4% | 1.5GiB | NEW | NEW |
| unary-grpc-tls | 256 | 4,331,789 | — | — | — | — | 3283.1% | 873MiB | NEW | NEW |
| unary-grpc-tls | 1024 | 4,233,517 | — | — | — | — | 3353.1% | 1.5GiB | NEW | NEW |
| latency-1m | 1024 | 997,842 | 0.9978 | 30.34us | 196.0us | 5578.0us | 3405.1% | 442MiB | +48.4% | +229.9% |
| latency-10k | 1024 | 9,983 | 0.9983 | 47.59us | 105.0us | 287.0us | 45.8% | 413MiB | ~0% | +742.9% |
| latency-500k-8cpu | 1024 | 449,686 | 0.8994 |
17.63us | 2837504.0us | 3032064.0us | 792.4% | 75MiB | +30.5% | +56.2% |
| async | 32000 | 1,825,758 | — | — | — | — | 5560.7% | 2.2GiB | +98.8% | +453.5% |
✅ fib-websocket
| Test | Conn | RPS | CPU | Mem | Δ RPS | Δ Mem |
|---|---|---|---|---|---|---|
| echo-ws | 512 | 3,539,300 | 6064.5% | 99MiB | +197.9% | +98.0% |
| echo-ws | 4096 | 3,972,953 | 6458.6% | 133MiB | +192.5% | +44.6% |
| echo-ws | 16384 | 3,748,427 | 6142.8% | 266MiB | +229.1% | +35.7% |
| echo-ws-pipeline | 512 | 46,941,744 | 6273.1% | 100MiB | +218.8% | +81.8% |
| echo-ws-pipeline | 4096 | 51,031,532 | 6617.6% | 134MiB | +178.1% | +48.9% |
| echo-ws-pipeline | 16384 | 48,606,448 | 6206.1% | 268MiB | +209.2% | +32.0% |
| echo-ws-limited | 512 | 2,053,428 | 5403.9% | 246MiB | +380.7% | +261.8% |
| echo-ws-limited | 4096 | 2,613,729 | 5887.0% | 277MiB | +443.8% | +123.4% |
MDA2AV
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
PR Commands — comment on this PR to trigger (requires collaborator approval):
/benchmark -f <framework>/benchmark -f <framework> -t <test>/benchmark -f <framework> --save/benchmark -f <framework> -t <test> --save/benchmark -f <framework> --compare <other>/benchmark-multiple -f <fw1>,<fw2>,...-tand--savetoo; saved results land in a single commit/benchmark-multiple --save-fneeded: benchmark and save every framework the PR touches/benchmark-test -t <test><test>and save the resultsFor
/benchmark, always specify-f <framework>; the flags combine in any order. Results come back as a comment with a per-profile table of RPS, p99, CPU and memory — one table per framework on multi runs. A new benchmark comment while a run is in flight queues behind it (one deep) instead of cancelling it. For multi-framework PRs (dependency bumps, same-language refactors) prefer/benchmark-multiple, which runs everything in a single job and commits all saved results together, so no run overwrites another.--compareworks on single-framework runs only.What the deltas are measured against. By default, this framework's own results published on
main- answering "did this change help?". When you are tuning a variant or a successor entry,--comparere-bases them on another entry instead:The reply states which baseline it used, and profiles the other framework does not run show
n/arather than a delta.Run benchmarks locally
You can validate and benchmark your framework locally with the lite script — no CPU pinning, fixed connection counts, all load generators run in Docker.
Requirements: Docker Engine on Linux. Load generators (gcannon, h2load, h2load-h3, wrk) are built as self-contained Docker images on first run.