HOLD (RSS trade): perf(regex): stop charging operation scratch to GC pressure; grow lent registers (bounded) - #11612
proggeramlug wants to merge 2 commits into
Conversation
… grow the lent registers to a fixed bound Regex operation scratch (perex_memory::Buffer / Reservation: the owned search path's match buffers, compile scratch, KMP failure tables, replacer argument slots) is freed by the operation that allocated it and bounded by that operation's MemoryBudget. Reporting it through gc_note_external_side_alloc/free put every per-call release into GC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure until the next full, so a loop over a program with more than 32 registers (dotenv's LINE has 42) ran a full mark-sweep every few hundred calls on bytes no collection could free. The lent per-thread scratch cell now grows its registers on demand up to LENT_REGISTERS = 1024 (8 KiB per thread), so such programs stop building per-call buffers at all. Programs past the bound keep the owned path. Subject-proportional storage (replace Spans, native piece records) is unchanged and still reported. Part of #11549
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
The survival-aware nursery pacing this PR was waiting for is in draft #11645. It carries this PR's two commits (and #11634's) unchanged, so they can land together. It also repairs a gate that this PR alone turns red on current main: |
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
…rs inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer.
|
This lands via #11645, which carries its two commits unchanged ( |
…nflux ladder (includes #11612) (#11645) * perf(regex): stop charging operation scratch to GC external pressure; grow the lent registers to a fixed bound Regex operation scratch (perex_memory::Buffer / Reservation: the owned search path's match buffers, compile scratch, KMP failure tables, replacer argument slots) is freed by the operation that allocated it and bounded by that operation's MemoryBudget. Reporting it through gc_note_external_side_alloc/free put every per-call release into GC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure until the next full, so a loop over a program with more than 32 registers (dotenv's LINE has 42) ran a full mark-sweep every few hundred calls on bytes no collection could free. The lent per-thread scratch cell now grows its registers on demand up to LENT_REGISTERS = 1024 (8 KiB per thread), so such programs stop building per-call buffers at all. Programs past the bound keep the owned path. Subject-proportional storage (replace Spans, native piece records) is unchanged and still reported. Part of #11549 * changelog: #11612 regex scratch is not GC pressure * perf(gc): survival-aware nursery pacing: the ladder starts at a quarter of the base The scavenge nursery cap's influx-driven ladder (#7377) now counts in quarters of the base cap and powers on at a quarter of it: 4 MB instead of 16 MB. A thread stays there while its minors find little alive and climbs toward the unchanged 64 MB top while survivor influx stays above 4% of the cap, with the same debounced 4%/1% band. A move now takes every step the one-step rule would take in a row on the same reading, so a survivor-heavy program settles at the size it settled at before without two extra minors per step. The landing is inside the dead band, so it cannot oscillate. The survivor target, the promoted-cohort floor, the allocation-census seed point and the JSON-leaf gate stay keyed on the 16 MB base: they measure lifetimes in bytes allocated, which a smaller Eden must not shorten. Part of #11549. * chore(gc): classify the external-side byte counters in the root-holders inventory #11612 stops regex operation scratch from reporting to GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the reachability walk no longer reaches them from a registered scanner and gc_runtime_root_holders.py flags both as new rule-T holders. They are usize byte counters and cannot hold a GC pointer. * changelog: #11645 survival-aware nursery pacing --------- Co-authored-by: Perry Bot <perry-bot@users.noreply.github.com>
Part of #11549
What changed
regex/perex_memory.rs).BufferandReservationcover the owned search path's match buffers, the compiler's node and range scratch, the KMP failure table and the replacer argument slots. They used to callgc_note_external_side_allocandgc_note_external_side_free. Every one of them is freed by the operation that allocated it, when it returns or unwinds. No collection can reclaim a byte of it. But each release was added toGC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure until the next full collection. So a loop over a program with more than 32 registers ran a full mark-sweep every few hundred calls on phantom bytes. InPERRY_GC_DIAGoutput that shows asexternal_drained=33553968,arena_total=3145728,old_in_use=0, 81 times in one run.MemoryBudgetlimit (perex_api::SCRATCH_BYTES) before it exists. That check is unchanged.perex_replace_direct::Spansand the native piece records inperex_replace_storagestill report, because their size follows the input rather than a fixed limit. Map, Set, JSON-tape and N-API external bytes are untouched.StorageError::Abruptis deleted. Its only producer was the collection that the note could trigger.regex/perex_runtime.rs). Registers used to be a fixed[usize; 32]. They are now aVecgrown on demand toLENT_REGISTERS = 1024, so the cell retains at most 8 KiB per thread whatever programs run. dotenv'sLINEneeds 42 registers, so with the fixed 32 every search built and dropped owned buffers. It now builds them once per thread. A register count is a property of the program, so the choice of path is made before any work is done, never through a failed attempt. A program past 1024 registers still takes the owned path, which the budget bounds and the operation frees.framesandundovectors already grew on demand and were never reported. They are bounded only indirectly, by the per-operationChargeover the whole cell. I left that alone here.Measurements (perrymaster, Linux x86_64; both arms built here with
--releasefrom 5af043c, differing only by this diff;PERRY_NO_AUTO_OPTIMIZE=1)Instruction counts are
perf stat -e instructions:u, a two-N differential with the median of 3 at each N, taken under/tmp/perry-bench-lock.d. The host load average was about 23. Instructions per iteration:RE.testuuid, 2-register program (control)LINE.execover 40 lines (42 registers)s.replace(/(\d+)/g, fn)Package workloads (
scripts/package_bench.py run --arms perry --modes instr). Every run's output matched Node 26.5.1 byte for byte:Collections (
PERRY_GC_DIAG=1, one run at n2) and peak RSS (/usr/bin/time -f %M, median of 11 runs, with the [min..max] range):main's RSS on dotenv is bimodal and depends on load, because the budgeted fulls are time-sliced. In an earlier, quieter set of 11 runs its median was 31.5 MB against this PR's 37.8 MB. The PR arm is stable in both sets.
Where the RSS goes, and what recovers it
Without the phantom fulls these loops pace the way every other allocating Perry program already does: the young generation fills to the 16 MB scavenge nursery cap before a copying minor runs. main was collecting them at about 3 MB of young data, using full collections that cost 12.7% of dotenv's wall time (
[gc-time] share_permille=127, against 10 with this PR).PERRY_GC_SCAVENGE_NURSERY_MBmeasures the lever directly, on this PR's binaries (instructions per iteration, RSS from one run):A 4 MB nursery would give both wins on the regex rows: −29.8% and −13.6% instructions at main's RSS. But a smaller global nursery costs +3% instructions on an allocation-bound loop. So lowering the nursery cap is itself a trade, and it does not belong in this PR.
Per minor, the fixed cost is about 310–380k instructions on the alloc loop and about 720k on dotenv. A symbolized profile at a 1 MB nursery puts most of the minor-side cost in
RuntimeRootVisitor::visit_tagged_raw_addr(8.7% of the run) andarray_tail_transition::{scan_table, prune_table}(5.3%). Cutting that fixed cost is what would make a smaller or work-aware nursery free, and it is the natural next step for direction 2.Tests
gc/tests/runtime_roots/perex_scratch_pressure.rshas three tests:Bufferleavesexternal_side_live_bytesand the drained counter untouched, and the budget still bounds it.#[cfg(test)] OWNED_SEARCHEScounter.perex_memory.rs, tests 1 and 3 fail (left: (32768, 0) right: (0, 0); drained200256vs161792). WithLENT_REGISTERS = 32, test 2 fails (65 owned searches, not 1). All three pass with the fix.perex_execution::perex_host_buffers_account_overlap_failure_and_unwindasserted the old contract, that a buffer's bytes appear as external side bytes. It now asserts the budget's overlap accounting and an unchanged external reading.RUST_TEST_THREADS=1 cargo test --release -p perry-runtime(codegen-units 16): 4701 passed, 0 failed, 5 ignored.RUSTFLAGS="-D warnings" cargo check -p perry-runtime --all-targets(dev profile): clean.cargo fmt --all -- --check: clean.scripts/check_file_size.sh: OK./opt/node-v26.5.1-linux-x64),PERRY_SKIP_BUILD=1 PERRY_NO_AUTO_OPTIMIZE=1, withPERRY_BINandPERRY_RUNTIME_DIRpinned to each arm's own-p perry -p perry-runtime-static -p perry-stdlib-staticrelease build. Filters weretest_gap_combined with each of regex, regexp, split, replace, string and gc: 139 distinct tests. Both arms: 134 PASS and the same 5 pre-existing COMPILE_FAILs (test_gap_regex_replace_dyn_regex_with_http,test_gap_11258_eventemitter_async_resource_subclass_gc,test_gap_9552_cross_thread_promise_survives_gc,test_gap_gc_http2_pending_event_callback_rooting,test_gap_gc_net_once_flags_rekey). Per-test results identical.SKIP_COMPILE_GATES=1 scripts/run_lint_gates.sh: 97 of 100 script gates passed; the compile tier was not run. The 3 failures are pre-existing:cargo xwinis not installed on this host.gc_runtime_root_holders.pyflagsGC_EXTERNAL_SIDE_ALLOC_PENDINGandGC_EXTERNAL_SIDE_LIVE_BYTESingc/policy.rs, which this PR does not touch. It fails identically on 5af043c.git diff --statwas clean afterwards.Not run: the full gap sweep,
cargo test --workspace, macOS, the default auto-optimizeperry compilearm, wall-clock measurements (this host is shared and loaded), the Windows xwin check, and the compile tier of the lint gates.No test outside
perry-runtime's regex and GC test modules is expected to change.