gc(tenuring): the survival-rate lock latches S=1 only on steady-state evidence, never from the startup cohort - #9949
proggeramlug wants to merge 3 commits into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
2f7d203 to
a771c33
Compare
|
Closing (owner-approved stale-draft cleanup, 2026-09-24). It measured −7.7 % over 4 cc turns but was held for an unresolved intermittent turn-3 crash, and it now conflicts with main. If the tenuring-lock idea is revisited, re-derive it on the current tenuring code rather than porting this branch. |
|
Reopened (owner decision, 2026-09-24). My earlier close undersold this PR: it measured −7.7 % over 4 cc turns (8.87 s vs 9.61 s main-adaptive, lowest promoted bytes t2–t4, peak/settled equal to S=2). It was held only by cc's intermittent turn-3 TypeError, and on 09-09 that crash was attributed away from this PR to two move-sensitive rooting bugs in
Plan: fix (A) and (B) in one PR with a moving-GC test first, then rebase and re-validate this one. It conflicts with main now. |
a771c33 to
5c16a19
Compare
Exclude startup cohorts from promote-on-first-copy decisions and require three consecutive substantial high-survival rounds before latching S=1. Gate sweep seeding on the same startup boundary and expose cumulative copy and promotion price inputs under GC diagnostics.
Record the lock map, diagnostic schema, sabotage cases, disk-blocked gates, and the exact TN follow-up request with falsifiable predictions.
…SS1_MARKED window The rebase onto main re-introduces two thread-locals the root-holder gate wants a verdict for, and again moves a source the PASS1_MARKED pin covers. `RATED_ROUNDS` (Cell<u64>) and `PROMOTE_LOCK_STREAK` (Cell<u8>) hold a rated cohort count and a streak count; neither is ever an address or a NaN-boxed value, so both are `not_a_gc_pointer`. The PASS1_MARKED pin covers `gc/mod.rs`, which this branch touches. Re-audited rather than repinned blind: the change is confined to `emit_gc_time_share_diag`, which reads four relaxed atomic counters and widens one `eprintln!` -- it allocates nothing, collects nothing and calls no JS. Its only caller is `emit_incremental_liveness_diag`, gated on `telemetry::gc_diag_enabled()` (default off), main-thread only, latched to emit once per process, and reached only from `js_gc_release_current_thread_collection_side_allocations`, the every-process-exit funnel. The added code therefore cannot run inside `cycle.rs::run_to_completion`, and the mark-complete to sweep-entry window is unchanged. Only that one hash moved; the other four pinned sources still match, which is how the hash computation itself was checked before the pin was touched.
5c16a19 to
4b912d9
Compare
Re-measured on current main: the mechanism still fires. The CPU delta is not reportable from this run.Rebased onto main The mechanism is confirmed, and it is unambiguous
Per-turn copying is These two signals are load-independent, which matters on this host (below). Correction to an earlier reading: the "+8% full collections" is an artefact, not an effectI previously reported arm B at 346 full collections against arm A's 321. That was my aggregation error — unbalanced n across rounds. Full collections are wall-time-proportional at the same rate in both arms:
The trigger mix is also the same: Why no CPU numberThe host was shared and heavily oversubscribed: load ran 84 → 141 on 48 threads and rose monotonically through the run. The harness's own contamination stamp makes the scale concrete — So September's −7.7% four-turn figure is neither confirmed nor refuted here. Instruction-count rows ( |
Instruction counts: more expensive in all seven paired rounds, and the cost is bimodalWhole-process counters now, on the same rig. Paired rounds only — a round enters the table only if both arms produced a complete four-turn row, so n is balanced by construction — arm order alternating, Read the instruction column, not the cycle column: arm A's instruction count varies by 1.6% across rounds while its cycle count varies by 51.8% under the host's load. That is precisely why the cycle and wall-clock numbers were not reportable before, and why these are. Set 1 — five paired rounds
Paired delta: mean +11.48%, median +16.11%, range +1.39% to +19.60%. Arm B is never cheaper. Set 2 — two more pairs with a graceful exit, load-matchedCtrl-C twice instead of SIGKILL, so the process-exit funnel runs and prints
The delta is not a spread around one value — it is two modes, and the collection count picks itAcross all seven rounds:
The extra collections are all So the split is: when arm B takes the same number of full collections as arm A it costs +1.4% and +2.3% instructions; when it takes ~51 more it costs +16% to +20%, and +18% to +23% CPU. The couple of percent at equal collection count is the copying work itself — 11.5 MB evacuated per four turns — plus the added instruments. What sets the mode is not load, and not arm orderI first suspected load, because in set 1 the high mode fell in the three quiet rounds. Set 2 refutes that: at load 139/139 and 130/144, arm B is firmly in the high mode, while in set 1 it was in the low mode at loads 92 and 139. Arm order does not explain it either — round 2 ran arm B first and went high, round 4 ran arm B first and went low. Arm A never shows a second mode in any of the seven rounds. The practical consequence is that this change cannot be evaluated from a small number of runs. Two honest four-turn measurements of the same two binaries can land 16 points apart depending on which mode arm B picks, which is a plausible account of how the original −7.7% and my +11.5% mean can both be real measurements of the same code. The mode needs identifying before an accept/reject decision means anything. Corrections to my previous commentTwo things in it were wrong, both mine:
Memory, and the mechanism
The mechanism is intact and reproduces in every single round: the control evacuates nothing, arm B copies ~11.5 MB per four turns and promotes ~18% less, and settled RSS is ~1.4% lower. That memory win is real but small at this workload, and it does not pay for a 16–20% instruction cost in the rounds where the extra collections appear. For context on why collection count dominates everything here: One rig note for anyone reproducingWrapping the stream harness in Raw rows, per-run diags, the pinned base, the patch and the patched harness are preserved and I can hand them over. |
|
An eighth pair finished just after I posted, and it refines the CPU claim above, so here it is rather than left out.
Arm B is in the low mode here, and in that mode it is at parity: +0.73% instructions and −0.47% CPU, i.e. marginally faster. So the "+18% to +23% CPU" from the two load-matched pairs above is the cost of the high mode specifically, not of the change in general — the corrected reading is:
Final census over all eight pairs — arm B fulls That makes the question sharper rather than softer. The change is either free-with-a-memory-win or a 20% regression depending on a mode nobody has yet identified, and it flips between them on identical binaries and identical input. Worth finding the switch before deciding on the PR — if the low mode can be made the only mode, the original case for this change stands up. |
Bimodality root-caused: it is not a latch and not the tenuring lock — the control's collection count is bounded, this PR's tracks churn, and the workload straddles a stepInvestigated on the 16 runs already captured (8 per arm, same pinned base It is not the tenuring lock, and there is a clean counterexampleI first went after the lock and the nursery ladder, because the first divergence between a low-mode and a high-mode run looks like one: the first five That is a red herring. Across all eight armB runs none of the lock-side variables separate the modes:
What does separate them: total bytes through the collectorSumming
The control is bounded-cost. Its collection count is 334 in seven of eight runs (332 once) while its churn varies 14.8% — it absorbs the entire variation in how much each collection frees (16.6 → 19.0 MB). This PR removes that property: the count follows churn at r = +0.88. And the response is a step, not a slope:
Two discrete operating points. Within each, churn moves 7% and 12% while the count barely moves; across the gap between 8504 and 8898 MB — 4.6% more churn — the count jumps 15.3%. Per-collection work drops from 25.5 to 23.1 MB across the step, i.e. above it the collector does more, smaller collections. So the mechanism is
This is why the mode looked load- and order-independent: churn is neither. It is not a one-time latch — nothing needs to flip once per process. Implication for the PR, and the fix directionThe −18% promotion and −1.4% settled RSS are real and reproduce in every round, so the goal Ralph set — keep the promotion win, drop the extra collections — looks reachable, and the two candidate directions are different in kind:
Acceptance for either should be the same instrument used here, since it is what made the effect legible: report One honest limit: |
Correction, and the actual cause: the PR's evidence gate can never fire on this workload, so the threshold lands on 3 instead of 1 — and that doubles the nurseryMy churn explanation in the comment above was wrong and I want it corrected before anyone acts on it. I said the extra collector traffic was "consistent with copied survivors being re-swept". It is not: copying moves only ~35 MB over a whole run, while the freed difference is 2.4–3.6 GB. Two orders of magnitude off. The step and the correlation stand; the mechanism I offered for them does not. Here is what the data actually says. One line, present in all 16 runs without exception:
The control settles on 1 — promote on first copy, which is why its That single policy difference explains every number in this thread. Comparing the pacer state at
Holding three cohorts instead of one needs twice the nursery, so every full collection sweeps 2.2x more young data — and Why it lands on 3 instead of 1From the PR's own documentation of the lock: the first two rated cohorts are excluded as startup, then three consecutive cohorts must each be ≥90% surviving to lock the threshold to 1. Every That is 22% to 38% survival — the ≥90% bar (900 permille) is never met, so So this is not a latch that flips once, and not noise amplification. It is a gate whose entry condition is unreachable for a workload with ~30% survivor survival, with a fallback that moves in the opposite direction to the gate's intent. The bimodality is then just the churn step from the earlier comment acting on a 2x-larger young generation. Revised fix direction"Bound the re-copying" was aimed at the wrong thing, since re-copying is 35 MB. The lever is the threshold itself:
Option 1 is the one that matches the measurements: it should keep most of the −18% promotion reduction (promotion is already down at threshold 3 mainly because 0.3³ of a cohort reaches old age) while returning |
Correction: with collection counts matched, this PR costs +1.55% instructions, not +16–20% — and my proposed fix is refutedI ran the threshold matrix Ralph asked for, and it inverts my two comments above. Posting it immediately because those comments would sink this PR on wrong grounds. Method. Five cells from the two existing binaries, no rebuild, using the
Against the control: B_auto is +1.55% instructions, −18.87% promoted, −1.71% settled RSS. Collection counts are 331–334 in every cell, so this is a like-for-like comparison, and it is a materially better result for the PR than anything I posted earlier. What I got wrong
What survives, and is the real open questionThe +16–20% I measured earlier was real, but it is an intermittent excursion, not the steady state. Collection counts across every run I have:
The excursion is worth +16% instructions when it happens, it is specific to the armB binary — never once in 16 control runs, and it is not selected by the threshold (it hit the pinned cells in this matrix and the adaptive cell in the earlier set) nor cleanly by load, though its rate was higher in the loaded set (5 of 8) than the quiet one (2 of 12). Taking the observed rate at face value, expected cost is roughly 0.35 × 16% + 0.65 × 1.5% ≈ 6%, with a bimodal distribution rather than a spread. So the decision Ralph asked me to inform is narrower than I made it look: the steady state is a good trade (+1.5% compute for −18.9% promotion and −1.7% settled RSS), and the open risk is a ~35%-frequency excursion that doubles the collection count and that the control never shows. That excursion is what still needs root-causing, and my threshold hypothesis for it is dead. It is also why the same binaries gave −7.7% and +11.5% in different hands. Method note for anyone re-running: report the collection count next to every timing number. Every wrong conclusion in this thread, mine included, came from comparing two runs that had silently landed on different collection counts. Raw rows, per-run diags and the matrix script are preserved with the rest at |
Runtime-only, stacked on #9861 (
2248fba56). Written by codex from the campaign's TN4 measurement; not yet compiled — perrymaster's gate ladder and the TN5 rows will be appended here.Why
TN4 (main + #9861, 4-turn cc replies, quiet host): the adaptive tenuring loop's only S change is
survivals 2 -> 1 (lock, eden_live_bytes=3.2 MB desired=2.1 MB)at startup, right after the second copying minor — the survival-rate lock rates the process's startup cohort (copied once at S=2, returned nearly intact) and latches promote-on-first-copy before the first turn. From then on it promotes ~7 MB per turn where S=2 pinned promotes 4.3, and turn 1 costs +0.3…0.4 s against S=2 pinned (and +0.5 against S=1 pinned) with extra full-cycle work; turns 2–4 are equal across arms; S=2 pinned is the best arm on four-turn CPU (−5…6 %), peak and settled RSS. The power-on fix in #9861 gated the occupancy rule from claiming the ceiling before a round is measured; the lock still concluded from the first rated round. This is the campaign design note's item 1: the first rateable cohort of a process is the least representative one.What changes
prev_cohort_copied > 0(the cohort-scoped signal from fix(gc): the tenuring occupancy rule may not claim promote-on-first-copy (#9851) #9861). The lock's entry now requires the round to be post-startup (startup = the process's first two rated survivor cohorts — the loop's own threshold-invariant evidence; the allocation census is deliberately not the marker because it seeds before the first copying minor), with fresh intake ≥desired / 4(the existing substantial-volume bar) and survival ≥ 90 % (the existing bar), for K = 3 consecutive rated rounds (PROMOTE_LOCK_STREAK); any rated round failing those resets the streak. K = 3 is the smallest window that rejects a one- or two-cycle phase boundary while a truly non-dying workload still reaches S=1 within five rated cohorts.seed_promote_lock_from_sweepkeeps its two conditions but refuses while fewer than two cohorts are rated. The unlock path is unchanged except that unlocking clears the entry streak.PERRY_GC_DIAGeach copying minor accumulates copied and promoted bytes; with the cumulative minor pause andstep_us + remark_usthese givecopy_cost = copy_pause_us / tenuring_copied_bytesandpromote_cost = promote_us / tenuring_promoted_bytes, the two terms of the design's decision rule (keep aging while mortality > copy_cost / promote_cost). The 90 % constant stays until TN5 supplies both prices aligned.rounds_rated= streak= survival_permille= copied_bytes= startup= copy_pause_us= tenuring_copied_bytes= promote_us= tenuring_promoted_bytes=; the[gc-time]exit block gains the four cumulative fields. No cost when diagnostics are off.Tests (named; sabotage stated in the campaign report; not yet executed)
startup_shaped_survivors_do_not_contribute_to_the_lock_streak,k_steady_fully_surviving_rounds_latch_promote_on_first_copy(the kill condition: a non-dying workload must still latch),mortality_inside_the_steady_window_resets_the_lock_streak,sweep_seed_cannot_latch_from_a_startup_census; #9861's pinned-S tests andoccupancy_may_not_claim_the_ceiling_before_any_round_is_measuredunchanged.Predictions for TN5 (falsifiers)
On cc, the adaptive arm shows no latch during startup; turn 1 equals S=2 pinned within noise; turns 2–4 equal S=2 pinned; promoted bytes in turns 2–4 ≈ 12 MB; peak and settled RSS equal S=2 pinned's. Gates run so far: rustfmt, diff check, file-size gate, root-holder inventory + self-test. Not run (disk): the runtime suite, the archive build, the default compiler build. Draft until perrymaster's ladder and rows are on this PR.
Measured — TN5 (perrymaster, main
504e180d0+ #9861 + this commit, runtime relinked on main's cache; gate: default build, archive feature set, runtime suite 3238 passed / 0 failed / 4 ignored with all 20 named tests green; 4-turn graceful runs, quiet box, stamped rows)400: adaptive 0.73, 0.34, 0.53, 0.36 (S2:7) = pinned S=2 0.76, 0.34, 0.58, 0.37; pinned S=1 0.70, 0.40, 0.47, 0.36. Second adaptive round (re-run clean, load 0.07): 3.07, 1.94, 1.97, 2.03 → 9.01; S-history S2:16 S3:7 S4:2, every transition occupancy-driven with
streak=0(rated survival 10–699 ‰, so the lock never has evidence on cc); promoted t2–4 9.8 MB; peak 618 / settled 523 — the same as round 1 and as S=2 pinned.Reading: the predictions hold. The adaptive arm never latches promote-on-first-copy on cc; it climbs the occupancy ladder (2 → 3 → 4 → 2 …) on measured rounds, turn 1 equals S=2 pinned, turns 2–4 equal S=2 pinned, steady promotion is the lowest of all arms (10.0 MB), and the four-turn total (8.87 s) beats every pin and improves on main's adaptive by −7.7 %. Peak RSS within S=2 pinned's range, settled equal. Prices from the new counters (per run): copy ≈ 10.5 ns/byte at S=2 and ≈ 22.8 at the adaptive arm's higher ages, promote ≈ 32–57 ns/byte; the design's decision rule now has both terms on file.
Local gates on a771c33 (rebased onto main 8b7dc33, patch-identical; macOS arm64, 2026-09-08): runtime lib suite one thread 3,267 passed / 0 failed / 4 ignored;
cargo build --release -p perry-runtime --features wasm-hostok. Rows (TN5, perrymaster): adaptive no longer latches S=1 from the startup cohort; 4-turn 3300 8.87 s vs main-adaptive 9.61 (−7.7 %), promoted t2–4 10 MB, peak/settled = S=2 pinned.HOLD (2026-09-08): an intermittent turn-3 cc crash reproduces on this PR's promotion regime
On perrymaster, a 4-turn cc run on this tree (with the #9965 verifier fixes, verifier live on all 17 minors, zero verifier findings) died at turn 3 with
Cannot convert undefined or null to objectin the React commit's Yoga layout pass (WN8.setGapPercent → sJ7.hook → get years). The same family shows 2/7 on #9951's regime and 0/5 on main6, so it correlates with non-default promotion regimes rather than with any single PR; the evacuation verifier does not see it, which points at a runtime-side registry holding a heap pointer outside its walk. Until the mechanism is found (from-space protection runs and a Yoga-registry GC audit are in progress), this PR is not a landing candidate regardless of its rows.https://claude.ai/code/session_011dhBmdn4vGgNibjo3oZqTo