Skip to content

gc(tenuring): the survival-rate lock latches S=1 only on steady-state evidence, never from the startup cohort - #9949

Draft
proggeramlug wants to merge 3 commits into
PerryTS:mainfrom
proggeramlug:perf/tenuring-evidence-lock
Draft

proggeramlug wants to merge 3 commits into
PerryTS:mainfrom
proggeramlug:perf/tenuring-evidence-lock

Conversation

@proggeramlug

@proggeramlug proggeramlug commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Runtime-only, stacked on #9861 (2248fba56). Written by codex from the campaign's TN4 measurement; not yet compiled — perrymaster's gate ladder and the TN5 rows will be appended here.

Why

TN4 (main + #9861, 4-turn cc replies, quiet host): the adaptive tenuring loop's only S change is survivals 2 -> 1 (lock, eden_live_bytes=3.2 MB desired=2.1 MB) at startup, right after the second copying minor — the survival-rate lock rates the process's startup cohort (copied once at S=2, returned nearly intact) and latches promote-on-first-copy before the first turn. From then on it promotes ~7 MB per turn where S=2 pinned promotes 4.3, and turn 1 costs +0.3…0.4 s against S=2 pinned (and +0.5 against S=1 pinned) with extra full-cycle work; turns 2–4 are equal across arms; S=2 pinned is the best arm on four-turn CPU (−5…6 %), peak and settled RSS. The power-on fix in #9861 gated the occupancy rule from claiming the ceiling before a round is measured; the lock still concluded from the first rated round. This is the campaign design note's item 1: the first rateable cohort of a process is the least representative one.

What changes

  • A rated round still means prev_cohort_copied > 0 (the cohort-scoped signal from fix(gc): the tenuring occupancy rule may not claim promote-on-first-copy (#9851) #9861). The lock's entry now requires the round to be post-startup (startup = the process's first two rated survivor cohorts — the loop's own threshold-invariant evidence; the allocation census is deliberately not the marker because it seeds before the first copying minor), with fresh intake ≥ desired / 4 (the existing substantial-volume bar) and survival ≥ 90 % (the existing bar), for K = 3 consecutive rated rounds (PROMOTE_LOCK_STREAK); any rated round failing those resets the streak. K = 3 is the smallest window that rejects a one- or two-cycle phase boundary while a truly non-dying workload still reaches S=1 within five rated cohorts.
  • seed_promote_lock_from_sweep keeps its two conditions but refuses while fewer than two cohorts are rated. The unlock path is unchanged except that unlocking clears the entry streak.
  • Prices recorded, not yet deciding: under PERRY_GC_DIAG each copying minor accumulates copied and promoted bytes; with the cumulative minor pause and step_us + remark_us these give copy_cost = copy_pause_us / tenuring_copied_bytes and promote_cost = promote_us / tenuring_promoted_bytes, the two terms of the design's decision rule (keep aging while mortality > copy_cost / promote_cost). The 90 % constant stays until TN5 supplies both prices aligned.
  • Diag: every S transition prints rounds_rated= streak= survival_permille= copied_bytes= startup= copy_pause_us= tenuring_copied_bytes= promote_us= tenuring_promoted_bytes=; the [gc-time] exit block gains the four cumulative fields. No cost when diagnostics are off.

Tests (named; sabotage stated in the campaign report; not yet executed)

startup_shaped_survivors_do_not_contribute_to_the_lock_streak, k_steady_fully_surviving_rounds_latch_promote_on_first_copy (the kill condition: a non-dying workload must still latch), mortality_inside_the_steady_window_resets_the_lock_streak, sweep_seed_cannot_latch_from_a_startup_census; #9861's pinned-S tests and occupancy_may_not_claim_the_ceiling_before_any_round_is_measured unchanged.

Predictions for TN5 (falsifiers)

On cc, the adaptive arm shows no latch during startup; turn 1 equals S=2 pinned within noise; turns 2–4 equal S=2 pinned; promoted bytes in turns 2–4 ≈ 12 MB; peak and settled RSS equal S=2 pinned's. Gates run so far: rustfmt, diff check, file-size gate, root-holder inventory + self-test. Not run (disk): the runtime suite, the archive build, the default compiler build. Draft until perrymaster's ladder and rows are on this PR.

Measured — TN5 (perrymaster, main 504e180d0 + #9861 + this commit, runtime relinked on main's cache; gate: default build, archive feature set, runtime suite 3238 passed / 0 failed / 4 ignored with all 20 named tests green; 4-turn graceful runs, quiet box, stamped rows)

3300, turn CPU t1…t4 → sum S-history promoted t2–4 MB peak / settled MB
adaptive, this PR r1 3.03, 1.83, 1.93, 2.08 → 8.87 S2:17 S3:6 S4:2 — no S=1 latch 10.0
pinned S=2 r1 / r2 3.02, 1.93, 1.94, 2.01 → 8.90 / 3.01, 1.95, 1.95, 2.10 → 9.01 S2 11.6 / 11.7
pinned S=1 r1 / r2 2.92, 2.01, 2.18, 2.02 → 9.13 / 2.87, 2.24, 2.19, 2.01 → 9.31 S1 17.9 / 17.5
adaptive on main + #9861 only (TN4) 3.47, 2.24, 1.91, 1.99 → 9.61 / 3.28, 2.04, 2.17, 2.11 → 9.60 S1:25 S2:2 (latched in startup) 17.5 / 17.7

400: adaptive 0.73, 0.34, 0.53, 0.36 (S2:7) = pinned S=2 0.76, 0.34, 0.58, 0.37; pinned S=1 0.70, 0.40, 0.47, 0.36. Second adaptive round (re-run clean, load 0.07): 3.07, 1.94, 1.97, 2.03 → 9.01; S-history S2:16 S3:7 S4:2, every transition occupancy-driven with streak=0 (rated survival 10–699 ‰, so the lock never has evidence on cc); promoted t2–4 9.8 MB; peak 618 / settled 523 — the same as round 1 and as S=2 pinned.

Reading: the predictions hold. The adaptive arm never latches promote-on-first-copy on cc; it climbs the occupancy ladder (2 → 3 → 4 → 2 …) on measured rounds, turn 1 equals S=2 pinned, turns 2–4 equal S=2 pinned, steady promotion is the lowest of all arms (10.0 MB), and the four-turn total (8.87 s) beats every pin and improves on main's adaptive by −7.7 %. Peak RSS within S=2 pinned's range, settled equal. Prices from the new counters (per run): copy ≈ 10.5 ns/byte at S=2 and ≈ 22.8 at the adaptive arm's higher ages, promote ≈ 32–57 ns/byte; the design's decision rule now has both terms on file.

Local gates on a771c33 (rebased onto main 8b7dc33, patch-identical; macOS arm64, 2026-09-08): runtime lib suite one thread 3,267 passed / 0 failed / 4 ignored; cargo build --release -p perry-runtime --features wasm-host ok. Rows (TN5, perrymaster): adaptive no longer latches S=1 from the startup cohort; 4-turn 3300 8.87 s vs main-adaptive 9.61 (−7.7 %), promoted t2–4 10 MB, peak/settled = S=2 pinned.

HOLD (2026-09-08): an intermittent turn-3 cc crash reproduces on this PR's promotion regime

On perrymaster, a 4-turn cc run on this tree (with the #9965 verifier fixes, verifier live on all 17 minors, zero verifier findings) died at turn 3 with Cannot convert undefined or null to object in the React commit's Yoga layout pass (WN8.setGapPercent → sJ7.hook → get years). The same family shows 2/7 on #9951's regime and 0/5 on main6, so it correlates with non-default promotion regimes rather than with any single PR; the evacuation verifier does not see it, which points at a runtime-side registry holding a heap pointer outside its walk. Until the mechanism is found (from-space protection runs and a Yoga-registry GC audit are in progress), this PR is not a landing candidate regardless of its rows.

https://claude.ai/code/session_011dhBmdn4vGgNibjo3oZqTo

@coderabbitai

coderabbitai Bot commented Sep 7, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Closing (owner-approved stale-draft cleanup, 2026-09-24). It measured −7.7 % over 4 cc turns but was held for an unresolved intermittent turn-3 crash, and it now conflicts with main. If the tenuring-lock idea is revisited, re-derive it on the current tenuring code rather than porting this branch.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Reopened (owner decision, 2026-09-24). My earlier close undersold this PR: it measured −7.7 % over 4 cc turns (8.87 s vs 9.61 s main-adaptive, lowest promoted bytes t2–t4, peak/settled equal to S=2). It was held only by cc's intermittent turn-3 TypeError, and on 09-09 that crash was attributed away from this PR to two move-sensitive rooting bugs in crates/perry-runtime/src/yoga.rs, both still present on main:

Plan: fix (A) and (B) in one PR with a moving-GC test first, then rebase and re-validate this one. It conflicts with main now.

Ralph Küpper added 3 commits September 28, 2026 14:47
Exclude startup cohorts from promote-on-first-copy decisions and require
three consecutive substantial high-survival rounds before latching S=1.
Gate sweep seeding on the same startup boundary and expose cumulative copy
and promotion price inputs under GC diagnostics.
Record the lock map, diagnostic schema, sabotage cases, disk-blocked gates,
and the exact TN follow-up request with falsifiable predictions.
…SS1_MARKED window

The rebase onto main re-introduces two thread-locals the root-holder gate wants
a verdict for, and again moves a source the PASS1_MARKED pin covers.

`RATED_ROUNDS` (Cell<u64>) and `PROMOTE_LOCK_STREAK` (Cell<u8>) hold a rated
cohort count and a streak count; neither is ever an address or a NaN-boxed
value, so both are `not_a_gc_pointer`.

The PASS1_MARKED pin covers `gc/mod.rs`, which this branch touches. Re-audited
rather than repinned blind: the change is confined to `emit_gc_time_share_diag`,
which reads four relaxed atomic counters and widens one `eprintln!` -- it
allocates nothing, collects nothing and calls no JS. Its only caller is
`emit_incremental_liveness_diag`, gated on `telemetry::gc_diag_enabled()`
(default off), main-thread only, latched to emit once per process, and reached
only from `js_gc_release_current_thread_collection_side_allocations`, the
every-process-exit funnel. The added code therefore cannot run inside
`cycle.rs::run_to_completion`, and the mark-complete to sweep-entry window is
unchanged. Only that one hash moved; the other four pinned sources still match,
which is how the hash computation itself was checked before the pin was touched.
@proggeramlug
proggeramlug force-pushed the perf/tenuring-evidence-lock branch from 5c16a19 to 4b912d9 Compare September 28, 2026 13:34
@proggeramlug

Copy link
Copy Markdown
Contributor Author

Re-measured on current main: the mechanism still fires. The CPU delta is not reportable from this run.

Rebased onto main 295473cd5 and measured paired on a temporary 48-thread host. Both arms sit on that one pinned base; arm B is that base plus only this PR's four GC files (copying.rs, instruments.rs, mod.rs, tenuring.rs, +332/−66), applied clean. The two cc binaries differ by sha256. Four turns in one cc process per run (stream_multi2.py), arms alternating each round, PERRY_GC_DIAG=1 captured per run. Raw rows, per-run diags, the pinned base and the exact patch are preserved at perrymaster:/root/rig9831/n9949/.

The mechanism is confirmed, and it is unambiguous

per 4 turns, median arm A (control) arm B (this PR)
copied_mb 0.00 11.50
promoted_mb 10.70 8.60
minors 5 5

Per-turn copying is [0, 0, 0, 0] for the control in every round and [6.3, 1.9, 1.5, 1.8] for arm B in every round. That is exactly the claim: the control latches the survival threshold to in-place promotion and never evacuates; the evidence-gated lock keeps copying minors, and as a direct consequence promotes about 20% less. This is the part that could plausibly have rotted across ~900 commits of drift, and it has not.

These two signals are load-independent, which matters on this host (below).

Correction to an earlier reading: the "+8% full collections" is an artefact, not an effect

I previously reported arm B at 346 full collections against arm A's 321. That was my aggregation error — unbalanced n across rounds. Full collections are wall-time-proportional at the same rate in both arms:

row fulls wall ms fulls/sec
armA r1 321 61 020 5.26
armA r2 323 61 176 5.28
armB r1 320 63 240 5.06
armB r2 320 60 708 5.27
armB r3 371 73 557 5.04
armB r4 372 72 431 5.14

The trigger mix is also the same: kind=due dominates in both (660 vs 653 across two rounds), with ArenaBytes 18 vs 19 and OldReclaim 4 vs 5. due fires at site=budgeted_start on a schedule, so a run that takes longer in wall time accumulates more of them — and arm B's rounds 3 and 4 simply happened to run under the heaviest foreign load. No extra full collections are attributable to this change, so there is no hidden cost there to eat the win.

Why no CPU number

The host was shared and heavily oversubscribed: load ran 84 → 141 on 48 threads and rose monotonically through the run. The harness's own contamination stamp makes the scale concrete — foreign_cpu_s per turn was ~600–800 s in most rows but ~1296 s in arm B's rounds 3 and 4, i.e. the rounds that produced its slowest times. A naive read gives arm B +11% (68.7 s vs 61.9 s median); that number is confounded with load and arm order and should not be used. One arm A row was also short (2 of 4 turns) and was discarded, leaving n unbalanced at 3 vs 4.

So September's −7.7% four-turn figure is neither confirmed nor refuted here. Instruction-count rows (perf stat -e instructions,cycles, whole 4-turn process, balanced n, alternating rounds) are running now, since instruction counts are insensitive to contention in a way CPU seconds are not. Peak and settled RSS per arm will come with them — that is where this change should show a win, given it promotes ~20% less.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Instruction counts: more expensive in all seven paired rounds, and the cost is bimodal

Whole-process counters now, on the same rig. Paired rounds only — a round enters the table only if both arms produced a complete four-turn row, so n is balanced by construction — arm order alternating, perf stat -e instructions,cycles,task-clock wrapping the entire four-turn process. Both arms on pinned base 295473cd5; diff -rq over crates/ shows exactly this PR's four files differing (copying.rs 4, instruments.rs 31, mod.rs 4, tenuring.rs 359 changed lines), the two cc binaries differ by sha256, and the cc compile line is byte-identical. Every row landed on its first attempt.

Read the instruction column, not the cycle column: arm A's instruction count varies by 1.6% across rounds while its cycle count varies by 51.8% under the host's load. That is precisely why the cycle and wall-clock numbers were not reportable before, and why these are.

Set 1 — five paired rounds

round A instr B instr Δ instr A fulls B fulls load A/B
1 530.79 G 616.30 G +16.11% 334 385 33 / 30
2 526.77 G 621.69 G +18.02% 334 386 25 / 25
3 522.45 G 624.87 G +19.60% 334 385 29 / 43
4 529.65 G 537.04 G +1.39% 332 334 93 / 92
5 526.89 G 538.79 G +2.26% 334 333 105 / 139

Paired delta: mean +11.48%, median +16.11%, range +1.39% to +19.60%. Arm B is never cheaper.

Set 2 — two more pairs with a graceful exit, load-matched

Ctrl-C twice instead of SIGKILL, so the process-exit funnel runs and prints [gc-time]. Both arms here sat at essentially the same load, so CPU is comparable within a pair:

round A instr B instr Δ instr Δ CPU A fulls B fulls load A/B GC share A/B
1 528.85 G 625.04 G +18.19% +23.02% 334 386 139 / 139 65.4% / 68.6%
2 534.05 G 628.21 G +17.63% +17.72% 334 384 130 / 144 66.9% / 68.4%

The delta is not a spread around one value — it is two modes, and the collection count picks it

Across all seven rounds:

  • arm A fulls: 332, 334, 334, 334, 334, 334, 334 — one mode.
  • arm B fulls: 333, 334, 384, 385, 385, 386, 386 — two, cleanly separated, nothing in between.

The extra collections are all kind=due at site=budgeted_start (381–382 against arm A's 323–328); ArenaBytes and OldReclaim are unchanged. Pricing arm B's own two modes against each other — 83.77 G instructions over 51 extra collections — gives 1.63 G instructions per full collection, which accounts for the entire +16–20%.

So the split is: when arm B takes the same number of full collections as arm A it costs +1.4% and +2.3% instructions; when it takes ~51 more it costs +16% to +20%, and +18% to +23% CPU. The couple of percent at equal collection count is the copying work itself — 11.5 MB evacuated per four turns — plus the added instruments.

What sets the mode is not load, and not arm order

I first suspected load, because in set 1 the high mode fell in the three quiet rounds. Set 2 refutes that: at load 139/139 and 130/144, arm B is firmly in the high mode, while in set 1 it was in the low mode at loads 92 and 139. Arm order does not explain it either — round 2 ran arm B first and went high, round 4 ran arm B first and went low. Arm A never shows a second mode in any of the seven rounds.

The practical consequence is that this change cannot be evaluated from a small number of runs. Two honest four-turn measurements of the same two binaries can land 16 points apart depending on which mode arm B picks, which is a plausible account of how the original −7.7% and my +11.5% mean can both be real measurements of the same code. The mode needs identifying before an accept/reject decision means anything.

Corrections to my previous comment

Two things in it were wrong, both mine:

  1. I said full-collection counts are wall-time-proportional at ~5.1–5.3 per second in both arms. They are not. Arm A's count is essentially constant per unit of work (332–334) across wall times from 36.9 s to 59.3 s.
  2. On that wrong model I dismissed the "+8% fulls" as an artefact of unbalanced n. It is not an artefact — with seven balanced pairs, arm B really does take ~51 more full collections in five of seven rounds.

Memory, and the mechanism

median over the 5 paired rounds arm A arm B Δ
peak RSS 443 MB 448 MB +1.1%
settled RSS (after the inter-turn gap) 439 MB 433 MB −1.4%
promoted 10.50 MB 8.60 MB −18.1%
copied 0.00 MB 11.50 MB —

The mechanism is intact and reproduces in every single round: the control evacuates nothing, arm B copies ~11.5 MB per four turns and promotes ~18% less, and settled RSS is ~1.4% lower. That memory win is real but small at this workload, and it does not pay for a 16–20% instruction cost in the rounds where the extra collections appear.

For context on why collection count dominates everything here: [gc-time] wall_us=80475180 step_us=48255771 share_permille=654 — GC is 65–67% of wall time on this workload in the control and 68–69% with the change.

One rig note for anyone reproducing

Wrapping the stream harness in perf stat silently invalidates its per-turn instrumentation: perf becomes the direct child, so every /proc/<pid> read measures perf rather than the program — turn_cpu_s reads 0.0 and RSS reads 11 MB while all four turns in fact ran. And a teardown that SIGKILLs that child loses every counter, because perf writes its -o file only when the program it launched goes away. I measured through a patched copy that follows the wrapper down to the real pid and kills the program first; the rows above carry two witness fields (measure_pid_is_descendant, and the parsed counters) so a row that lost either cannot enter the table.

Raw rows, per-run diags, the pinned base, the patch and the patched harness are preserved and I can hand them over.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

An eighth pair finished just after I posted, and it refines the CPU claim above, so here it is rather than left out.

round A instr B instr Δ instr Δ CPU A fulls B fulls load A/B GC share A/B
set 2, round 3 532.16 G 536.06 G +0.73% −0.47% 334 334 140 / 106 67.0% / 66.5%

Arm B is in the low mode here, and in that mode it is at parity: +0.73% instructions and −0.47% CPU, i.e. marginally faster. So the "+18% to +23% CPU" from the two load-matched pairs above is the cost of the high mode specifically, not of the change in general — the corrected reading is:

  • low mode (3 of 8 rounds, arm B at 333–334 collections): +0.7% to +2.3% instructions, −0.5% to +1.5% CPU — parity, with promotion down ~18% and settled RSS down ~1.4%. On these rounds the change is a small, genuine win on memory for no measurable time cost.
  • high mode (5 of 8 rounds, arm B at 384–386 collections): +16% to +20% instructions, +18% to +23% CPU.

Final census over all eight pairs — arm B fulls 333, 334, 334, 384, 385, 385, 386, 386, arm A fulls 332, 334, 334, 334, 334, 334, 334, 334. Load still does not select the mode: arm B went high at 139 and low at 140.

That makes the question sharper rather than softer. The change is either free-with-a-memory-win or a 20% regression depending on a mode nobody has yet identified, and it flips between them on identical binaries and identical input. Worth finding the switch before deciding on the PR — if the low mode can be made the only mode, the original case for this change stands up.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Bimodality root-caused: it is not a latch and not the tenuring lock — the control's collection count is bounded, this PR's tracks churn, and the workload straddles a step

Investigated on the 16 runs already captured (8 per arm, same pinned base 295473cd5, same binaries), so no new measurement was needed to get here.

It is not the tenuring lock, and there is a clean counterexample

I first went after the lock and the nursery ladder, because the first divergence between a low-mode and a high-mode run looks like one: the first five [gc-tenuring] lines are byte-identical, then the low run emits an extra sweep-seed evaluation and the nursery bands split (high 32145145, low 32614907, off a mean_surviving_object_bytes reading of 69 against 70).

That is a red herring. Across all eight armB runs none of the lock-side variables separate the modes:

low mode (3 runs) high mode (5 runs)
sweep-seed present 1, 1, 1 1, 1, 0, 1, 0
rounds_rated at survivals 2 -> 3 7, 7, 8 7, 8, 6, 6, 8
first three mean readings identical in all 8 runs identical in all 8 runs
1x -> 2x scale step eden_live 4118656 … 4120120 4119320 … 4120184

g_armB_r2_a1 settles it: its nursery band ladder, its sweep-seed line and its mean readings are identical to the three low-mode runs, and it took 384 collections. The lock and the nursery pacing do not pick the mode. [gc-survival] is exactly 405 in all 16 runs, both arms, and armB's copied_bytes (≈35 MB) and promoted_bytes (≈50 MB) totals are the same in both modes — the young-generation side is not where the modes differ. The arena ceiling does not separate them either (low 83/84/86 MB, high 80/82/83/86/89 MB), nor does nursery_cap (64.0 MB in all 16).

What does separate them: total bytes through the collector

Summing freed_bytes over every [gc] blocks line as a churn proxy:

fulls spread freed spread MB per full corr(fulls, freed)
armA (control) 332–334 0.6% 5535–6353 MB 14.8% 16.57–19.02 −0.55
armB (#9949) 333–386 15.9% 7917–9995 MB 26.3% 23.11–26.03 +0.88

The control is bounded-cost. Its collection count is 334 in seven of eight runs (332 once) while its churn varies 14.8% — it absorbs the entire variation in how much each collection frees (16.6 → 19.0 MB). This PR removes that property: the count follows churn at r = +0.88.

And the response is a step, not a slope:

armB freed (MB) 7917 7925 8504 8898 9349 9419 9938 9995
full collections 333 334 334 385 386 385 386 384

Two discrete operating points. Within each, churn moves 7% and 12% while the count barely moves; across the gap between 8504 and 8898 MB — 4.6% more churn — the count jumps 15.3%. Per-collection work drops from 25.5 to 23.1 MB across the step, i.e. above it the collector does more, smaller collections.

So the mechanism is

  1. The change raises total bytes through the collector by 35–58% against the control at the same workload (7917–9995 MB against 5535–6353 MB). Young-generation survivor traffic is identical, so this is old-generation/block-sweep traffic — consistent with copied survivors staying in the young generation and being re-swept on later cycles instead of being promoted out once.
  2. cc's own garbage volume varies ±12% run to run for reasons unrelated to the change (the control shows the same 14.8% spread).
  3. That combination lands the workload on both sides of a discrete step in the full-collection pacer, at roughly 8.7 GB cumulative. The control never reaches the step because its churn is 35–58% lower and its count is pinned regardless.
  4. Above the step: +51 collections × ~1.63 G instructions = +83 G instructions, which is the +16–20% I measured.

This is why the mode looked load- and order-independent: churn is neither. It is not a one-time latch — nothing needs to flip once per process.

Implication for the PR, and the fix direction

The −18% promotion and −1.4% settled RSS are real and reproduce in every round, so the goal Ralph set — keep the promotion win, drop the extra collections — looks reachable, and the two candidate directions are different in kind:

  • Bound the re-copying. Cap how many times a survivor may be evacuated before it is promoted in place. That attacks step 1 directly: it should cut the 35–58% churn excess, which is the thing that pushes the workload over the step at all, while keeping most of the promotion reduction. This is the one I would try first, because it removes the cause rather than the symptom.
  • Restore the bounded-cost pacer. Make the full-collection count insensitive to churn the way the control's is (332–334 across a 14.8% churn range). This would remove the bimodality even if the churn excess stays, but it leaves the extra collector traffic in place.

Acceptance for either should be the same instrument used here, since it is what made the effect legible: report corr(fulls, freed) and the fulls spread over ≥6 paired rounds, not a single CPU delta. A fix is working when armB's fulls spread comes back under ~1% and its total freed drops toward the control's, with promotion still ~18% down.

One honest limit: freed_bytes on [gc] blocks is a block-sweep counter and I am using it as a churn proxy. I have not separately verified that it does not double-count a copied survivor swept in more than one cycle — if it does, the 35–58% figure overstates the true byte traffic, though it would not change the ranking or the step, both of which show up in collection counts too.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Correction, and the actual cause: the PR's evidence gate can never fire on this workload, so the threshold lands on 3 instead of 1 — and that doubles the nursery

My churn explanation in the comment above was wrong and I want it corrected before anyone acts on it. I said the extra collector traffic was "consistent with copied survivors being re-swept". It is not: copying moves only ~35 MB over a whole run, while the freed difference is 2.4–3.6 GB. Two orders of magnitude off. The step and the correlation stand; the mechanism I offered for them does not.

Here is what the data actually says. One line, present in all 16 runs without exception:

tenuring threshold transition
armA (control), all 8 runs [gc-tenuring] survivals 2 -> 1
armB (#9949), all 8 runs [gc-tenuring] survivals 2 -> 3

The control settles on 1 — promote on first copy, which is why its copied_mb is exactly 0.00 in every round. This PR settles on 3 — an object must survive three minors before it is promoted. Not once in eight runs does it reach 1.

That single policy difference explains every number in this thread. Comparing the pacer state at budgeted_start (medians over a run, control against PR):

armA (control) armB (#9949) ratio
nursery_cap 33.5 MB 66.2 MB 2.0x
from_space (live young data at the trigger) 12.9 MB 28.4 MB 2.2x
next_base (where the next collection is armed) 90.2 MB 81.8 MB armed earlier
arena_total 54.5 MB 67.1 MB 1.2x
old_in_use 32.7 MB 31.2 MB 0.95x — the promotion win

Holding three cohorts instead of one needs twice the nursery, so every full collection sweeps 2.2x more young data — and freed_bytes per collection is 25.5 MB against the control's 17.5 MB, a ratio of 1.46 that tracks the from_space ratio, not the 35 MB of copying. The nursery is larger and next_base is lower, so collections are armed sooner against a bigger young generation. That is where the 35–58% extra byte traffic comes from, and it is also why the collection count becomes churn-sensitive (r = +0.88) while the control's stays pinned at 332–334.

Why it lands on 3 instead of 1

From the PR's own documentation of the lock: the first two rated cohorts are excluded as startup, then three consecutive cohorts must each be ≥90% surviving to lock the threshold to 1. Every survivals line in every armB run reports the measured rate:

survival_permille = 218, 225, 227, 320, 375, 377, 378

That is 22% to 38% survival — the ≥90% bar (900 permille) is never met, so PROMOTE_LOCK_STREAK reads streak=0 in every run, at rounds_rated 6, 7 and 8 alike. The lock cannot fire on this workload, and with it unable to fire the occupancy path takes the threshold the other way, to 3. cc gets the worst of both: no lock to 1, and a doubled nursery.

So this is not a latch that flips once, and not noise amplification. It is a gate whose entry condition is unreachable for a workload with ~30% survivor survival, with a fallback that moves in the opposite direction to the gate's intent. The bimodality is then just the churn step from the earlier comment acting on a 2x-larger young generation.

Revised fix direction

"Bound the re-copying" was aimed at the wrong thing, since re-copying is 35 MB. The lever is the threshold itself:

  1. Make the gate reachable, or make its failure land on 1 rather than 3. A workload at 30% survival is one where copying repeatedly is especially poor value — 70% of each cohort dies anyway, so the third cohort's retention buys little and costs a doubled nursery. The control's choice of 1 is the right answer here and the PR's gate should be able to arrive at it. Either lower the ≥90% bar to something a real workload reaches, or invert the default so that an unproven cohort promotes rather than accumulates.
  2. Cap the nursery independently of the threshold. If threshold 3 is genuinely wanted for high-survival workloads, it should not be free to double the nursery, since that is what turns a promotion win into a 43%-per-collection sweep cost.

Option 1 is the one that matches the measurements: it should keep most of the −18% promotion reduction (promotion is already down at threshold 3 mainly because 0.3³ of a cohort reaches old age) while returning from_space and the collection count to the control's range. The acceptance instrument from the previous comment still applies — fulls spread back under ~1%, corr(fulls, freed) near zero, total freed near the control's, promotion still down — plus, now, nursery_cap and from_space medians back near 33.5 MB and 12.9 MB, and the survivals line reading 2 -> 1.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Correction: with collection counts matched, this PR costs +1.55% instructions, not +16–20% — and my proposed fix is refuted

I ran the threshold matrix Ralph asked for, and it inverts my two comments above. Posting it immediately because those comments would sink this PR on wrong grounds.

Method. Five cells from the two existing binaries, no rebuild, using the PERRY_GC_TENURING_SURVIVALS knob that both arms already carry: control adaptive, control pinned to 3, PR adaptive, PR pinned to 1, PR pinned to 2. Four rounds, all five cells per round or the round is dropped (4 complete rounds, n=4 each). The host was quiet this time — loads 14–110, against 84–141 for everything I posted earlier.

cell threshold instructions fulls freed nursery med from_space copied promoted settled RSS
A_auto control → 1 529.62 G 334 5872 MB 32.0 MB 12.8 MB 0.00 10.60 438
A_S3 control pinned 3 544.01 G 331 9705 MB 63.1 MB 32.4 MB 16.10 12.60 433
B_auto this PR → 3 537.85 G 334 7656 MB 62.7 MB 26.2 MB 12.20 8.60 431
B_S1 this PR pinned 1 559.38 G 332 8722 MB 64.0 MB 26.2 MB 0.00 8.20 430
B_S2 this PR pinned 2 543.57 G 334 7980 MB 63.1 MB 26.4 MB 11.30 8.95 426

Against the control: B_auto is +1.55% instructions, −18.87% promoted, −1.71% settled RSS. Collection counts are 331–334 in every cell, so this is a like-for-like comparison, and it is a materially better result for the PR than anything I posted earlier.

What I got wrong

  1. "The entire cost difference is the tenuring threshold (1 vs 3)." Refuted. I based that on a single A_S3 round showing 382 collections; at n=4 A_S3 sits at 331. Forcing the control to threshold 3 costs +2.72% instructions, not +18%.
  2. "Make the gate reachable so the lock lands on 1." Refuted, and backwards. Pinning this PR to threshold 1 is the most expensive cell in the matrix (+5.62%), and threshold 2 (+2.63%) is also worse than the PR's own adaptive choice of 3 (+1.55%). The PR's adaptive threshold is the cheapest armB configuration available. There is nothing to fix in the lock's entry condition on cost grounds.
  3. The nursery doubling is not caused by the threshold. B_S1 runs at threshold 1 and still holds a 64.0 MB nursery, where the control at threshold 1 holds 32.0 MB. So the 2x nursery comes from the PR's other changes, not from holding three cohorts — although in the control the threshold does drive it (A_S3 → 63.1 MB). That decoupling is worth a maintainer's eye, but it is not a cost regression: B_auto carries the doubled nursery at +1.55%.
  4. Also note A_S3 promotes 19% more than the control, not less — so threshold 3 by itself does not buy the promotion reduction. The PR's −18.9% is not simply "threshold 3"; it survives at threshold 1 too (B_S1, −22.6%), which again points at the nursery rather than the copying policy.

What survives, and is the real open question

The +16–20% I measured earlier was real, but it is an intermittent excursion, not the steady state. Collection counts across every run I have:

binary runs collection counts excursions
armA (control) 16 330–335, every run 0 of 16
armB (this PR) 20 331–335, except 382, 384, 385, 385, 385, 386, 386 7 of 20

The excursion is worth +16% instructions when it happens, it is specific to the armB binary — never once in 16 control runs, and it is not selected by the threshold (it hit the pinned cells in this matrix and the adaptive cell in the earlier set) nor cleanly by load, though its rate was higher in the loaded set (5 of 8) than the quiet one (2 of 12). Taking the observed rate at face value, expected cost is roughly 0.35 × 16% + 0.65 × 1.5% ≈ 6%, with a bimodal distribution rather than a spread.

So the decision Ralph asked me to inform is narrower than I made it look: the steady state is a good trade (+1.5% compute for −18.9% promotion and −1.7% settled RSS), and the open risk is a ~35%-frequency excursion that doubles the collection count and that the control never shows. That excursion is what still needs root-causing, and my threshold hypothesis for it is dead. It is also why the same binaries gave −7.7% and +11.5% in different hands.

Method note for anyone re-running: report the collection count next to every timing number. Every wrong conclusion in this thread, mine included, came from comparing two runs that had silently landed on different collection counts.

Raw rows, per-run diags and the matrix script are preserved with the rest at perrymaster:/root/rig9831/n9949/.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant