Skip to content

docs(gc): RFC — deferred collection (L2b) - #11528

Draft
proggeramlug wants to merge 2 commits into
mainfrom
docs/rfc-deferred-collection
Draft

proggeramlug wants to merge 2 commits into
mainfrom
docs/rfc-deferred-collection

Conversation

@proggeramlug

Copy link
Copy Markdown
Contributor

Design only, for owner review. No code changes. Adds docs/src/internals/rfc-deferred-collection.md and links it from SUMMARY.md. This is a docs-only PR, so no changelog fragment is needed (the fragment rule applies to crates/ changes).

Proposal (lever L2b)

  • Allocation never begins a precise collection. When the nursery is full it takes another block and arms the existing poll word.
  • Precise collections begin only at declared polls. These are loop back-edges; an indirect-entry prologue on closures, methods and callbacks; one entry poll per recursive SCC; runtime_poll() inside garbage-heavy helpers; gc(); and the pump boundaries.
  • A call is a safepoint only if its callee may reach a poll. A static call-graph checker over the linked runtime archives computes this and emits a committed, generated table with four classes: Leaf, AllocOnly, ThrowOnly, Reenters. CI fails if a leaf-classified symbol reaches a JS-invoking symbol (the fix(gc): js_array_length is classified AllocNoReentry but its Proxy arm runs user JS (get trap + number coercion) #11522 class). A runtime unmapped-frame verifier, run together with the seeded schedule, is the second layer.
  • Throw cut and fast/slow IC splits.
  • Memory bound. The existing conservative, non-moving slack valve stays as the only collection that can start at an allocation, and it is gated to fire zero times on the corpus. Heavy runtime helpers are classified as output-dominated, garbage-producing or JS-calling, and each kind is bounded separately.

Key finding

The nursery half of this already ships. What is new is making it an invariant: the OldReclaim alloc-point arm, budgeted root-scan and remark in mutator assists, and root-lock flushes (#11523) move to polls. Codegen then exploits the invariant.

Measured (claude-code bundle)

Migration

S0 #11500 → S1 soundness fixes + checker + L2 table (sound today) → S2 linear liveness estimator → S3 runtime invariant + polls → S4 L2b → S5 throw cut → S6 fast/slow splits.

Deviation from the suggested order: the throw cut has to wait for S3. The throw helpers allocate, and today that allocation can reach a precise root scan at an unmapped frame.

Open questions

Ten, in §8 of the RFC. The main ones:

  • Accept deferring OldReclaim (the Memory not free in benchmark #5476 guarantee)?
  • Poll placement: indirect-entry plus SCC polls, or polls at every function entry (12.18M vs 7.39M relocations)?
  • Table authority and which targets the checker runs on.
  • The valve slack.
  • The RSS acceptance bar.

Design only, for owner review. Proposes that allocation never begins a
precise collection, that precise collections begin only at declared polls,
and that a call is a safepoint only if its callee may reach a poll, computed
by a call-graph checker over the linked runtime archives. Includes the
claude-code bundle relocation census, the memory-bound design, interactions,
a six-step migration plan and open questions.
@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Owner decisions on §8 (2026-09-27)

  1. Memory not free in benchmark #5476 / OldReclaim: not moved for now. OldReclaim stays at the allocation point (non-moving, so sound under D1). Revisit only once measurement shows no RSS cost.
  2. Poll placement: as proposed (indirect-entry prologue polls + one poll per recursive SCC + loop back-edge polls). Measured 7.39 M vs 12.18 M relocations for all-function entry polls.
  3. Straight-line polls: measure first, with a hard rule: if a module-init body or large entry chunk exceeds the RSS bar (item 8), codegen adds statement-boundary polls to those bodies only.
  4. Table authority: a generated, committed table plus the call-graph checker in CI, run on Linux x86-64, macOS aarch64, and Windows via cargo xwin.
  5. Valve: keep 64 MiB slack. A valve firing on the gap suite or the ratchet probes is a hard pr-gate failure, with a counter asserting the check actually ran.
  6. Budgeted-cycle latency: accepted. Measure pause and latency on a server fixture.
  7. Fatal-throw carve-out: accepted, with a test (uncaught throw + an allocation-heavy exit listener).
  8. RSS bar for S3: ≤ 2 % peak RSS on any single probe AND no regression of the median across the ratchet corpus and the cc bundle. Compute must not regress either.
  9. Census as CI: a ratchet (relocation and safepoint totals may only go down). A small representative corpus in regular CI, and the full cc census nightly through /root/perry-heavy.sh.
  10. GC_UNSAFE_ZONES: out of scope, but add a diagnostic counter for growth inside unsafe zones.

Proceeding: S0 = #11500, S1 = soundness fixes #11522/#11523 + the call-graph checker + the generated table (census tooling moves into scripts/).

…he deferred-collection RFC

Adds the owner's 2026-09-27 decisions as section 8, resolving the ten open
questions (OldReclaim stays at the allocation point, three-target checker,
valve as a hard pr-gate failure, RSS bar, census ratchet). Adds the
fast-path-split census (13.89M -> 3.23M relocations when stacked on L2b plus
the throw cut) with its caveats. Reorders the migration plan to S0 #11500,
S1 table and checker, S2 fast/slow IC splits (argued sound on today's
runtime, helper by helper), S3 rematerialization, then the estimator, the
runtime invariant, L2b and the throw cut.
@proggeramlug

Copy link
Copy Markdown
Contributor Author

Owner decision on the S2 static-reduction question (2026-09-27)

Yes: add a step after S4. S2 (fast/slow IC splits) ships as a runtime win: the hot path is gc-leaf, and the slow call stays a statepoint. A new step S4b follows the liveness estimator: at each out-of-line slow call site, the live GC values are rooted in plain stack slots recorded in the GC map and updated in place by the collector. That means no gc.relocate SSA fan-out, and no runtime calls (explicitly not the shadow-frame root stack). The aim is the static reduction (relocations, IR size, RS4GC time) that the census's −77% assumed. Measure the static and runtime effect separately.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

S4 landed as a PR: the design changed in three places

S4 (the linear liveness estimator) is PR #11624. The RFC's S4 row says "backward liveness of root allocas across the classified safepoints of the pre-RS4GC IR, O(instructions + roots)". The implementation differs in three ways, each forced by measurement. I am recording them here rather than editing the RFC branch.

  1. It runs at LLVM level, on RS4GC's exact input, not on root allocas. The backend now runs the statepoint pipeline in two halves: always-inline,function(mem2reg,sccp), then the count, then rewrite-statepoints-for-gc. The halves print the same module as the one-shot pipeline, and a test pins that. The count is over SSA values after mem2reg/sccp, so it needs no model of Perry's lowering. S0/S1/S2 leaf marks are already attributes on the calls, and S3's rematerialized roots are already fresh loads. It is exact, not a bound: on the gap suite (16,024 functions) and on the first 48 of 128 claude-code units (27,773 functions, including __25747 and GW7), the prediction equals the gc.relocate count RS4GC emits in every function. Exactness needed three RS4GC behaviours the census plugin does not model:
    • a call's GC-pointer arguments are live across it (findLiveSetAtInst walks the call itself);
    • RS4GC's CFG cleanup, the single-use icmp sink, and nounwind/noreturn invokes;
    • findBasePointer. A phi that merges a NaN-box tag constant (inttoptr (i64 0x7FFC…)) with a heap pointer gets a fresh .base phi, so it costs two relocations per crossing. Leaving this out undercounted 175 of 7,692 gap-suite functions.
  2. The cost is O(IR + Σ live blocks), not O(instructions + roots). It is a per-value backward walk (the census algorithm) with difference arrays for the per-safepoint counts. That is linear in the IR plus the liveness it reports, and never more than RS4GC's own liveness. Measured on the bundle: 128 ms on the slowest function (IoK, 148 k instructions), 59 ms on __25747, and 4.8 s in total over 48 units, where RS4GC took 338 s.
  3. The HIR estimate is gone, not replaced. Every function lowers natively first. A function over the budget takes the existing codegen: reconcile the root-spill threshold (#8623, 32M) with the RS4GC instruction budget (#8586, 1.57M) — budget-gap functions get refused instead of spilled #8679 retry and is re-lowered onto a shadow frame before RS4GC runs. The budget is the post-RS4GC instruction budget (1.5 Mi), because every relocation is one instruction of the rewritten body.

Two findings that bear on S4b and later steps:

  • RS4GC time does not track relocations. On the bundle, the invoke-heavy units cost more RS4GC time than __25747's unit, which has seven times the relocations: __84092's unit takes 51 s for 65 k relocations, against 27 s for 501 k. On a synthetic chain of if/else diamonds, 500 values × 250 safepoints (0.25 M relocations) take 540 s in RS4GC, while the same count in straight-line code takes about 4 s. Neither budget catches that shape. S4b's stack-slot rooting removes the join phis that drive it. The per-safepoint counts are exposed for S4b as gc_liveness::FunctionLiveness::safepoints.
  • The former spill list is gone. __87158, __84092 and __85198 now stay on statepoints, with 42.9 k, 25.7 k and 17.4 k real relocations.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

S5 (runtime invariant + new polls): implementation notes and design changes

S5 is split into two PRs: #11630 (runtime invariant, safety net, gates) and #11631 (entry polls, stacked on #11630). The design differs from §2 in the ways below.

  1. D2 is enforced at the allocation point, not through a "declared safepoint" flag. An allocation point is the dynamic extent of gc_check_trigger, which every slow path funnels into.

    • A synchronous collection started there must have requested the conservative scan. The check reads the request, not the resulting scan decision, because the unit-test isolation guards and PERRY_CONSERVATIVE_STACK_SCAN=off override the decision on purpose. It panics in every build.
    • A budgeted cycle whose next step reads frame roots (RootScan, FinalRootRemark) is parked, the poll is armed, and the next declared poll serves the phase.
    • The runtime half of PERRY_GC_SAFEPOINT_ONLY (heal/strict) is deleted. The codegen half, the AllocNoReentry research switch, is left for S6 to replace with the table.
  2. New: a parked-cycle valve (owner question below). An active budgeted cycle blocks every other collection, including the nursery valve. So a program that parks a cycle and then allocates without reaching a poll would grow without bound. After the nursery slack (64 MiB, budget-scaled), the phase is served at the allocation point, precisely. The valve is counted and gated to zero on the gap suite and the ratchet probes.

    • It is sound today only because every allocating call is still a statepoint.
    • It becomes unsound at S6.
    • Budgeted cycles are classifier-mode, so they cannot take the conservative scan themselves.
  3. Indirect-entry polls are in-body, not a stub, for closures, methods (instance, static, constructors, accessors) and value wrappers. Two findings forced this: a stub's arguments are not roots, so a moving poll there would leave them stale, and a dozen runtime tables are keyed on the body's address.

    • Closures and methods poll at the point where their parameters are rooted.
    • The existing __perry_wrap_* forwarders are the stub for top-level functions: they spill their arguments to a buffer that js_gc_entry_safepoint_args roots.
    • Scaffolds are emitted during lowering and kept only in functions that the module leaf analysis (with entry polls ignored) already proves collecting, so no call changes classification in S5.
    • Consequence for S6: a direct call to an allocating method reaches that method's poll, so it cannot become a leaf. If the census shows that matters, S6 needs a real stub for methods.
  4. Unmapped-frame verifier without a range table. Function membership comes from the unwinder's region start (_Unwind_GetRegionStart, _Unwind_Find_FDE on the x29 walk, RUNTIME_FUNCTION on Windows), looked up in the GC map's function list. Instrumented compiles also list zero-record statepoint functions, which v6 runtimes already skip, so the map format did not change.

    • Armed by PERRY_GC_VERIFY_FRAMES=1, a schedule seed, and debug_assertions.
    • Off in release: every unmatched frame costs an FDE lookup, and the listing costs map bytes.
  5. Proposed refinement of decision 2, needs an owner ruling. Decision 2 says one poll per recursive SCC. Taken literally, that gave benchmarks/suite/05_fibonacci +59.98% instructions: fib's numeric fast clone and its boxed fallback form an SCC that allocates only through the cold fallback, so every call paid a poll. gc: function-entry polls at indirect entries and recursive SCCs (RFC deferred collection S5, polls) #11631 polls an SCC only when the recursion itself allocates without a poll: a member with its own collecting edge, or a call out of the SCC to a function that may collect and has no entry poll. A callee that polls at entry bounds its own allocation per call. With that rule, fib is back to +0.00%, and a loop-free recursive allocator still gets its poll (−52.7% instructions, peak RSS −60 to −74% against main, where it waited for the 64 MiB valve). The coverage checker applies the same definition. If you want the literal rule instead, it is a one-line change.

  6. Poll-coverage checker is scripts/gc_poll_coverage_check.py over --trace llvm output, reporting per loop and per SCC, with a ratchet.

Owner question (blocks S6, not S5): before AllocOnly becomes leaf, the parked-cycle valve must either be proven unreachable (today it is gated to zero on the gap suite and the ratchet probes) or be replaced by an abort of the parked budgeted cycle followed by the conservative valve. Abort does not exist today: the parked cycle has enabled the barrier, allocate-black and the remembered-set snapshot, so abort is a real design item. Which do you want?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant