Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions changelog.d/11624-s4-exact-relocation-count.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
- **Codegen: the shadow-frame spill is decided on RS4GC's exact relocation count (RFC deferred collection, S4).** #8583 moved a function's GC roots to a shadow frame when `(root slots + call sites) × call sites` exceeded 32 M. On the claude-code bundle that estimate overshot the real `gc.relocate` count about 100×, so it spilled functions RS4GC handles in seconds (`__87158`: estimated 134.5 M, real 42.9 k).
- **The count.** `inprocess/gc_liveness.rs` runs between the two halves of the statepoint pipeline: `always-inline,function(mem2reg,sccp)` first, then the count, then `rewrite-statepoints-for-gc`. So it sees exactly RS4GC's input, with every `gc-leaf-function` mark (S0, S1, S2) already on the calls and every S3-rematerialized root already a fresh load. It models RS4GC's liveness (a call's GC-pointer arguments are live across it), its CFG cleanup (`noreturn` cuts, `nounwind` invokes becoming calls, constant branches, phi folding), the single-use `icmp` sink, both relocation edges of an `invoke`, `TargetLibraryInfo` leaf calls, and `findBasePointer` for phis and selects. A phi that merges a NaN-box tag constant with a heap pointer gets a fresh `.base` phi, so it costs two relocations per crossing. The cost is linear in the IR plus the liveness it reports.
- **Exact against RS4GC.** Under `PERRY_CODEGEN_UNIT_TIMINGS` the backend counts the `gc.relocate`s RS4GC really emitted and prints one audit line per function. On the gap suite and on the whole claude-code bundle, prediction and RS4GC agree on every function (see the PR for the numbers).
- **The decision.** The HIR-level estimate (`maybe_spill_roots_to_shadow_frame`) and the constructed-IR `(allocas + sites) × sites` preflight are gone. A function over the budget gets the existing retry, re-lowered onto a shadow frame before RS4GC runs. Its log line now prints the real count, the statepoints (and how many are invokes) and the largest live set. The budget, `PERRY_ROOT_SPILL_RELOCATIONS` (default 1.5 Mi), is now the post-RS4GC instruction budget. Every relocation is one instruction of the rewritten body, so a larger count is over that backstop by construction.
- **The fast-emit cliff (found in validation).** Relocations aren't the only thing that grows the rewritten body past a cliff: `default_fast_emit_max_instrs` (600 k on x86-64, 100 k elsewhere; `PERRY_LL_FAST_EMIT_MAX_INSTRS`) is the point where LLVM's optimized machine pipeline gets bounded to an O0 fallback for instruction selection and register allocation (#10586). The relocation cap alone missed it: `__25747` sits at 0.4 M relocations, comfortably under the 1.5 Mi cap, but its rewritten body crossed 600 k instructions and fell into O0, growing its `.text` ~11×. The preflight now also predicts a function's post-RS4GC instruction count (pre-rewrite size plus a growth factor times the predicted relocations, calibrated against a 128-unit claude-code audit — see `POST_RS4GC_GROWTH_FACTOR`) and spills when *either* the relocation cap or this fast-emit prediction is exceeded. The prediction is deliberately conservative: it estimates the raw post-rewrite size, which is an upper bound on the further-optimized size the real fast-emit decision measures, so it can spill early but never miss a real cliff. On the bundle, two of the previously-cited functions (`__87158`, `__85198`) no longer spill under either check; `__25747` spills again (correctly, for the fast-emit reason this time) and `__84092` still doesn't (its relocation and predicted-instruction counts both stay well under budget).
- **For S4b.** `FunctionLiveness::safepoints` lists every call RS4GC will turn into a statepoint, with its per-edge relocation count. It is an internal API; nothing uses it yet.
6 changes: 0 additions & 6 deletions crates/perry-codegen/src/codegen/closure.rs
Original file line number Diff line number Diff line change
Expand Up @@ -616,12 +616,6 @@ pub(super) fn compile_closure(
let capture_root_slots =
u32::from(captures_this || enclosing_class.is_some() || entry_bound_this)
+ u32::from(captures_new_target);
crate::codegen::helpers::maybe_spill_roots_to_shadow_frame(
lf,
&llvm_name,
m.len() + capture_root_slots as usize,
body,
);
lf.enable_shadow_frame(m.len() as u32 + capture_root_slots);
m
} else {
Expand Down
6 changes: 0 additions & 6 deletions crates/perry-codegen/src/codegen/function.rs
Original file line number Diff line number Diff line change
Expand Up @@ -746,12 +746,6 @@ pub(super) fn compile_function(
);
// Root the entry `this` slot of a this-reading body.
let this_root_slots = usize::from(reads_this);
crate::codegen::helpers::maybe_spill_roots_to_shadow_frame(
lf,
&llvm_name,
m.len() + this_root_slots,
&f.body,
);
lf.enable_shadow_frame((m.len() + this_root_slots) as u32);
m
} else {
Expand Down
164 changes: 0 additions & 164 deletions crates/perry-codegen/src/codegen/helpers.rs
Original file line number Diff line number Diff line change
Expand Up @@ -502,166 +502,6 @@ pub(crate) fn inline_hot_small_max_call_sites() -> u32 {
})
}

/// #8583: statepoint relocation estimate above which a function keeps its GC
/// roots in a shadow frame instead of native statepoints.
///
/// `rewrite-statepoints-for-gc` adds one relocation per GC value live across
/// each safepoint, so the optimizer's post-rewrite cost scales with
/// `live_roots × safepoints`. Past a point that fan-out makes the `-Os`/`-O3`
/// middle-end super-linear and the compile does not finish (the Claude Code
/// bundle's 68 MB entry body measured 795 root slots × ~106k safepoints ≈ 8.4e7
/// and grew 439k → 6.5M instructions under RS4GC; without RS4GC the same unit
/// optimized at `-Os` in ~5s). Real functions sit orders of magnitude below
/// this: hundreds of call sites times tens of slots is ~1e4–1e5.
///
/// The default (#8620) is measured, not guessed. Synthetic entry functions with
/// a controlled `slots × safepoints` estimate were compiled at `-Os` with
/// spilling OFF (pure RS4GC fan-out) and the `@main` codegen unit timed:
///
/// | estimate | fan-out finish |
/// |---------:|---------------:|
/// | 8.0M | ~325 s |
/// | 16.0M | ~235 s |
/// | 32.0M | ~511 s (8.5m) |
/// | 40.0M | did not finish in 20 min |
/// | 48.0M | did not finish in 20 min |
///
/// The fan-out cliff sits between 32M and 40M, so the default is the largest
/// estimate whose fan-out still finished in bounded time. Below it fan-out is
/// the cheaper lowering — spilling a moderate function costs more than the
/// fan-out it avoids (an ~8M function spilled in 303 s vs 180 s fanned out,
/// #8620) — and above it fan-out risks not finishing and the shadow frame wins.
/// The former 4M default fired on ~8M functions that fan out fine in minutes.
/// The post-RS4GC instruction budget (#8586/#8679, inprocess.rs) backstops any
/// function this estimate misses: it re-lowers that function onto a precise
/// shadow frame and retries before LLVM's optimizer can hang, so raising the
/// estimate threshold is safe.
///
/// `PERRY_ROOT_SPILL_RELOCATIONS=<n>` overrides it; `0` disables spilling
/// (every function stays on native statepoints, the pre-#8583 behavior).
const DEFAULT_ROOT_SPILL_RELOCATIONS: usize = 32_000_000;

pub(crate) fn root_spill_relocation_threshold() -> usize {
std::env::var("PERRY_ROOT_SPILL_RELOCATIONS")
.ok()
.and_then(|v| v.trim().parse::<usize>().ok())
.unwrap_or(DEFAULT_ROOT_SPILL_RELOCATIONS)
}

/// The relocation estimate for a function with `slot_count` GC-root slots and
/// a body containing `safepoint_sites` call-like expressions. Saturating so a
/// pathological product cannot wrap.
/// The root population RS4GC actually relocates: named pointer locals plus
/// ~one live pointer temporary per call result (#8583). Production and the
/// threshold tests must agree on this composition — computing it in only one
/// of the two is how the endpoint tests silently stop guarding the real
/// formula.
pub(crate) fn spill_live_root_count(slot_count: usize, safepoint_sites: usize) -> usize {
slot_count.saturating_add(safepoint_sites)
}

pub(crate) fn root_relocation_estimate(slot_count: usize, safepoint_sites: usize) -> usize {
slot_count.saturating_mul(safepoint_sites)
}

#[cfg(test)]
mod root_spill_default_tests {
use super::{root_relocation_estimate, spill_live_root_count, DEFAULT_ROOT_SPILL_RELOCATIONS};

/// Exactly what `maybe_spill_roots_to_shadow_frame` computes, so these
/// endpoint tests track the production formula instead of a stale copy of
/// it (#8633 changed the composition; before this helper the tests still
/// asserted on the pre-#8633 `slot_count x sites`).
fn production_estimate(slot_count: usize, sites: usize) -> usize {
root_relocation_estimate(spill_live_root_count(slot_count, sites), sites)
}

/// #8620: the default is pinned to the measured RS4GC fan-out cliff — the
/// largest estimate whose fan-out finished in bounded time (32M finished in
/// ~8.5 min; 40M/48M did not finish in 20 min). Change it only with fresh
/// measurement.
#[test]
fn default_sits_at_the_measured_fan_out_cliff() {
assert_eq!(DEFAULT_ROOT_SPILL_RELOCATIONS, 32_000_000);
}

/// The moderate case the old 4M default wrongly spilled (#8620): ~8M
/// relocations (4000 root slots × ~2001 safepoints) fans out in minutes, so
/// under the new default it stays on native statepoints.
#[test]
fn moderate_fan_out_stays_on_statepoints() {
let est = production_estimate(4000, 2001);
assert_eq!(est, 12_008_001);
assert!(
est <= DEFAULT_ROOT_SPILL_RELOCATIONS,
"moderate estimate {est} must not exceed the default (would spill)",
);
}

/// The genuinely-catastrophic case (Claude Code `cli.js` `@main`,
/// ~795 slots × ~106k safepoints ≈ 8.4e7, never finishes at `-Os`) must
/// still spill under the new default.
#[test]
fn catastrophic_fan_out_still_spills() {
let est = production_estimate(795, 106_000);
assert!(
est > DEFAULT_ROOT_SPILL_RELOCATIONS,
"catastrophic estimate {est} must exceed the default (should spill)",
);
}
}

/// Decide whether `func` should spill its roots to the shadow frame, and if so
/// mark it (BEFORE its `enable_*_shadow_frame` call) and report it. Only
/// meaningful under native stack-map roots — the shadow frame is already the
/// lowering otherwise. Reporting is at default verbosity because #8421 requires
/// that a change to how a function is compiled is never silent; the message
/// states that the optimization level is unchanged.
pub(super) fn maybe_spill_roots_to_shadow_frame(
func: &mut crate::function::LlFunction,
fn_name: &str,
slot_count: usize,
body: &[perry_hir::Stmt],
) {
if !native_stack_roots_enabled() {
return;
}
let threshold = root_spill_relocation_threshold();
if threshold == 0 {
return;
}
let sites = crate::collectors::count_safepoint_sites(body);
// #8583 (unit-4 / `__33499` of the Claude Code bundle): `slot_count` is the
// shadow-slot map size — the count of *named* pointer-typed locals — but
// that is NOT the root population RS4GC relocates. A call-heavy minified
// closure produces one pointer-typed *temporary* per call result (the
// constructed IR carries ~one `alloca ptr addrspace(1)` per call), and each
// is live across the later safepoints; those temporaries dominate the true
// root count yet are invisible to `collect_pointer_typed_locals`. `__33499`
// measured ~20.3k named-and-anonymous pointer roots × ~20.3k safepoints, but
// its `slot_count` alone was ~100x smaller, so `slot_count × sites` fell
// under the threshold, the function stayed on statepoints, and RS4GC then
// fanned out for >3 h / ~30 GiB (never reaching the #8586 post-rewrite
// budget assertion, which only fires *after* the rewrite it never finishes).
// Count each safepoint as contributing ~one live pointer temporary. This is
// an over-approximation biased toward spilling — the intended direction (a
// false-positive shadow frame is cheap; a missed fan-out is not).
let live_roots = spill_live_root_count(slot_count, sites);
let estimate = root_relocation_estimate(live_roots, sites);
if estimate <= threshold {
return;
}
func.request_shadow_frame_spill();
eprintln!(
"perry: `{fn_name}` keeps its {live_roots} GC roots (incl. call-result temporaries) in a \
shadow frame instead of statepoints: an estimated {estimate} relocations ({live_roots} \
roots × {sites} safepoints) would make rewrite-statepoints-for-gc fan-out super-linear in \
the optimizer (> {threshold}). The function is still compiled at the requested \
optimization level; only its GC-root representation changes, and its roots stay \
precise (#8583). Override with PERRY_ROOT_SPILL_RELOCATIONS."
);
}

/// #10663: a function body with at least this many property stores outside
/// any loop outlines those stores' inline caches.
///
Expand Down Expand Up @@ -721,10 +561,6 @@ pub(super) fn enable_module_init_shadow_frame(

let shadow_slot_map =
crate::collectors::collect_pointer_typed_locals(&[], stmts, flat_const_ids);
// #8583: the module-entry body is the minified-bundle IIFE — the function
// that fans out catastrophically under RS4GC. Decide its root lowering
// before the frame is built.
maybe_spill_roots_to_shadow_frame(func, "main", shadow_slot_map.len(), stmts);
func.enable_post_init_shadow_frame(shadow_slot_map.len() as u32);
let shadow_slot_clears_after_stmt =
crate::collectors::collect_shadow_slot_clear_points(stmts, &shadow_slot_map);
Expand Down
6 changes: 0 additions & 6 deletions crates/perry-codegen/src/codegen/method.rs
Original file line number Diff line number Diff line change
Expand Up @@ -368,12 +368,6 @@ pub(super) fn compile_method(
),
&cross_module.scope_map,
);
crate::codegen::helpers::maybe_spill_roots_to_shadow_frame(
lf,
&llvm_name,
m.len() + 1,
method_body,
);
lf.enable_shadow_frame(m.len() as u32 + 1);
m
} else {
Expand Down
6 changes: 0 additions & 6 deletions crates/perry-codegen/src/codegen/method_static.rs
Original file line number Diff line number Diff line change
Expand Up @@ -70,12 +70,6 @@ pub(in crate::codegen) fn compile_static_method(
crate::collectors::collect_pointer_typed_locals(&f.params, &f.body, &flat_const_ids),
&cross_module.scope_map,
);
crate::codegen::helpers::maybe_spill_roots_to_shadow_frame(
lf,
&llvm_name,
m.len() + 1,
&f.body,
);
lf.enable_shadow_frame(m.len() as u32 + 1);
m
} else {
Expand Down
14 changes: 7 additions & 7 deletions crates/perry-codegen/src/collectors/safepoint_sites.rs
Original file line number Diff line number Diff line change
Expand Up @@ -7,15 +7,15 @@
//! body of the Claude Code bundle measured 795 root slots × ~106k safepoints
//! and grew 439k → 6.5M instructions under RS4GC, and a single `-Os` pass on
//! the result did not finish in practical time (#8583).
//! `codegen/helpers::maybe_spill_roots_to_shadow_frame` multiplies this count
//! by the function's root-slot count and, past a threshold, keeps that
//! function's roots in a shadow frame instead of statepoints.
//! Entry outlining (#8595, `codegen/entry_outline.rs`) uses this count to cut
//! a huge entry body into chunks before any of it is lowered. It is a
//! source-level proxy only: whether a function keeps its roots on statepoints
//! is decided on the exact relocation count of the lowered IR
//! (`inprocess::gc_liveness`, RFC deferred collection S4), which replaced the
//! `(slots + sites) × sites` estimate this count used to feed.
//!
//! A safepoint is any call-like expression: a call can re-enter the runtime
//! and collect. The count is an over-approximation biased toward spilling —
//! a false positive is a shadow frame on a function that would have been fine
//! (cheap; the shadow lowering is the pre-#7370 default), while a false
//! negative would let relocation fan-out reach the optimizer. Nested closures
//! and collect. The count over-approximates what RS4GC will see. Nested closures
//! are NOT counted: each compiles to its own `LlFunction` with its own frame,
//! so its safepoints belong to it (`walk_expr_children` does not descend into
//! a closure's body, only its parameter defaults).
Expand Down
15 changes: 8 additions & 7 deletions crates/perry-codegen/src/function.rs
Original file line number Diff line number Diff line change
Expand Up @@ -163,10 +163,10 @@ pub struct LlFunction {
/// `js_shadow_slot_bind` calls, removes the calls, and emits stack maps.
stack_map_slot_count: u32,
/// #8583: force this function onto the heap-backed shadow frame even when
/// native stack-map roots are the build default. Set for a function whose
/// estimated statepoint relocation count (`live_roots × safepoints`) would
/// make `rewrite-statepoints-for-gc` fan-out super-linear in the optimizer
/// (`codegen/helpers::maybe_spill_roots_to_shadow_frame`). The shadow-frame
/// native stack-map roots are the build default. Set by the backend's
/// retry for a function whose exact statepoint relocation count
/// (`inprocess::gc_liveness`, RFC deferred collection S4) or post-RS4GC
/// size exceeds its budget (`apply_budget_spill_retry`). The shadow-frame
/// lowering is the pre-#7370 default, walked by the same runtime root scan
/// as stack maps, so a spilled function's roots stay precise — it simply
/// carries no `gc "statepoint-example"` strategy and RS4GC skips it. The
Expand Down Expand Up @@ -335,9 +335,10 @@ impl LlFunction {
/// #8583/#8679: route this function's precise roots through the heap
/// shadow frame instead of native statepoints.
///
/// The estimate-driven path calls this before `enable_shadow_frame`, while
/// the post-RS4GC budget retry calls it after lowering is complete. In the
/// latter case the native-root path deliberately retained the original
/// The backend's budget retry calls this after lowering is complete (the
/// pre-RS4GC relocation count and the post-RS4GC size both decide there);
/// tests may also call it before `enable_shadow_frame`. After lowering,
/// the native-root path has deliberately retained the original
/// `js_shadow_slot_bind` calls until final rendering, so converting the
/// recorded stack-map request back into a shadow-frame push is a complete
/// re-lowering: final rendering keeps those binds, adds the matching pops,
Expand Down
Loading
Loading