diff --git a/changelog.d/11624-s4-exact-relocation-count.md b/changelog.d/11624-s4-exact-relocation-count.md new file mode 100644 index 0000000000..fdb86bb39d --- /dev/null +++ b/changelog.d/11624-s4-exact-relocation-count.md @@ -0,0 +1,6 @@ +- **Codegen: the shadow-frame spill is decided on RS4GC's exact relocation count (RFC deferred collection, S4).** #8583 moved a function's GC roots to a shadow frame when `(root slots + call sites) × call sites` exceeded 32 M. On the claude-code bundle that estimate overshot the real `gc.relocate` count about 100×, so it spilled functions RS4GC handles in seconds (`__87158`: estimated 134.5 M, real 42.9 k). + - **The count.** `inprocess/gc_liveness.rs` runs between the two halves of the statepoint pipeline: `always-inline,function(mem2reg,sccp)` first, then the count, then `rewrite-statepoints-for-gc`. So it sees exactly RS4GC's input, with every `gc-leaf-function` mark (S0, S1, S2) already on the calls and every S3-rematerialized root already a fresh load. It models RS4GC's liveness (a call's GC-pointer arguments are live across it), its CFG cleanup (`noreturn` cuts, `nounwind` invokes becoming calls, constant branches, phi folding), the single-use `icmp` sink, both relocation edges of an `invoke`, `TargetLibraryInfo` leaf calls, and `findBasePointer` for phis and selects. A phi that merges a NaN-box tag constant with a heap pointer gets a fresh `.base` phi, so it costs two relocations per crossing. The cost is linear in the IR plus the liveness it reports. + - **Exact against RS4GC.** Under `PERRY_CODEGEN_UNIT_TIMINGS` the backend counts the `gc.relocate`s RS4GC really emitted and prints one audit line per function. On the gap suite and on the whole claude-code bundle, prediction and RS4GC agree on every function (see the PR for the numbers). + - **The decision.** The HIR-level estimate (`maybe_spill_roots_to_shadow_frame`) and the constructed-IR `(allocas + sites) × sites` preflight are gone. A function over the budget gets the existing retry, re-lowered onto a shadow frame before RS4GC runs. Its log line now prints the real count, the statepoints (and how many are invokes) and the largest live set. The budget, `PERRY_ROOT_SPILL_RELOCATIONS` (default 1.5 Mi), is now the post-RS4GC instruction budget. Every relocation is one instruction of the rewritten body, so a larger count is over that backstop by construction. + - **The fast-emit cliff (found in validation).** Relocations aren't the only thing that grows the rewritten body past a cliff: `default_fast_emit_max_instrs` (600 k on x86-64, 100 k elsewhere; `PERRY_LL_FAST_EMIT_MAX_INSTRS`) is the point where LLVM's optimized machine pipeline gets bounded to an O0 fallback for instruction selection and register allocation (#10586). The relocation cap alone missed it: `__25747` sits at 0.4 M relocations, comfortably under the 1.5 Mi cap, but its rewritten body crossed 600 k instructions and fell into O0, growing its `.text` ~11×. The preflight now also predicts a function's post-RS4GC instruction count (pre-rewrite size plus a growth factor times the predicted relocations, calibrated against a 128-unit claude-code audit — see `POST_RS4GC_GROWTH_FACTOR`) and spills when *either* the relocation cap or this fast-emit prediction is exceeded. The prediction is deliberately conservative: it estimates the raw post-rewrite size, which is an upper bound on the further-optimized size the real fast-emit decision measures, so it can spill early but never miss a real cliff. On the bundle, two of the previously-cited functions (`__87158`, `__85198`) no longer spill under either check; `__25747` spills again (correctly, for the fast-emit reason this time) and `__84092` still doesn't (its relocation and predicted-instruction counts both stay well under budget). + - **For S4b.** `FunctionLiveness::safepoints` lists every call RS4GC will turn into a statepoint, with its per-edge relocation count. It is an internal API; nothing uses it yet. diff --git a/crates/perry-codegen/src/codegen/closure.rs b/crates/perry-codegen/src/codegen/closure.rs index 1bdeb86124..320741e455 100644 --- a/crates/perry-codegen/src/codegen/closure.rs +++ b/crates/perry-codegen/src/codegen/closure.rs @@ -616,12 +616,6 @@ pub(super) fn compile_closure( let capture_root_slots = u32::from(captures_this || enclosing_class.is_some() || entry_bound_this) + u32::from(captures_new_target); - crate::codegen::helpers::maybe_spill_roots_to_shadow_frame( - lf, - &llvm_name, - m.len() + capture_root_slots as usize, - body, - ); lf.enable_shadow_frame(m.len() as u32 + capture_root_slots); m } else { diff --git a/crates/perry-codegen/src/codegen/function.rs b/crates/perry-codegen/src/codegen/function.rs index 775b0a5e55..7f1b7ce0db 100644 --- a/crates/perry-codegen/src/codegen/function.rs +++ b/crates/perry-codegen/src/codegen/function.rs @@ -746,12 +746,6 @@ pub(super) fn compile_function( ); // Root the entry `this` slot of a this-reading body. let this_root_slots = usize::from(reads_this); - crate::codegen::helpers::maybe_spill_roots_to_shadow_frame( - lf, - &llvm_name, - m.len() + this_root_slots, - &f.body, - ); lf.enable_shadow_frame((m.len() + this_root_slots) as u32); m } else { diff --git a/crates/perry-codegen/src/codegen/helpers.rs b/crates/perry-codegen/src/codegen/helpers.rs index 9b40f188dd..d13ce94e68 100644 --- a/crates/perry-codegen/src/codegen/helpers.rs +++ b/crates/perry-codegen/src/codegen/helpers.rs @@ -502,166 +502,6 @@ pub(crate) fn inline_hot_small_max_call_sites() -> u32 { }) } -/// #8583: statepoint relocation estimate above which a function keeps its GC -/// roots in a shadow frame instead of native statepoints. -/// -/// `rewrite-statepoints-for-gc` adds one relocation per GC value live across -/// each safepoint, so the optimizer's post-rewrite cost scales with -/// `live_roots × safepoints`. Past a point that fan-out makes the `-Os`/`-O3` -/// middle-end super-linear and the compile does not finish (the Claude Code -/// bundle's 68 MB entry body measured 795 root slots × ~106k safepoints ≈ 8.4e7 -/// and grew 439k → 6.5M instructions under RS4GC; without RS4GC the same unit -/// optimized at `-Os` in ~5s). Real functions sit orders of magnitude below -/// this: hundreds of call sites times tens of slots is ~1e4–1e5. -/// -/// The default (#8620) is measured, not guessed. Synthetic entry functions with -/// a controlled `slots × safepoints` estimate were compiled at `-Os` with -/// spilling OFF (pure RS4GC fan-out) and the `@main` codegen unit timed: -/// -/// | estimate | fan-out finish | -/// |---------:|---------------:| -/// | 8.0M | ~325 s | -/// | 16.0M | ~235 s | -/// | 32.0M | ~511 s (8.5m) | -/// | 40.0M | did not finish in 20 min | -/// | 48.0M | did not finish in 20 min | -/// -/// The fan-out cliff sits between 32M and 40M, so the default is the largest -/// estimate whose fan-out still finished in bounded time. Below it fan-out is -/// the cheaper lowering — spilling a moderate function costs more than the -/// fan-out it avoids (an ~8M function spilled in 303 s vs 180 s fanned out, -/// #8620) — and above it fan-out risks not finishing and the shadow frame wins. -/// The former 4M default fired on ~8M functions that fan out fine in minutes. -/// The post-RS4GC instruction budget (#8586/#8679, inprocess.rs) backstops any -/// function this estimate misses: it re-lowers that function onto a precise -/// shadow frame and retries before LLVM's optimizer can hang, so raising the -/// estimate threshold is safe. -/// -/// `PERRY_ROOT_SPILL_RELOCATIONS=` overrides it; `0` disables spilling -/// (every function stays on native statepoints, the pre-#8583 behavior). -const DEFAULT_ROOT_SPILL_RELOCATIONS: usize = 32_000_000; - -pub(crate) fn root_spill_relocation_threshold() -> usize { - std::env::var("PERRY_ROOT_SPILL_RELOCATIONS") - .ok() - .and_then(|v| v.trim().parse::().ok()) - .unwrap_or(DEFAULT_ROOT_SPILL_RELOCATIONS) -} - -/// The relocation estimate for a function with `slot_count` GC-root slots and -/// a body containing `safepoint_sites` call-like expressions. Saturating so a -/// pathological product cannot wrap. -/// The root population RS4GC actually relocates: named pointer locals plus -/// ~one live pointer temporary per call result (#8583). Production and the -/// threshold tests must agree on this composition — computing it in only one -/// of the two is how the endpoint tests silently stop guarding the real -/// formula. -pub(crate) fn spill_live_root_count(slot_count: usize, safepoint_sites: usize) -> usize { - slot_count.saturating_add(safepoint_sites) -} - -pub(crate) fn root_relocation_estimate(slot_count: usize, safepoint_sites: usize) -> usize { - slot_count.saturating_mul(safepoint_sites) -} - -#[cfg(test)] -mod root_spill_default_tests { - use super::{root_relocation_estimate, spill_live_root_count, DEFAULT_ROOT_SPILL_RELOCATIONS}; - - /// Exactly what `maybe_spill_roots_to_shadow_frame` computes, so these - /// endpoint tests track the production formula instead of a stale copy of - /// it (#8633 changed the composition; before this helper the tests still - /// asserted on the pre-#8633 `slot_count x sites`). - fn production_estimate(slot_count: usize, sites: usize) -> usize { - root_relocation_estimate(spill_live_root_count(slot_count, sites), sites) - } - - /// #8620: the default is pinned to the measured RS4GC fan-out cliff — the - /// largest estimate whose fan-out finished in bounded time (32M finished in - /// ~8.5 min; 40M/48M did not finish in 20 min). Change it only with fresh - /// measurement. - #[test] - fn default_sits_at_the_measured_fan_out_cliff() { - assert_eq!(DEFAULT_ROOT_SPILL_RELOCATIONS, 32_000_000); - } - - /// The moderate case the old 4M default wrongly spilled (#8620): ~8M - /// relocations (4000 root slots × ~2001 safepoints) fans out in minutes, so - /// under the new default it stays on native statepoints. - #[test] - fn moderate_fan_out_stays_on_statepoints() { - let est = production_estimate(4000, 2001); - assert_eq!(est, 12_008_001); - assert!( - est <= DEFAULT_ROOT_SPILL_RELOCATIONS, - "moderate estimate {est} must not exceed the default (would spill)", - ); - } - - /// The genuinely-catastrophic case (Claude Code `cli.js` `@main`, - /// ~795 slots × ~106k safepoints ≈ 8.4e7, never finishes at `-Os`) must - /// still spill under the new default. - #[test] - fn catastrophic_fan_out_still_spills() { - let est = production_estimate(795, 106_000); - assert!( - est > DEFAULT_ROOT_SPILL_RELOCATIONS, - "catastrophic estimate {est} must exceed the default (should spill)", - ); - } -} - -/// Decide whether `func` should spill its roots to the shadow frame, and if so -/// mark it (BEFORE its `enable_*_shadow_frame` call) and report it. Only -/// meaningful under native stack-map roots — the shadow frame is already the -/// lowering otherwise. Reporting is at default verbosity because #8421 requires -/// that a change to how a function is compiled is never silent; the message -/// states that the optimization level is unchanged. -pub(super) fn maybe_spill_roots_to_shadow_frame( - func: &mut crate::function::LlFunction, - fn_name: &str, - slot_count: usize, - body: &[perry_hir::Stmt], -) { - if !native_stack_roots_enabled() { - return; - } - let threshold = root_spill_relocation_threshold(); - if threshold == 0 { - return; - } - let sites = crate::collectors::count_safepoint_sites(body); - // #8583 (unit-4 / `__33499` of the Claude Code bundle): `slot_count` is the - // shadow-slot map size — the count of *named* pointer-typed locals — but - // that is NOT the root population RS4GC relocates. A call-heavy minified - // closure produces one pointer-typed *temporary* per call result (the - // constructed IR carries ~one `alloca ptr addrspace(1)` per call), and each - // is live across the later safepoints; those temporaries dominate the true - // root count yet are invisible to `collect_pointer_typed_locals`. `__33499` - // measured ~20.3k named-and-anonymous pointer roots × ~20.3k safepoints, but - // its `slot_count` alone was ~100x smaller, so `slot_count × sites` fell - // under the threshold, the function stayed on statepoints, and RS4GC then - // fanned out for >3 h / ~30 GiB (never reaching the #8586 post-rewrite - // budget assertion, which only fires *after* the rewrite it never finishes). - // Count each safepoint as contributing ~one live pointer temporary. This is - // an over-approximation biased toward spilling — the intended direction (a - // false-positive shadow frame is cheap; a missed fan-out is not). - let live_roots = spill_live_root_count(slot_count, sites); - let estimate = root_relocation_estimate(live_roots, sites); - if estimate <= threshold { - return; - } - func.request_shadow_frame_spill(); - eprintln!( - "perry: `{fn_name}` keeps its {live_roots} GC roots (incl. call-result temporaries) in a \ - shadow frame instead of statepoints: an estimated {estimate} relocations ({live_roots} \ - roots × {sites} safepoints) would make rewrite-statepoints-for-gc fan-out super-linear in \ - the optimizer (> {threshold}). The function is still compiled at the requested \ - optimization level; only its GC-root representation changes, and its roots stay \ - precise (#8583). Override with PERRY_ROOT_SPILL_RELOCATIONS." - ); -} - /// #10663: a function body with at least this many property stores outside /// any loop outlines those stores' inline caches. /// @@ -721,10 +561,6 @@ pub(super) fn enable_module_init_shadow_frame( let shadow_slot_map = crate::collectors::collect_pointer_typed_locals(&[], stmts, flat_const_ids); - // #8583: the module-entry body is the minified-bundle IIFE — the function - // that fans out catastrophically under RS4GC. Decide its root lowering - // before the frame is built. - maybe_spill_roots_to_shadow_frame(func, "main", shadow_slot_map.len(), stmts); func.enable_post_init_shadow_frame(shadow_slot_map.len() as u32); let shadow_slot_clears_after_stmt = crate::collectors::collect_shadow_slot_clear_points(stmts, &shadow_slot_map); diff --git a/crates/perry-codegen/src/codegen/method.rs b/crates/perry-codegen/src/codegen/method.rs index dfbb207579..d9ef95832c 100644 --- a/crates/perry-codegen/src/codegen/method.rs +++ b/crates/perry-codegen/src/codegen/method.rs @@ -368,12 +368,6 @@ pub(super) fn compile_method( ), &cross_module.scope_map, ); - crate::codegen::helpers::maybe_spill_roots_to_shadow_frame( - lf, - &llvm_name, - m.len() + 1, - method_body, - ); lf.enable_shadow_frame(m.len() as u32 + 1); m } else { diff --git a/crates/perry-codegen/src/codegen/method_static.rs b/crates/perry-codegen/src/codegen/method_static.rs index 9def7626f5..0674c60358 100644 --- a/crates/perry-codegen/src/codegen/method_static.rs +++ b/crates/perry-codegen/src/codegen/method_static.rs @@ -70,12 +70,6 @@ pub(in crate::codegen) fn compile_static_method( crate::collectors::collect_pointer_typed_locals(&f.params, &f.body, &flat_const_ids), &cross_module.scope_map, ); - crate::codegen::helpers::maybe_spill_roots_to_shadow_frame( - lf, - &llvm_name, - m.len() + 1, - &f.body, - ); lf.enable_shadow_frame(m.len() as u32 + 1); m } else { diff --git a/crates/perry-codegen/src/collectors/safepoint_sites.rs b/crates/perry-codegen/src/collectors/safepoint_sites.rs index d96f39e0f5..26bbf52eb2 100644 --- a/crates/perry-codegen/src/collectors/safepoint_sites.rs +++ b/crates/perry-codegen/src/collectors/safepoint_sites.rs @@ -7,15 +7,15 @@ //! body of the Claude Code bundle measured 795 root slots × ~106k safepoints //! and grew 439k → 6.5M instructions under RS4GC, and a single `-Os` pass on //! the result did not finish in practical time (#8583). -//! `codegen/helpers::maybe_spill_roots_to_shadow_frame` multiplies this count -//! by the function's root-slot count and, past a threshold, keeps that -//! function's roots in a shadow frame instead of statepoints. +//! Entry outlining (#8595, `codegen/entry_outline.rs`) uses this count to cut +//! a huge entry body into chunks before any of it is lowered. It is a +//! source-level proxy only: whether a function keeps its roots on statepoints +//! is decided on the exact relocation count of the lowered IR +//! (`inprocess::gc_liveness`, RFC deferred collection S4), which replaced the +//! `(slots + sites) × sites` estimate this count used to feed. //! //! A safepoint is any call-like expression: a call can re-enter the runtime -//! and collect. The count is an over-approximation biased toward spilling — -//! a false positive is a shadow frame on a function that would have been fine -//! (cheap; the shadow lowering is the pre-#7370 default), while a false -//! negative would let relocation fan-out reach the optimizer. Nested closures +//! and collect. The count over-approximates what RS4GC will see. Nested closures //! are NOT counted: each compiles to its own `LlFunction` with its own frame, //! so its safepoints belong to it (`walk_expr_children` does not descend into //! a closure's body, only its parameter defaults). diff --git a/crates/perry-codegen/src/function.rs b/crates/perry-codegen/src/function.rs index a06c11b98f..a8aecb2a6d 100644 --- a/crates/perry-codegen/src/function.rs +++ b/crates/perry-codegen/src/function.rs @@ -163,10 +163,10 @@ pub struct LlFunction { /// `js_shadow_slot_bind` calls, removes the calls, and emits stack maps. stack_map_slot_count: u32, /// #8583: force this function onto the heap-backed shadow frame even when - /// native stack-map roots are the build default. Set for a function whose - /// estimated statepoint relocation count (`live_roots × safepoints`) would - /// make `rewrite-statepoints-for-gc` fan-out super-linear in the optimizer - /// (`codegen/helpers::maybe_spill_roots_to_shadow_frame`). The shadow-frame + /// native stack-map roots are the build default. Set by the backend's + /// retry for a function whose exact statepoint relocation count + /// (`inprocess::gc_liveness`, RFC deferred collection S4) or post-RS4GC + /// size exceeds its budget (`apply_budget_spill_retry`). The shadow-frame /// lowering is the pre-#7370 default, walked by the same runtime root scan /// as stack maps, so a spilled function's roots stay precise — it simply /// carries no `gc "statepoint-example"` strategy and RS4GC skips it. The @@ -335,9 +335,10 @@ impl LlFunction { /// #8583/#8679: route this function's precise roots through the heap /// shadow frame instead of native statepoints. /// - /// The estimate-driven path calls this before `enable_shadow_frame`, while - /// the post-RS4GC budget retry calls it after lowering is complete. In the - /// latter case the native-root path deliberately retained the original + /// The backend's budget retry calls this after lowering is complete (the + /// pre-RS4GC relocation count and the post-RS4GC size both decide there); + /// tests may also call it before `enable_shadow_frame`. After lowering, + /// the native-root path has deliberately retained the original /// `js_shadow_slot_bind` calls until final rendering, so converting the /// recorded stack-map request back into a shadow-frame push is a complete /// re-lowering: final rendering keeps those binds, adds the matching pops, diff --git a/crates/perry-codegen/src/inprocess.rs b/crates/perry-codegen/src/inprocess.rs index dbbbeecc56..7a2b8f757b 100644 --- a/crates/perry-codegen/src/inprocess.rs +++ b/crates/perry-codegen/src/inprocess.rs @@ -17,6 +17,7 @@ //! IR and flags this pipeline produces objects byte-identical to Homebrew //! clang 22's `clang -c`. +pub(crate) mod gc_liveness; mod optimize_emit; mod split_emit; use optimize_emit::optimize_and_emit; @@ -35,6 +36,7 @@ use inkwell::targets::{ use inkwell::values::AsValueRef; use inkwell::OptimizationLevel; +#[cfg(test)] use crate::linker::STATEPOINT_REWRITE_PASSES; /// Test seam (#7502): parse `ll_text`, run [`STATEPOINT_REWRITE_PASSES`] for @@ -367,6 +369,16 @@ pub struct UnitCodegenStats { pub post_rewrite_instructions: usize, pub post_rewrite_widest: Option<(String, usize)>, pub rewrite_secs: f64, + /// Time in the S4 liveness count between RS4GC's canonicalization and + /// the rewrite itself. + pub liveness_secs: f64, + /// Statepoints the liveness count predicted for the unit. + pub statepoints: usize, + /// `gc.relocate`s predicted (upper bound) and produced for the unit. + pub relocations_predicted: u64, + pub relocations_actual: u64, + /// Functions whose prediction differed from RS4GC's output. + pub liveness_mismatches: usize, pub optimize_secs: f64, pub emit_secs: f64, /// Functions stamped `"disable-tail-calls"` because their alloca-walk @@ -686,6 +698,46 @@ fn fast_emit_fallbacks( /// lowers it, `warn:` only warns, and `0`/`off` disables the check. const DEFAULT_RS4GC_MAX_INSTRS: usize = 1_572_864; +/// Instructions RS4GC adds to a function's body per predicted relocation, +/// used to predict whether a function will cross the fast-emit budget +/// ([`default_fast_emit_max_instrs`]) *before* paying for the rewrite and +/// the IR optimizer (#11624 follow-up: the relocation cap alone let +/// `__25747` reach the claude-code bundle with 0.4 M relocations — comfortably +/// under the 1.5 Mi relocation cap — but its rewritten body crossed the +/// 600 k-instruction x86-64 fast-emit budget and fell back to LLVM's O0 +/// machine pipeline, growing its `.text` ~11x). +/// +/// Measured on the #11624 128-unit claude-code 2.1.112 audit +/// (`PERRY_CODEGEN_UNIT_TIMINGS`): summed over all 128 units, post-RS4GC +/// instructions exceeded pre-RS4GC instructions by 10,006,633 while RS4GC +/// emitted 8,751,060 relocations — a corpus-wide average of ~1.14 +/// instructions per relocation. Per-function samples (the widest function in +/// each of 5 audited units) ranged from 1.28x to 5.15x, so this constant is +/// rounded well above the corpus average for headroom. A low-relocation +/// function is insensitive to this factor's precision either way — its +/// pre-rewrite size already dominates the prediction and keeps it far under +/// budget — so the imprecision this rounds past only matters for the +/// high-relocation functions where the fast-emit cliff can actually happen, +/// and those are exactly the ones the corpus average describes. +const POST_RS4GC_GROWTH_FACTOR: f64 = 2.0; + +/// Predict a function's post-RS4GC instruction count from its pre-rewrite +/// size and its predicted relocation count (see [`POST_RS4GC_GROWTH_FACTOR`]). +/// +/// This deliberately predicts the RAW post-rewrite count, not the count +/// after the IR optimizer that runs on top of it — `fast_emit_fallbacks` +/// compares against the latter, which is measured after the pipeline has +/// had a chance to shrink the rewritten body (DCE, SimplifyCFG, and friends, +/// on IR that RS4GC's rewrite already canonicalized). The raw count is +/// therefore an upper bound on what `fast_emit_fallbacks` will see: this can +/// spill a function whose optimized size would have stayed under budget, but +/// never the reverse. Conservative in the safe direction, per the #11624 +/// follow-up. +fn predicted_post_rewrite_instructions(pre_instructions: usize, relocations: u64) -> usize { + let growth = (relocations as f64 * POST_RS4GC_GROWTH_FACTOR).ceil(); + pre_instructions.saturating_add(growth as usize) +} + #[derive(Debug, Clone, Copy, PartialEq, Eq)] enum RewriteBudget { Off, @@ -702,13 +754,40 @@ pub(crate) enum Rs4gcBudgetCause { /// non-leaf call sites LLVM will actually see, rather than another source /// syntax approximation. PreRewrite { - root_allocas: usize, - safepoints: usize, - estimated_relocations: usize, + /// Calls RS4GC turns into statepoints. + statepoints: usize, + /// Of those, `invoke`s: relocated on the normal and the unwind edge. + invoke_statepoints: usize, + /// Most GC values live across one statepoint. + max_live: u32, + /// The `gc.relocate`s RS4GC would emit (an upper bound when the + /// function has derived pointers; see `gc_liveness`). + relocations: u64, }, /// RS4GC finished, but its relocation fan-out made the rewritten body too /// large for the normal optimization pipeline. PostRewrite { post_instructions: usize }, + /// Predicted (before RS4GC runs) to cross the fast-emit machine-pipeline + /// budget — the cliff into LLVM's O0 instruction selection / register + /// allocation (#10586) — even though it is under the relocation cap. + /// See [`predicted_post_rewrite_instructions`]. + PredictedFastEmit { + /// The `gc.relocate`s RS4GC would emit. + relocations: u64, + /// [`predicted_post_rewrite_instructions`]'s estimate. + predicted_instructions: usize, + }, +} + +impl Rs4gcBudgetCause { + fn pre_rewrite(l: &gc_liveness::FunctionLiveness) -> Self { + Rs4gcBudgetCause::PreRewrite { + statepoints: l.statepoints(), + invoke_statepoints: l.invoke_statepoints(), + max_live: l.max_live(), + relocations: l.relocation_bound(), + } + } } #[derive(Debug, Clone, PartialEq, Eq)] @@ -849,113 +928,138 @@ fn rs4gc_functions(module: &inkwell::module::Module<'_>) -> std::collections::Ha names } -/// The two constructed-IR factors that bound RS4GC relocation fan-out. +/// #8583: relocation budget above which a function keeps its GC roots in a +/// shadow frame instead of native statepoints. /// -/// Count only allocas whose payload is a managed pointer and call sites which -/// are not explicitly marked as GC leaves. LLVM intrinsics are also leaves: -/// they cannot enter Perry's runtime or collect. This is deliberately the -/// same conservative model as the source-level spill estimate — each -/// safepoint can leave one additional pointer result live across later calls — -/// but it observes the calls codegen actually emitted. That closes estimator -/// holes where one source expression expands into several collecting helpers. -fn rs4gc_preflight_factors(function: inkwell::values::FunctionValue<'_>) -> (usize, usize) { - let mut root_allocas = 0usize; - let mut safepoints = 0usize; - for bb in function.get_basic_blocks() { - let mut inst = bb.get_first_instruction(); - while let Some(i) = inst { - match i.get_opcode() { - inkwell::values::InstructionOpcode::Alloca => { - if matches!( - i.get_allocated_type(), - Ok(inkwell::types::BasicTypeEnum::PointerType(ptr)) - if ptr.get_address_space() == inkwell::AddressSpace::from(1u16) - ) { - root_allocas += 1; - } - } - inkwell::values::InstructionOpcode::Call - | inkwell::values::InstructionOpcode::CallBr - | inkwell::values::InstructionOpcode::Invoke => { - // Call, invoke and callbr are all LLVM CallBase values, so - // the call-site attribute API is valid for each opcode. - let call = unsafe { inkwell::values::CallSiteValue::new(i.as_value_ref()) }; - let gc_leaf = call - .get_string_attribute( - inkwell::attributes::AttributeLoc::Function, - "gc-leaf-function", - ) - .is_some(); - let intrinsic = call - .get_called_fn_value() - .map_or(false, |callee| callee.get_intrinsic_id() != 0); - if !gc_leaf && !intrinsic { - safepoints += 1; - } - } - _ => {} - } - inst = i.get_next_instruction(); +/// The budget is on the number RS4GC will really produce: the exact count of +/// `gc.relocate`s ([`gc_liveness`], RFC deferred collection S4), measured on +/// RS4GC's own input just before it runs. It used to be on an estimate, +/// `(root slots + call sites) × call sites`, which overshot the claude-code +/// bundle about 100× and moved functions to shadow frames that RS4GC would +/// have handled in seconds. +/// +/// The default is the post-RS4GC instruction budget +/// ([`DEFAULT_RS4GC_MAX_INSTRS`], 1.5 Mi), and that is derived, not tuned. +/// Every relocation is one `gc.relocate` instruction in the rewritten body, +/// so a function over this many relocations is over the post-RS4GC budget by +/// construction: that backstop would spill it anyway, after RS4GC had spent +/// its time on it. Any lower default would spill functions the measured +/// optimizer limit (#8128) accepts. What the pre-rewrite check adds is that +/// the decision now costs milliseconds instead of an RS4GC run. +/// +/// Measured (LLVM 22 `opt`, perrymaster) on synthetic straight-line functions +/// of known count: 0.37 M relocations → RS4GC 4 s; 1.1 M → 19 s; 2.5 M → 51 s; +/// 4.5 M → 122 s, so RS4GC stays in bounded time right up to this budget. +/// On the claude-code 2.1.112 bundle after S0–S3 the largest real count is +/// 0.47 M (`__25747`), so nothing spills there. The old estimate spilled +/// `__87158`, `__84092` and `__85198` (42.9 k, 25.7 k and 17.4 k real +/// relocations). +/// +/// Relocations are not RS4GC's only cost driver. A branchy function with +/// hundreds of values live through many small conditional blocks costs +/// minutes at a few hundred thousand relocations (a synthetic diamond chain, +/// 500 values × 250 safepoints: 0.25 M relocations, 540 s), and neither this +/// budget nor the post-RS4GC one sees it. No real function in the bundle is +/// that shape today; the per-safepoint counts ([`gc_liveness`]) are what a +/// shape-aware budget would read. +/// +/// `PERRY_ROOT_SPILL_RELOCATIONS=` overrides it; `0` disables spilling +/// (every function stays on native statepoints, the pre-#8583 behavior). +pub(crate) const DEFAULT_ROOT_SPILL_RELOCATIONS: u64 = DEFAULT_RS4GC_MAX_INSTRS as u64; + +pub(crate) fn root_spill_relocation_threshold() -> u64 { + #[cfg(test)] + if let Some(cap) = TEST_ROOT_SPILL_RELOCATIONS.with(std::cell::Cell::get) { + return cap; + } + std::env::var("PERRY_ROOT_SPILL_RELOCATIONS") + .ok() + .and_then(|v| v.trim().parse::().ok()) + .unwrap_or(DEFAULT_ROOT_SPILL_RELOCATIONS) +} + +#[cfg(test)] +thread_local! { + static TEST_ROOT_SPILL_RELOCATIONS: std::cell::Cell> = const { + std::cell::Cell::new(None) + }; +} + +/// Thread-local seam for the relocation budget, so a test can move the +/// threshold without mutating `PERRY_ROOT_SPILL_RELOCATIONS` under other +/// concurrently running LLVM tests. +#[cfg(test)] +pub(crate) fn with_test_root_spill_threshold(cap: u64, run: impl FnOnce() -> T) -> T { + struct Restore(Option); + impl Drop for Restore { + fn drop(&mut self) { + TEST_ROOT_SPILL_RELOCATIONS.set(self.0); } } - (root_allocas, safepoints) + let old = TEST_ROOT_SPILL_RELOCATIONS.replace(Some(cap)); + let _restore = Restore(old); + run() } -/// Every RS4GC-participating function whose constructed IR predicts more -/// relocation work than the source-level spill budget permits. +/// Every RS4GC-participating function whose exact relocation count exceeds +/// `cap`, OR — under `cap` but predicted to cross the fast-emit +/// machine-pipeline budget (`fast_emit_cap`, `None` when that fallback is +/// disabled) — whose predicted post-rewrite instruction count exceeds it. +/// The relocation cap is checked first: it already implies a spill, and its +/// message is the more specific one when both would apply. `cap == 0` +/// disables spilling entirely (`PERRY_ROOT_SPILL_RELOCATIONS=0`), including +/// the fast-emit prediction. fn rs4gc_preflight_violations( - module: &inkwell::module::Module<'_>, - cap: usize, - rewritten_functions: &std::collections::HashSet, -) -> Vec<(String, usize, usize, usize)> { + liveness: &[(String, gc_liveness::FunctionLiveness)], + cap: u64, + pre: &std::collections::HashMap, + fast_emit_cap: Option, +) -> Vec { if cap == 0 { return Vec::new(); } - let mut over = Vec::new(); - let mut function = module.get_first_function(); - while let Some(f) = function { - if f.count_basic_blocks() > 0 { - let name = f.get_name().to_string_lossy().into_owned(); - if rewritten_functions.contains(&name) { - let (root_allocas, safepoints) = rs4gc_preflight_factors(f); - let live_roots = - crate::codegen::helpers::spill_live_root_count(root_allocas, safepoints); - let estimate = - crate::codegen::helpers::root_relocation_estimate(live_roots, safepoints); - if estimate > cap { - over.push((name, root_allocas, safepoints, estimate)); - } + liveness + .iter() + .filter_map(|(name, l)| { + let relocations = l.relocation_bound(); + if relocations > cap { + return Some(Rs4gcBudgetViolation { + name: name.clone(), + pre_instructions: pre.get(name).copied(), + cause: Rs4gcBudgetCause::pre_rewrite(l), + cap: cap as usize, + }); } - } - function = f.get_next_function(); - } - over + let fast_emit_cap = fast_emit_cap?; + let pre_instructions = pre.get(name).copied().unwrap_or(0); + let predicted_instructions = + predicted_post_rewrite_instructions(pre_instructions, relocations); + if predicted_instructions > fast_emit_cap { + Some(Rs4gcBudgetViolation { + name: name.clone(), + pre_instructions: Some(pre_instructions), + cause: Rs4gcBudgetCause::PredictedFastEmit { + relocations, + predicted_instructions, + }, + cap: fast_emit_cap, + }) + } else { + None + } + }) + .collect() } -/// Stop before RS4GC itself enters its super-linear liveness/rewrite walk and -/// ask codegen to re-lower the named functions with precise shadow roots. +/// Stop before RS4GC runs and ask codegen to re-lower the named functions +/// with precise shadow roots. fn enforce_rs4gc_preflight_budget( - module: &inkwell::module::Module<'_>, - cap: usize, + liveness: &[(String, gc_liveness::FunctionLiveness)], + cap: u64, pre: &std::collections::HashMap, - rewritten_functions: &std::collections::HashSet, + fast_emit_cap: Option, ) -> Result<()> { - let violations: Vec = - rs4gc_preflight_violations(module, cap, rewritten_functions) - .into_iter() - .map( - |(name, root_allocas, safepoints, estimated_relocations)| Rs4gcBudgetViolation { - pre_instructions: pre.get(&name).copied(), - name, - cause: Rs4gcBudgetCause::PreRewrite { - root_allocas, - safepoints, - estimated_relocations, - }, - cap, - }, - ) - .collect(); + let violations = rs4gc_preflight_violations(liveness, cap, pre, fast_emit_cap); if violations.is_empty() { Ok(()) } else { @@ -993,15 +1097,17 @@ fn rewrite_budget_message(violation: &Rs4gcBudgetViolation, retry: bool) -> Stri }; match &violation.cause { Rs4gcBudgetCause::PreRewrite { - root_allocas, - safepoints, - estimated_relocations, + statepoints, + invoke_statepoints, + max_live, + relocations, } => format!( - "before rewrite-statepoints-for-gc, `{}` has {root_allocas} managed-root allocas and \ - {safepoints} non-leaf call sites; accounting for call-result temporaries predicts \ - {estimated_relocations} relocations, above the pre-rewrite budget {}. RS4GC's own \ - liveness/rewrite walk is super-linear on fan-out of this size; {outcome} (#8583). \ - Override with PERRY_ROOT_SPILL_RELOCATIONS= (raise) or =0 (disable).", + "before rewrite-statepoints-for-gc, `{}` has {statepoints} statepoints \ + ({invoke_statepoints} of them invokes, relocated on both edges) with up to \ + {max_live} GC values live across one; RS4GC would emit {relocations} relocations, \ + above the pre-rewrite budget {}. RS4GC and the optimizer after it are super-linear \ + on fan-out of this size; {outcome} (#8583). Override with \ + PERRY_ROOT_SPILL_RELOCATIONS= (raise) or =0 (disable).", violation.name, violation.cap ), Rs4gcBudgetCause::PostRewrite { post_instructions } => { @@ -1018,6 +1124,26 @@ fn rewrite_budget_message(violation: &Rs4gcBudgetViolation, retry: bool) -> Stri violation.name, violation.cap ) } + Rs4gcBudgetCause::PredictedFastEmit { + relocations, + predicted_instructions, + } => { + let pre = violation + .pre_instructions + .map(|n| format!(" ({n} before the rewrite)")) + .unwrap_or_default(); + format!( + "before rewrite-statepoints-for-gc, `{}` would emit {relocations} relocations, \ + under the relocation cap but predicted to grow the function to about \ + {predicted_instructions} instructions{pre} — above the fast-emit \ + machine-pipeline budget {}. Past that budget LLVM keeps the optimized IR but \ + falls back to its O0 machine pipeline for instruction selection and register \ + allocation, which is size, not correctness, but real (#11624); {outcome}. \ + Override with PERRY_LL_FAST_EMIT_MAX_INSTRS= (raise) or =0 (disable this \ + check and the O0 fallback it predicts).", + violation.name, violation.cap + ) + } } } diff --git a/crates/perry-codegen/src/inprocess/gc_liveness.rs b/crates/perry-codegen/src/inprocess/gc_liveness.rs new file mode 100644 index 0000000000..5619b045d8 --- /dev/null +++ b/crates/perry-codegen/src/inprocess/gc_liveness.rs @@ -0,0 +1,1285 @@ +//! Exact liveness of GC values across RS4GC safepoints (RFC "deferred +//! collection", step S4; #11528). +//! +//! `rewrite-statepoints-for-gc` (RS4GC) emits one `gc.relocate` per GC value +//! live across each safepoint, and two for a safepoint that is an `invoke`: +//! one on the normal edge and one in the landing pad. That relocation count is +//! what makes RS4GC and the optimizer after it slow on large functions (#8583). +//! This module computes it **before** RS4GC runs, on the module RS4GC is about +//! to rewrite, so the shadow-frame spill decision is a budget on the real +//! number instead of on `(slots + sites) × sites`, which overshot the +//! claude-code bundle by about 100× (1.34e9 estimated vs 13.9 M real). +//! +//! # Where it runs, and why there +//! +//! The production pipeline is `always-inline,function(mem2reg,sccp), +//! rewrite-statepoints-for-gc` ([`crate::linker::STATEPOINT_REWRITE_PASSES`]). +//! The backend runs its first half ([`crate::linker::STATEPOINT_PREPARE_PASSES`]), +//! calls [`analyze_function`] on every function that carries the statepoint GC +//! strategy, and only then runs RS4GC itself. The input is therefore exactly +//! RS4GC's input: after inlining, after `mem2reg` turned root allocas into SSA +//! values, and after `sccp` folded constants. Everything that decides whether a +//! call is a safepoint is already an attribute on the call by then — the +//! audited `gc-leaf-function` marking (S0, S1's generated call-effects table), +//! the IC fast paths (S2) — and a root that S3 rematerializes from its global +//! is a fresh load after each safepoint, not a value live across it. So this +//! analysis needs no knowledge of Perry's lowering, and cannot drift from it. +//! Shadow-stack targets never get here: their functions carry no GC strategy. +//! +//! Computing it on the textual IR before `mem2reg` (in `precise_roots.rs`) was +//! the alternative. It would have to re-derive `mem2reg` and `sccp` — which +//! slot stores fold to constants, which call results are dead — and would +//! miss the always-inlined helpers' calls. The LLVM-level count is exact at +//! the cost of splitting one `run_passes` call in two. +//! +//! # What it models +//! +//! RS4GC's liveness (`computeLiveInValues`/`findLiveSetAtInst`): a value of a +//! GC pointer type (`ptr addrspace(1)`, or a vector of them) that is not a +//! `Constant` is live across a safepoint when it is live *into* the call: a +//! use is reachable from the call without passing the value's definition. A +//! phi's operand is a use at the end of the incoming block. The call's own +//! result is not live across it, but its GC-pointer arguments are, even when +//! the call is their last use (RS4GC's `findLiveSetAtInst` walks the call's +//! operands too; Perry passes NaN-boxed `i64`s, so this is rare in its IR). +//! Before computing +//! liveness RS4GC also edits the function, and each edit changes the count: +//! +//! - `removeUnreachableBlocks` (`markAliveBlocks` in `Local.cpp`): code after +//! a `noreturn` call, a store to null/undef, an `assume(false)` and a call +//! through null/undef is deleted; an `invoke` of a `nounwind` callee becomes +//! a `call` (one relocation edge, not two), or disappears when its result is +//! unused and it has no side effects; an `invoke` of a `noreturn` callee +//! loses its normal successor; constant branch conditions and identical +//! branch targets are folded; unreachable blocks are deleted, and a phi that +//! lost an incoming edge folds when its remaining inputs agree; +//! - `FoldSingleEntryPHINodes` on every block with a unique predecessor; +//! - moving a single-use `icmp` that feeds a conditional branch down to the +//! branch, which extends its operands' live ranges to the terminator. +//! +//! A call is a safepoint unless it is inline asm, carries (or its callee +//! carries) `"gc-leaf-function"`, is an intrinsic other than +//! `gc.statepoint`'s deopt/element-atomic memcpy family, or calls a C library +//! function LLVM's `TargetLibraryInfo` knows ([`LIBFUNCS`]). +//! +//! Base pointers (`findBasePointer`): RS4GC relocates a value's *base* too. +//! Perry's phis and selects often merge a NaN-box tag constant +//! (`inttoptr (i64 0x7FFC… to ptr addrspace(1))`) with a heap pointer. The +//! constant's base is `null`, the pointer's is itself, so RS4GC clones the +//! phi into a `.base` phi that is live wherever the phi is: two relocations +//! per crossing. A phi of constants only has a constant base and is dropped +//! from the live set; a phi whose inputs are all bases (or `null`) is its own +//! base. The pruning and the optimistic meet are modeled as RS4GC runs them. +//! +//! One thing is bounded rather than exact: a `gep`, `addrspacecast`, +//! `bitcast` or `freeze` of a GC pointer (or a phi/select whose inputs agree +//! on one *other* existing base) makes RS4GC add that existing base to the +//! live set, where it may already be, or rematerialize the derived value. +//! Each adds at most one relocation per safepoint it crosses, so +//! [`FunctionLiveness::relocation_bound`] is an upper bound and +//! [`FunctionLiveness::is_exact`] says whether it is also exact. Perry's +//! lowering emits none of these today. The audit +//! (`PERRY_CODEGEN_UNIT_TIMINGS`) marks any function where the count is only +//! a bound, and every function where the prediction and RS4GC's real output +//! disagree. +//! +//! The count only chooses a *compile-time* lowering. RS4GC still decides, on +//! its own, what to relocate, so an error here can cost compile time but can +//! never make a root imprecise. +//! +//! # Cost +//! +//! One scan of the instructions, then a backward walk per GC value over the +//! blocks where it is live (the census's `gcm-stats` algorithm). That is +//! `O(instructions + Σ_v |live blocks of v|)`: linear in the IR plus the +//! liveness it reports, never more than RS4GC's own liveness, which +//! materializes the same (value, block) pairs as sets. Per-safepoint counts +//! come from a difference array, so a value live through a block with many +//! safepoints costs O(log n) there, not O(n). + +use std::collections::HashMap; +use std::hash::{BuildHasherDefault, Hasher}; +use std::sync::OnceLock; + +use llvm_sys::core::*; +use llvm_sys::prelude::{LLVMBasicBlockRef, LLVMTypeRef, LLVMValueRef}; +use llvm_sys::{LLVMOpcode, LLVMTailCallKind, LLVMTypeKind}; + +/// Live GC values across one safepoint. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) struct SafepointLiveness { + /// Index of the block in function order. + pub block: u32, + /// Index of the call within its block, in the analyzed (post-`sccp`) IR. + pub index: u32, + /// Relocation edges: 1 for a call, 2 for an `invoke` (normal + unwind). + pub edges: u8, + /// Values RS4GC relocates across the call, once per edge: the GC values + /// live across it, plus the fresh base of each live phi/select that + /// needs one, minus live values whose base is a constant. + pub live: u32, +} + +/// The liveness census of one function, as RS4GC will see it. +/// +/// `safepoints` is the internal API for S4b (rooting slow-path call sites in +/// GC-map stack slots): it lists every call RS4GC will turn into a statepoint, +/// in block order, with the number of values live across it. +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub(crate) struct FunctionLiveness { + pub instructions: usize, + pub blocks: usize, + /// GC-pointer SSA values (instructions and arguments) after RS4GC's edits. + pub gc_values: usize, + /// `gep`/cast/`freeze` values of a GC pointer, whose base RS4GC must + /// relocate alongside them (see the module docs). + pub derived_values: usize, + /// Phis/selects RS4GC gives a fresh base phi/select (a phi that merges a + /// tagged constant with a heap pointer, say): each crossing costs two + /// relocations, the value and its base. + pub fresh_bases: usize, + pub safepoints: Vec, + /// `Σ live × edges` over the safepoints: RS4GC's `gc.relocate` count + /// when [`Self::is_exact`]. + pub relocations: u64, + /// `Σ edges` over (possibly-derived value, safepoint it crosses) pairs: + /// the most base pointers RS4GC can add on top of `relocations`. + pub derived_crossings: u64, + /// Steps taken by the liveness walk (uses + live blocks + pred edges). + pub work: u64, + /// Wall time of the analysis, in microseconds (set by + /// [`analyze_module`]; reported by the audit). + pub micros: u64, +} + +impl FunctionLiveness { + pub fn statepoints(&self) -> usize { + self.safepoints.len() + } + + pub fn invoke_statepoints(&self) -> usize { + self.safepoints.iter().filter(|s| s.edges == 2).count() + } + + pub fn max_live(&self) -> u32 { + self.safepoints.iter().map(|s| s.live).max().unwrap_or(0) + } + + /// No derived pointer crosses a safepoint, so `relocations` is exact. + pub fn is_exact(&self) -> bool { + self.derived_crossings == 0 + } + + /// Upper bound on the relocations RS4GC emits; equal to `relocations` + /// when [`Self::is_exact`]. The spill decision budgets this number. + pub fn relocation_bound(&self) -> u64 { + self.relocations.saturating_add(self.derived_crossings) + } +} + +/// Every function of `module` named in `functions` (the ones carrying the +/// statepoint GC strategy), with its liveness. Order follows the module. +pub(crate) fn analyze_module( + module: &inkwell::module::Module<'_>, + functions: &std::collections::HashSet, +) -> Vec<(String, FunctionLiveness)> { + let mut out = Vec::new(); + let mut function = module.get_first_function(); + while let Some(f) = function { + if f.count_basic_blocks() > 0 { + let name = f.get_name().to_string_lossy(); + if functions.contains(name.as_ref()) { + use inkwell::values::AsValueRef; + let started = std::time::Instant::now(); + let mut l = analyze_function(f.as_value_ref()); + l.micros = started.elapsed().as_micros() as u64; + out.push((name.into_owned(), l)); + } + } + function = f.get_next_function(); + } + out +} + +/// Count `gc.relocate` calls in a rewritten function: the ground truth the +/// prediction is audited against. +pub(crate) fn count_relocates(function: LLVMValueRef) -> u64 { + let ids = ids(); + let mut n = 0u64; + unsafe { + let mut bb = LLVMGetFirstBasicBlock(function); + while !bb.is_null() { + let mut i = LLVMGetFirstInstruction(bb); + while !i.is_null() { + if LLVMGetInstructionOpcode(i) == LLVMOpcode::LLVMCall { + let callee = LLVMIsAFunction(LLVMGetCalledValue(i)); + if !callee.is_null() && LLVMGetIntrinsicID(callee) == ids.gc_relocate { + n += 1; + } + } + i = LLVMGetNextInstruction(i); + } + bb = LLVMGetNextBasicBlock(bb); + } + } + n +} + +/// Analyze one function. `function` must be a defined function of a live +/// module. +pub(crate) fn analyze_function(function: LLVMValueRef) -> FunctionLiveness { + // SAFETY: the caller hands us a function of a module it owns and does not + // mutate during the call; every C API below only reads the IR. + unsafe { analyze(function) } +} + +// --------------------------------------------------------------------------- +// Attribute kinds and intrinsic ids (process-global in LLVM). + +struct Ids { + noreturn: u32, + nounwind: u32, + willreturn: u32, + memory: u32, + assume: u32, + deoptimize: u32, + memcpy_atomic: u32, + memmove_atomic: u32, + gc_relocate: u32, +} + +fn ids() -> &'static Ids { + static IDS: OnceLock = OnceLock::new(); + IDS.get_or_init(|| unsafe { + let attr = |n: &str| LLVMGetEnumAttributeKindForName(n.as_ptr().cast(), n.len()); + let intr = |n: &str| LLVMLookupIntrinsicID(n.as_ptr().cast(), n.len()); + Ids { + noreturn: attr("noreturn"), + nounwind: attr("nounwind"), + willreturn: attr("willreturn"), + memory: attr("memory"), + assume: intr("llvm.assume"), + deoptimize: intr("llvm.experimental.deoptimize"), + memcpy_atomic: intr("llvm.memcpy.element.unordered.atomic"), + memmove_atomic: intr("llvm.memmove.element.unordered.atomic"), + gc_relocate: intr("llvm.experimental.gc.relocate"), + } + }) +} + +// --------------------------------------------------------------------------- +// Small helpers over the C API. + +/// Pointer-identity hashing for `LLVMValueRef`/`LLVMBasicBlockRef` keys. +#[derive(Default)] +struct PtrHasher(u64); + +impl Hasher for PtrHasher { + fn finish(&self) -> u64 { + self.0 + } + fn write(&mut self, bytes: &[u8]) { + for &b in bytes { + self.0 = (self.0.rotate_left(8) ^ u64::from(b)).wrapping_mul(0x9E37_79B9_7F4A_7C15); + } + } + fn write_usize(&mut self, v: usize) { + self.0 = ((v as u64) >> 3).wrapping_mul(0x9E37_79B9_7F4A_7C15); + } +} + +type PtrMap = HashMap>; + +unsafe fn is_gc_type(ty: LLVMTypeRef) -> bool { + match LLVMGetTypeKind(ty) { + LLVMTypeKind::LLVMPointerTypeKind => LLVMGetPointerAddressSpace(ty) == 1, + LLVMTypeKind::LLVMVectorTypeKind | LLVMTypeKind::LLVMScalableVectorTypeKind => { + let e = LLVMGetElementType(ty); + LLVMGetTypeKind(e) == LLVMTypeKind::LLVMPointerTypeKind + && LLVMGetPointerAddressSpace(e) == 1 + } + _ => false, + } +} + +unsafe fn called_function(call: LLVMValueRef) -> LLVMValueRef { + LLVMIsAFunction(LLVMGetCalledValue(call)) +} + +/// `CallBase::hasFnAttr(kind)`: on the call site or on the called function. +unsafe fn has_fn_attr(call: LLVMValueRef, kind: u32) -> bool { + if !LLVMGetCallSiteEnumAttribute(call, llvm_sys::LLVMAttributeFunctionIndex, kind).is_null() { + return true; + } + let f = called_function(call); + !f.is_null() + && !LLVMGetEnumAttributeAtIndex(f, llvm_sys::LLVMAttributeFunctionIndex, kind).is_null() +} + +unsafe fn has_fn_string_attr(call: LLVMValueRef, name: &str) -> bool { + let idx = llvm_sys::LLVMAttributeFunctionIndex; + let (p, n) = (name.as_ptr().cast(), name.len() as u32); + if !LLVMGetCallSiteStringAttribute(call, idx, p, n).is_null() { + return true; + } + let f = called_function(call); + !f.is_null() && !LLVMGetStringAttributeAtIndex(f, idx, p, n).is_null() +} + +unsafe fn value_name<'a>(v: LLVMValueRef) -> &'a [u8] { + let mut len = 0usize; + let p = LLVMGetValueName2(v, &mut len); + if p.is_null() { + &[] + } else { + std::slice::from_raw_parts(p.cast(), len) + } +} + +/// A null pointer in address space 0 (where LLVM treats null as undefined). +unsafe fn is_null_as0(v: LLVMValueRef) -> bool { + let ty = LLVMTypeOf(v); + LLVMGetTypeKind(ty) == LLVMTypeKind::LLVMPointerTypeKind + && LLVMGetPointerAddressSpace(ty) == 0 + && LLVMIsConstant(v) != 0 + && LLVMIsNull(v) != 0 +} + +unsafe fn is_false_or_undef(v: LLVMValueRef) -> bool { + LLVMIsUndef(v) != 0 || (!LLVMIsAConstantInt(v).is_null() && LLVMConstIntGetZExtValue(v) == 0) +} + +/// `CallBase::getMemoryEffects().onlyReadsMemory()`: the `memory` attribute +/// packs two bits per location, `Ref` = 1 and `Mod` = 2. +unsafe fn only_reads_memory(call: LLVMValueRef) -> bool { + let ids = ids(); + let unknown = u64::MAX; + let read = |a: llvm_sys::prelude::LLVMAttributeRef| { + if a.is_null() { + unknown + } else { + LLVMGetEnumAttributeValue(a) + } + }; + let site = read(LLVMGetCallSiteEnumAttribute( + call, + llvm_sys::LLVMAttributeFunctionIndex, + ids.memory, + )); + let f = called_function(call); + let callee = if f.is_null() { + unknown + } else { + read(LLVMGetEnumAttributeAtIndex( + f, + llvm_sys::LLVMAttributeFunctionIndex, + ids.memory, + )) + }; + (site & callee) & 0xAAAA_AAAA_AAAA_AAAA == 0 +} + +/// Personalities that catch asynchronous exceptions keep `invoke`s of +/// `nounwind` callees (`canSimplifyInvokeNoUnwind`). +unsafe fn can_simplify_invoke_nounwind(function: LLVMValueRef) -> bool { + if LLVMHasPersonalityFn(function) == 0 { + return true; + } + let p = LLVMGetPersonalityFn(function); + let name = value_name(p); + !matches!( + name, + b"_except_handler3" | b"_except_handler4" | b"__C_specific_handler" + ) +} + +/// A compare whose only use is the branch being folded: deleted with it. +unsafe fn trivially_dead_after_branch(cond: LLVMValueRef) -> bool { + if LLVMIsAInstruction(cond).is_null() + || !matches!( + LLVMGetInstructionOpcode(cond), + LLVMOpcode::LLVMICmp | LLVMOpcode::LLVMFCmp + ) + { + return false; + } + let first = LLVMGetFirstUse(cond); + !first.is_null() && LLVMGetNextUse(first).is_null() +} + +/// RS4GC's `NeedsRewrite` for a call site that survives its CFG cleanup. +unsafe fn needs_statepoint(call: LLVMValueRef) -> bool { + let ids = ids(); + if !LLVMIsAInlineAsm(LLVMGetCalledValue(call)).is_null() { + return false; + } + if has_fn_string_attr(call, "gc-leaf-function") { + return false; + } + let f = called_function(call); + if !f.is_null() { + let iid = LLVMGetIntrinsicID(f); + if iid != 0 { + // Most intrinsics are leaves. `gc.statepoint` itself is skipped + // by `NeedsRewrite`; these three are not leaves. + return iid == ids.deoptimize || iid == ids.memcpy_atomic || iid == ids.memmove_atomic; + } + if is_libfunc(value_name(f)) { + return false; + } + } + true +} + +fn is_libfunc(name: &[u8]) -> bool { + LIBFUNCS + .binary_search_by(|probe| probe.as_bytes().cmp(name)) + .is_ok() +} + +// --------------------------------------------------------------------------- +// The analysis. + +const LIVE_OUT: u32 = u32::MAX; + +/// What RS4GC's cleanup makes of a block's terminating `invoke`. +#[derive(Clone, Copy, PartialEq, Eq)] +enum InvokeFate { + /// Not an invoke (or removed as unreachable). + None, + /// Still a safepoint, relocated on this many edges. + Edges(u8), + /// A side-effect-free `nounwind` invoke whose result is unused: deleted. + Erased, +} + +unsafe fn analyze(function: LLVMValueRef) -> FunctionLiveness { + let ids = ids(); + let mut out = FunctionLiveness::default(); + + // Blocks and instructions, CSR by block. + let mut blocks: Vec = Vec::new(); + let mut bb = LLVMGetFirstBasicBlock(function); + while !bb.is_null() { + blocks.push(bb); + bb = LLVMGetNextBasicBlock(bb); + } + let nb = blocks.len(); + if nb == 0 { + return out; + } + let bidx: PtrMap = blocks + .iter() + .enumerate() + .map(|(i, &b)| (b, i as u32)) + .collect(); + let mut insts: Vec = Vec::new(); + let mut start: Vec = Vec::with_capacity(nb + 1); + for &b in &blocks { + start.push(insts.len()); + let mut i = LLVMGetFirstInstruction(b); + while !i.is_null() { + insts.push(i); + i = LLVMGetNextInstruction(i); + } + } + start.push(insts.len()); + out.instructions = insts.len(); + out.blocks = nb; + + // ---- markAliveBlocks: surviving prefix of each block, and its + // successors after constant folding (sorted, deduplicated). + let simplify_nounwind = can_simplify_invoke_nounwind(function); + let mut live_len = vec![0u32; nb]; + let mut fate = vec![InvokeFate::None; nb]; + let mut succ_start: Vec = Vec::with_capacity(nb + 1); + let mut succs: Vec = Vec::new(); + // Conditions deleted with a branch whose two targets are equal. + let mut deleted: PtrMap = PtrMap::default(); + // Conditional branches that survive: (block, condition) — the `icmp` + // move candidates. + let mut cond_branches: Vec<(u32, LLVMValueRef)> = Vec::new(); + for b in 0..nb { + succ_start.push(succs.len()); + let (s, e) = (start[b], start[b + 1]); + let mut cut: Option = None; + for k in s..e { + let i = insts[k]; + match LLVMGetInstructionOpcode(i) { + LLVMOpcode::LLVMCall => { + let callee = LLVMGetCalledValue(i); + let f = LLVMIsAFunction(callee); + if !f.is_null() { + if LLVMGetIntrinsicID(f) == ids.assume + && is_false_or_undef(LLVMGetOperand(i, 0)) + { + cut = Some(k - s); + break; + } + } else if is_null_as0(callee) || LLVMIsUndef(callee) != 0 { + cut = Some(k - s); + break; + } + if has_fn_attr(i, ids.noreturn) + && LLVMGetTailCallKind(i) != LLVMTailCallKind::LLVMTailCallKindMustTail + { + cut = Some(k - s + 1); + break; + } + } + LLVMOpcode::LLVMStore => { + if LLVMGetVolatile(i) == 0 { + let p = LLVMGetOperand(i, 1); + if LLVMIsUndef(p) != 0 || is_null_as0(p) { + cut = Some(k - s); + break; + } + } + } + _ => {} + } + } + if let Some(c) = cut { + live_len[b] = c as u32; + continue; + } + live_len[b] = (e - s) as u32; + if e == s { + continue; + } + let term = insts[e - 1]; + let first = succs.len(); + let push = |succs: &mut Vec, target: LLVMBasicBlockRef| { + succs.push(bidx[&target]); + }; + match LLVMGetInstructionOpcode(term) { + LLVMOpcode::LLVMInvoke => { + let callee = LLVMGetCalledValue(term); + if LLVMIsAFunction(callee).is_null() + && (is_null_as0(callee) || LLVMIsUndef(callee) != 0) + { + live_len[b] -= 1; + continue; + } + let noreturn = has_fn_attr(term, ids.noreturn); + let nounwind = simplify_nounwind && has_fn_attr(term, ids.nounwind); + if nounwind + && LLVMGetFirstUse(term).is_null() + && has_fn_attr(term, ids.willreturn) + && only_reads_memory(term) + { + fate[b] = InvokeFate::Erased; + } else { + fate[b] = InvokeFate::Edges(if nounwind { 1 } else { 2 }); + } + if !noreturn { + push(&mut succs, LLVMGetNormalDest(term)); + } + if !nounwind { + push(&mut succs, LLVMGetUnwindDest(term)); + } + } + LLVMOpcode::LLVMBr => { + if LLVMIsConditional(term) != 0 { + let cond = LLVMGetCondition(term); + let (t, f) = (LLVMGetSuccessor(term, 0), LLVMGetSuccessor(term, 1)); + if !LLVMIsAConstantInt(cond).is_null() { + push( + &mut succs, + if LLVMConstIntGetZExtValue(cond) != 0 { + t + } else { + f + }, + ); + } else if t == f { + // ConstantFoldTerminator(DeleteDeadConditions): the + // condition goes too when the branch was its only use. + push(&mut succs, t); + if trivially_dead_after_branch(cond) { + deleted.insert(cond, ()); + } + } else { + push(&mut succs, t); + push(&mut succs, f); + if !LLVMIsAInstruction(cond).is_null() + && LLVMGetInstructionOpcode(cond) == LLVMOpcode::LLVMICmp + { + cond_branches.push((b as u32, cond)); + } + } + } else { + push(&mut succs, LLVMGetSuccessor(term, 0)); + } + } + LLVMOpcode::LLVMSwitch => { + let cond = LLVMGetOperand(term, 0); + let n = LLVMGetNumSuccessors(term); + if !LLVMIsAConstantInt(cond).is_null() { + let mut dest = LLVMGetSuccessor(term, 0); + for c in 1..n { + if LLVMGetOperand(term, 2 * c) == cond { + dest = LLVMGetSuccessor(term, c); + break; + } + } + push(&mut succs, dest); + } else { + for c in 0..n { + push(&mut succs, LLVMGetSuccessor(term, c)); + } + } + } + _ => { + for c in 0..LLVMGetNumSuccessors(term) { + push(&mut succs, LLVMGetSuccessor(term, c)); + } + } + } + let tail = &mut succs[first..]; + tail.sort_unstable(); + let mut w = first; + for r in first..succs.len() { + if r == first || succs[r] != succs[w - 1] { + succs[w] = succs[r]; + w += 1; + } + } + succs.truncate(w); + } + succ_start.push(succs.len()); + let succ_of = |b: usize| &succs[succ_start[b]..succ_start[b + 1]]; + + // Reachability from the entry over the surviving edges. + let mut reach = vec![false; nb]; + let mut stack = vec![0usize]; + reach[0] = true; + while let Some(b) = stack.pop() { + for &s in succ_of(b) { + if !reach[s as usize] { + reach[s as usize] = true; + stack.push(s as usize); + } + } + } + // Predecessors, restricted to reachable blocks and surviving edges. + let mut pred_count = vec![0usize; nb + 1]; + for b in (0..nb).filter(|&b| reach[b]) { + for &s in succ_of(b) { + pred_count[s as usize + 1] += 1; + } + } + for b in 0..nb { + pred_count[b + 1] += pred_count[b]; + } + let pred_start = pred_count; + let mut preds = vec![0u32; pred_start[nb]]; + let mut fill = pred_start.clone(); + for b in (0..nb).filter(|&b| reach[b]) { + for &s in succ_of(b) { + preds[fill[s as usize]] = b as u32; + fill[s as usize] += 1; + } + } + let pred_of = |b: usize| &preds[pred_start[b]..pred_start[b + 1]]; + let has_edge = |p: usize, b: u32| succ_of(p).binary_search(&b).is_ok(); + + // ---- phi folding: removePredecessor's `hasConstantValue` fold for phis + // that lost an edge, then FoldSingleEntryPHINodes. A null alias target + // stands for `poison` (a phi of only itself). + let mut alias: PtrMap = PtrMap::default(); + let mut vals: Vec = Vec::new(); + for b in (0..nb).filter(|&b| reach[b]) { + let unique_pred = pred_of(b).len() == 1; + for k in start[b]..start[b] + live_len[b] as usize { + let phi = insts[k]; + if LLVMGetInstructionOpcode(phi) != LLVMOpcode::LLVMPHI { + break; + } + vals.clear(); + let mut lost = false; + for e in 0..LLVMCountIncoming(phi) { + let p = bidx[&LLVMGetIncomingBlock(phi, e)] as usize; + if reach[p] && has_edge(p, b as u32) { + vals.push(LLVMGetIncomingValue(phi, e)); + } else { + lost = true; + } + } + if vals.is_empty() { + continue; + } + if unique_pred { + let v = vals[0]; + alias.insert(phi, if v == phi { std::ptr::null_mut() } else { v }); + } else if lost { + let mut c = vals[0]; + let mut agree = true; + for &v in &vals[1..] { + if v != c && v != phi { + if c != phi { + agree = false; + break; + } + c = v; + } + } + if agree { + alias.insert(phi, if c == phi { std::ptr::null_mut() } else { c }); + } + } + } + } + let rep = |mut v: LLVMValueRef| -> LLVMValueRef { + for _ in 0..64 { + match alias.get(&v) { + Some(&n) if !n.is_null() => v = n, + Some(_) => return std::ptr::null_mut(), + None => return v, + } + } + v + }; + + // ---- GC values: arguments, then surviving instructions. + let mut id_of: PtrMap = PtrMap::default(); + let mut def_block: Vec = Vec::new(); + let mut def_pos: Vec = Vec::new(); + let mut maybe_derived: Vec = Vec::new(); + for p in 0..LLVMCountParams(function) { + let a = LLVMGetParam(function, p); + if is_gc_type(LLVMTypeOf(a)) { + id_of.insert(a, def_block.len() as u32); + def_block.push(0); + def_pos.push(-1); + maybe_derived.push(false); + } + } + let mut phi_or_select: Vec = Vec::new(); + for b in (0..nb).filter(|&b| reach[b]) { + for k in 0..live_len[b] as usize { + let i = insts[start[b] + k]; + if !is_gc_type(LLVMTypeOf(i)) || alias.contains_key(&i) { + continue; + } + let id = def_block.len() as u32; + id_of.insert(i, id); + def_block.push(b as u32); + def_pos.push(k as i64); + let op = LLVMGetInstructionOpcode(i); + let derived = matches!( + op, + LLVMOpcode::LLVMGetElementPtr + | LLVMOpcode::LLVMAddrSpaceCast + | LLVMOpcode::LLVMBitCast + | LLVMOpcode::LLVMFreeze + ); + if derived { + out.derived_values += 1; + } + if matches!(op, LLVMOpcode::LLVMPHI | LLVMOpcode::LLVMSelect) { + phi_or_select.push(id); + } + maybe_derived.push(derived); + } + } + let nv = def_block.len(); + // RS4GC's base pointers (`findBasePointer`) for phis and selects. + let base = if phi_or_select.is_empty() { + Vec::new() + } else { + phi_select_bases( + &phi_or_select, + nv, + &insts, + &start, + &def_block, + &def_pos, + &id_of, + &bidx, + &reach, + &|p, b| has_edge(p, b), + &rep, + ) + }; + // How many relocations one crossing of each value costs: 0 when its base + // is a constant (RS4GC drops it from the live set), 2 when RS4GC gives it + // a fresh base phi/select (live exactly where it is), else 1. A value + // whose base is some *other* existing value costs 1 plus at most 1 more + // (that base may already be live): it is counted in the bound. + let mut weight = vec![1u8; nv]; + for (v, kind) in base.iter().enumerate() { + match kind { + BaseKind::Own => {} + BaseKind::Constant => weight[v] = 0, + BaseKind::Fresh => weight[v] = 2, + BaseKind::Existing => maybe_derived[v] = true, + } + } + out.fresh_bases = base.iter().filter(|k| **k == BaseKind::Fresh).count(); + out.gc_values = nv; + + // ---- the `icmp` move: a single-use icmp feeding a surviving conditional + // branch is moved to just before that branch. + let mut moved: PtrMap = PtrMap::default(); + if !cond_branches.is_empty() { + let mut uses: PtrMap = + cond_branches.iter().map(|&(_, c)| (c, 0)).collect(); + for b in (0..nb).filter(|&b| reach[b]) { + for k in 0..live_len[b] as usize { + let i = insts[start[b] + k]; + if LLVMGetInstructionOpcode(i) == LLVMOpcode::LLVMPHI { + if alias.contains_key(&i) { + continue; + } + for e in 0..LLVMCountIncoming(i) { + let p = bidx[&LLVMGetIncomingBlock(i, e)] as usize; + if reach[p] && has_edge(p, b as u32) { + if let Some(n) = uses.get_mut(&LLVMGetIncomingValue(i, e)) { + *n += 1; + } + } + } + continue; + } + for o in 0..LLVMGetNumOperands(i) as u32 { + if let Some(n) = uses.get_mut(&LLVMGetOperand(i, o)) { + *n += 1; + } + } + } + } + for &(b, c) in &cond_branches { + if uses.get(&c) == Some(&1) { + moved.insert(c, (b, live_len[b as usize] - 1)); + } + } + } + + // ---- uses of GC values: (value, block, position or LIVE_OUT). + let mut uses: Vec<(u32, u32, u32)> = Vec::new(); + for b in (0..nb).filter(|&b| reach[b]) { + for k in 0..live_len[b] as usize { + let i = insts[start[b] + k]; + if LLVMGetInstructionOpcode(i) == LLVMOpcode::LLVMPHI { + if alias.contains_key(&i) || !is_gc_type(LLVMTypeOf(i)) { + continue; + } + for e in 0..LLVMCountIncoming(i) { + let p = bidx[&LLVMGetIncomingBlock(i, e)] as usize; + if !(reach[p] && has_edge(p, b as u32)) { + continue; + } + let v = rep(LLVMGetIncomingValue(i, e)); + if let Some(&id) = id_of.get(&v) { + uses.push((id, p as u32, LIVE_OUT)); + } + } + continue; + } + if deleted.contains_key(&i) + || (k + 1 == live_len[b] as usize && fate[b] == InvokeFate::Erased) + { + continue; + } + let (ub, up) = moved.get(&i).copied().unwrap_or((b as u32, k as u32)); + for o in 0..LLVMGetNumOperands(i) as u32 { + let v = LLVMGetOperand(i, o); + if v.is_null() || !is_gc_type(LLVMTypeOf(v)) { + continue; + } + if let Some(&id) = id_of.get(&rep(v)) { + uses.push((id, ub, up)); + } + } + } + } + // CSR by value. + let mut use_start = vec![0usize; nv + 1]; + for &(v, _, _) in &uses { + use_start[v as usize + 1] += 1; + } + for v in 0..nv { + use_start[v + 1] += use_start[v]; + } + let mut by_value = vec![(0u32, 0u32); uses.len()]; + let mut fill = use_start.clone(); + for &(v, b, p) in &uses { + by_value[fill[v as usize]] = (b, p); + fill[v as usize] += 1; + } + drop(uses); + + // ---- safepoints, CSR by block, positions ascending. + let mut sp_start = vec![0usize; nb + 1]; + let mut sp_pos: Vec = Vec::new(); + for b in 0..nb { + sp_start[b] = sp_pos.len(); + if !reach[b] { + continue; + } + let n = live_len[b] as usize; + for k in 0..n { + let i = insts[start[b] + k]; + let edges = match LLVMGetInstructionOpcode(i) { + LLVMOpcode::LLVMCall | LLVMOpcode::LLVMCallBr => 1u8, + LLVMOpcode::LLVMInvoke => match fate[b] { + InvokeFate::Edges(e) if k + 1 == n => e, + _ => continue, + }, + _ => continue, + }; + if needs_statepoint(i) { + sp_pos.push(k as u32); + out.safepoints.push(SafepointLiveness { + block: b as u32, + index: k as u32, + edges, + live: 0, + }); + } + } + } + sp_start[nb] = sp_pos.len(); + let nsp = sp_pos.len(); + if nsp == 0 { + return out; + } + + // ---- per-value backward liveness walk, accumulated into difference + // arrays over the global safepoint order. + let mut diff = vec![0i64; nsp + 1]; + // Existing phi/select bases can need a second relocation without a GEP or cast. + let mut derived_diff = if maybe_derived.iter().any(|&derived| derived) { + vec![0i64; nsp + 1] + } else { + Vec::new() + }; + let mut in_stamp = vec![0u32; nb]; + let mut out_stamp = vec![0u32; nb]; + let mut lu_stamp = vec![0u32; nb]; + let mut lu = vec![0u32; nb]; + let mut work: Vec = Vec::new(); + let mut in_blocks: Vec = Vec::new(); + let mut steps = 0u64; + for v in 0..nv { + let (u0, u1) = (use_start[v], use_start[v + 1]); + if u0 == u1 { + continue; + } + let stamp = v as u32 + 1; + let defb = def_block[v] as usize; + work.clear(); + in_blocks.clear(); + steps += (u1 - u0) as u64; + let mut mark_in = |b: usize, work: &mut Vec, in_blocks: &mut Vec| { + if in_stamp[b] != stamp { + in_stamp[b] = stamp; + work.push(b as u32); + in_blocks.push(b as u32); + } + }; + for &(ub, up) in &by_value[u0..u1] { + let ub = ub as usize; + if up == LIVE_OUT { + if out_stamp[ub] != stamp { + out_stamp[ub] = stamp; + if ub != defb { + mark_in(ub, &mut work, &mut in_blocks); + } + } + } else { + if lu_stamp[ub] != stamp || up > lu[ub] { + lu_stamp[ub] = stamp; + lu[ub] = up; + } + if ub != defb { + mark_in(ub, &mut work, &mut in_blocks); + } + } + } + while let Some(b) = work.pop() { + for &p in pred_of(b as usize) { + steps += 1; + let p = p as usize; + if out_stamp[p] != stamp { + out_stamp[p] = stamp; + if p != defb { + mark_in(p, &mut work, &mut in_blocks); + } + } + } + } + steps += in_blocks.len() as u64 + 1; + let derived = maybe_derived[v]; + let weight = i64::from(weight[v]); + if weight == 0 { + continue; + } + let mut add = |b: usize, from: i64, to: Option| { + let sps = &sp_pos[sp_start[b]..sp_start[b + 1]]; + if sps.is_empty() { + return; + } + let lo = if from < 0 { + 0 + } else { + sps.partition_point(|&p| i64::from(p) <= from) + }; + let hi = match to { + None => sps.len(), + // A use by the safepoint itself keeps the value live across + // it: RS4GC takes the live set *into* the call, operands + // included (`findLiveSetAtInst` walks the call itself). + Some(end) => sps.partition_point(|&p| p <= end), + }; + if hi > lo { + let g = sp_start[b]; + diff[g + lo] += weight; + diff[g + hi] -= weight; + if derived { + derived_diff[g + lo] += 1; + derived_diff[g + hi] -= 1; + } + } + }; + // Definition block: from the definition to the last use, or through + // the end when the value is live out. + if out_stamp[defb] == stamp { + add(defb, def_pos[v], None); + } else if lu_stamp[defb] == stamp { + add(defb, def_pos[v], Some(lu[defb])); + } + for &b in &in_blocks { + let b = b as usize; + if out_stamp[b] == stamp { + add(b, -1, None); + } else if lu_stamp[b] == stamp { + add(b, -1, Some(lu[b])); + } + } + } + out.work = steps; + + let mut live = 0i64; + let mut dlive = 0i64; + for (g, sp) in out.safepoints.iter_mut().enumerate() { + live += diff[g]; + sp.live = live as u32; + out.relocations += live as u64 * u64::from(sp.edges); + if !derived_diff.is_empty() { + dlive += derived_diff[g]; + out.derived_crossings += dlive as u64 * u64::from(sp.edges); + } + } + out +} + +/// RS4GC's base of a phi/select value. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum BaseKind { + /// Its own base (every input is a base, or a `null`): relocated once. + Own, + /// All inputs are constants: the base is `null`, and RS4GC drops the + /// value from the live set. + Constant, + /// Inputs disagree: RS4GC clones it into a `.base` phi/select that is + /// live wherever it is. + Fresh, + /// Every input agrees on one other existing value. + Existing, +} + +/// One input of a phi/select, as `findBasePointer` sees it. +#[derive(Clone, Copy)] +enum BaseInput { + /// The node itself. + Myself, + /// `null`: its own base, so it does not block pruning. + Null, + /// Any other constant: base `null`, but not its own BDV, so it blocks + /// pruning (an `inttoptr` of a NaN-box tag, typically). + Constant, + /// A value that is its own base (call, load, argument, `inttoptr`...). + Base(usize), + /// A derived value: base elsewhere, and it blocks pruning. + Derived(usize), + /// Another phi/select, by node index. + Node(u32), +} + +#[derive(Clone, Copy, PartialEq, Eq)] +enum BaseState { + Unknown, + Base(usize), + Conflict, +} + +impl BaseState { + fn meet(self, other: BaseState) -> BaseState { + match (self, other) { + (BaseState::Unknown, x) | (x, BaseState::Unknown) => x, + (BaseState::Base(a), BaseState::Base(b)) if a == b => self, + _ => BaseState::Conflict, + } + } +} + +/// `findBasePointer` for every phi/select GC value: prune the nodes whose +/// inputs are all bases, then run the optimistic meet over the rest. Linear +/// in the phi/select inputs (each node's state changes at most twice). +#[allow(clippy::too_many_arguments)] +unsafe fn phi_select_bases( + nodes: &[u32], + nv: usize, + insts: &[LLVMValueRef], + start: &[usize], + def_block: &[u32], + def_pos: &[i64], + id_of: &PtrMap, + bidx: &PtrMap, + reach: &[bool], + has_edge: &dyn Fn(usize, u32) -> bool, + rep: &dyn Fn(LLVMValueRef) -> LLVMValueRef, +) -> Vec { + let mut node_of = vec![u32::MAX; nv]; + for (n, &id) in nodes.iter().enumerate() { + node_of[id as usize] = n as u32; + } + let classify = |me: u32, v: LLVMValueRef| -> BaseInput { + if v.is_null() || LLVMIsConstant(v) != 0 { + return if !v.is_null() && LLVMIsNull(v) != 0 { + BaseInput::Null + } else { + BaseInput::Constant + }; + } + let Some(&id) = id_of.get(&v) else { + // Not a GC value we track (it cannot be one: phi inputs share + // the phi's type); treat it as a base of its own. + return BaseInput::Base(v as usize); + }; + let n = node_of[id as usize]; + if n == me { + BaseInput::Myself + } else if n != u32::MAX { + BaseInput::Node(n) + } else if !LLVMIsAInstruction(v).is_null() + && matches!( + LLVMGetInstructionOpcode(v), + LLVMOpcode::LLVMGetElementPtr + | LLVMOpcode::LLVMAddrSpaceCast + | LLVMOpcode::LLVMBitCast + | LLVMOpcode::LLVMFreeze + ) + { + BaseInput::Derived(v as usize) + } else { + BaseInput::Base(v as usize) + } + }; + // Inputs, CSR by node, and users (for the worklists). + let mut in_start = Vec::with_capacity(nodes.len() + 1); + let mut inputs: Vec = Vec::new(); + for (n, &id) in nodes.iter().enumerate() { + in_start.push(inputs.len()); + let b = def_block[id as usize] as usize; + let inst = insts[start[b] + def_pos[id as usize] as usize]; + if LLVMGetInstructionOpcode(inst) == LLVMOpcode::LLVMPHI { + for e in 0..LLVMCountIncoming(inst) { + let p = bidx[&LLVMGetIncomingBlock(inst, e)] as usize; + if reach[p] && has_edge(p, b as u32) { + inputs.push(classify(n as u32, rep(LLVMGetIncomingValue(inst, e)))); + } + } + } else { + for o in [1, 2] { + inputs.push(classify(n as u32, rep(LLVMGetOperand(inst, o)))); + } + } + } + in_start.push(inputs.len()); + let mut users: Vec> = vec![Vec::new(); nodes.len()]; + for n in 0..nodes.len() { + for input in &inputs[in_start[n]..in_start[n + 1]] { + if let BaseInput::Node(m) = *input { + users[m as usize].push(n as u32); + } + } + } + + // Pruning: a node is its own base when every input is itself, `null`, + // a base, or an already-pruned node. + let mut live = vec![true; nodes.len()]; + let prunable = |n: usize, live: &[bool]| { + inputs[in_start[n]..in_start[n + 1]] + .iter() + .all(|input| match *input { + BaseInput::Myself | BaseInput::Null | BaseInput::Base(_) => true, + BaseInput::Node(m) => !live[m as usize], + BaseInput::Constant | BaseInput::Derived(_) => false, + }) + }; + let mut work: Vec = (0..nodes.len() as u32).collect(); + while let Some(n) = work.pop() { + let n = n as usize; + if live[n] && prunable(n, &live) { + live[n] = false; + work.extend(users[n].iter().copied().filter(|&u| live[u as usize])); + } + } + + // Optimistic meet over the nodes that remain. + let mut state = vec![BaseState::Unknown; nodes.len()]; + let mut work: Vec = (0..nodes.len() as u32) + .filter(|&n| live[n as usize]) + .collect(); + while let Some(n) = work.pop() { + let n = n as usize; + let mut s = BaseState::Unknown; + for input in &inputs[in_start[n]..in_start[n + 1]] { + let x = match *input { + BaseInput::Myself => BaseState::Unknown, + BaseInput::Null | BaseInput::Constant => BaseState::Base(0), + BaseInput::Base(p) | BaseInput::Derived(p) => BaseState::Base(p), + BaseInput::Node(m) if live[m as usize] => state[m as usize], + // A pruned node is its own base. Odd keys cannot collide + // with the (aligned) addresses of other bases, nor with 0. + BaseInput::Node(m) => BaseState::Base(((nodes[m as usize] as usize + 1) << 1) | 1), + }; + s = s.meet(x); + } + if s != state[n] { + state[n] = s; + work.extend(users[n].iter().copied().filter(|&u| live[u as usize])); + } + } + + let mut kind = vec![BaseKind::Own; nv]; + for (n, &id) in nodes.iter().enumerate() { + if !live[n] { + continue; + } + kind[id as usize] = match state[n] { + BaseState::Unknown => BaseKind::Own, + BaseState::Conflict => BaseKind::Fresh, + BaseState::Base(0) => BaseKind::Constant, + BaseState::Base(_) => BaseKind::Existing, + }; + } + kind +} + +/// C library functions LLVM's `TargetLibraryInfo` recognizes (LLVM 22, +/// `llvm/include/llvm/Analysis/TargetLibraryInfo.td`, every +/// `TargetLibCall<"name", ...>`), sorted bytewise. RS4GC's +/// `callsGCLeafFunction` treats a call to any of them as a GC leaf. The +/// prototype check `TargetLibraryInfo` also applies is not repeated here: a +/// declaration with a library name but a foreign signature would be counted +/// as a leaf, which the audit would report as a mismatch. +#[rustfmt::skip] +const LIBFUNCS: &[&str] = &include!("gc_liveness_libfuncs.in"); + +#[cfg(test)] +#[path = "gc_liveness_tests.rs"] +mod tests; diff --git a/crates/perry-codegen/src/inprocess/gc_liveness_libfuncs.in b/crates/perry-codegen/src/inprocess/gc_liveness_libfuncs.in new file mode 100644 index 0000000000..61f5629be0 --- /dev/null +++ b/crates/perry-codegen/src/inprocess/gc_liveness_libfuncs.in @@ -0,0 +1,528 @@ +// Generated from LLVM 22 llvm/include/llvm/Analysis/TargetLibraryInfo.td: +// every TargetLibCall<"name", ...>, sorted bytewise. See LIBFUNCS in gc_liveness.rs. +[ + "??2@YAPAXI@Z", + "??2@YAPAXIABUnothrow_t@std@@@Z", + "??2@YAPEAX_K@Z", + "??2@YAPEAX_KAEBUnothrow_t@std@@@Z", + "??3@YAXPAX@Z", + "??3@YAXPAXABUnothrow_t@std@@@Z", + "??3@YAXPAXI@Z", + "??3@YAXPEAX@Z", + "??3@YAXPEAXAEBUnothrow_t@std@@@Z", + "??3@YAXPEAX_K@Z", + "??_U@YAPAXI@Z", + "??_U@YAPAXIABUnothrow_t@std@@@Z", + "??_U@YAPEAX_K@Z", + "??_U@YAPEAX_KAEBUnothrow_t@std@@@Z", + "??_V@YAXPAX@Z", + "??_V@YAXPAXABUnothrow_t@std@@@Z", + "??_V@YAXPAXI@Z", + "??_V@YAXPEAX@Z", + "??_V@YAXPEAXAEBUnothrow_t@std@@@Z", + "??_V@YAXPEAX_K@Z", + "_Exit", + "_IO_getc", + "_IO_putc", + "_ZSt9terminatev", + "_ZdaPv", + "_ZdaPvRKSt9nothrow_t", + "_ZdaPvSt11align_val_t", + "_ZdaPvSt11align_val_tRKSt9nothrow_t", + "_ZdaPvj", + "_ZdaPvjSt11align_val_t", + "_ZdaPvm", + "_ZdaPvmSt11align_val_t", + "_ZdlPv", + "_ZdlPvRKSt9nothrow_t", + "_ZdlPvSt11align_val_t", + "_ZdlPvSt11align_val_tRKSt9nothrow_t", + "_ZdlPvj", + "_ZdlPvjSt11align_val_t", + "_ZdlPvm", + "_ZdlPvmSt11align_val_t", + "_Znaj", + "_ZnajRKSt9nothrow_t", + "_ZnajSt11align_val_t", + "_ZnajSt11align_val_tRKSt9nothrow_t", + "_Znam", + "_Znam12__hot_cold_t", + "_ZnamRKSt9nothrow_t", + "_ZnamRKSt9nothrow_t12__hot_cold_t", + "_ZnamSt11align_val_t", + "_ZnamSt11align_val_t12__hot_cold_t", + "_ZnamSt11align_val_tRKSt9nothrow_t", + "_ZnamSt11align_val_tRKSt9nothrow_t12__hot_cold_t", + "_Znwj", + "_ZnwjRKSt9nothrow_t", + "_ZnwjSt11align_val_t", + "_ZnwjSt11align_val_tRKSt9nothrow_t", + "_Znwm", + "_Znwm12__hot_cold_t", + "_ZnwmRKSt9nothrow_t", + "_ZnwmRKSt9nothrow_t12__hot_cold_t", + "_ZnwmSt11align_val_t", + "_ZnwmSt11align_val_t12__hot_cold_t", + "_ZnwmSt11align_val_tRKSt9nothrow_t", + "_ZnwmSt11align_val_tRKSt9nothrow_t12__hot_cold_t", + "__acos_finite", + "__acosf_finite", + "__acosh_finite", + "__acoshf_finite", + "__acoshl_finite", + "__acosl_finite", + "__asin_finite", + "__asinf_finite", + "__asinl_finite", + "__atan2_finite", + "__atan2f_finite", + "__atan2l_finite", + "__atanh_finite", + "__atanhf_finite", + "__atanhl_finite", + "__atomic_load", + "__atomic_store", + "__cosh_finite", + "__coshf_finite", + "__coshl_finite", + "__cospi", + "__cospif", + "__cxa_atexit", + "__cxa_guard_abort", + "__cxa_guard_acquire", + "__cxa_guard_release", + "__cxa_throw", + "__exp10_finite", + "__exp10f_finite", + "__exp10l_finite", + "__exp2_finite", + "__exp2f_finite", + "__exp2l_finite", + "__exp_finite", + "__expf_finite", + "__expl_finite", + "__isoc99_scanf", + "__isoc99_sscanf", + "__kmpc_alloc_shared", + "__kmpc_free_shared", + "__log10_finite", + "__log10f_finite", + "__log10l_finite", + "__log2_finite", + "__log2f_finite", + "__log2l_finite", + "__log_finite", + "__logf_finite", + "__logl_finite", + "__memccpy_chk", + "__memcpy_chk", + "__memmove_chk", + "__mempcpy_chk", + "__memset_chk", + "__nvvm_reflect", + "__pow_finite", + "__powf_finite", + "__powl_finite", + "__sincospi_stret", + "__sincospif_stret", + "__sinh_finite", + "__sinhf_finite", + "__sinhl_finite", + "__sinpi", + "__sinpif", + "__size_returning_new", + "__size_returning_new_aligned", + "__size_returning_new_aligned_hot_cold", + "__size_returning_new_hot_cold", + "__small_fprintf", + "__small_printf", + "__small_sprintf", + "__snprintf_chk", + "__sprintf_chk", + "__sqrt_finite", + "__sqrtf_finite", + "__sqrtl_finite", + "__stpcpy_chk", + "__stpncpy_chk", + "__strcat_chk", + "__strcpy_chk", + "__strdup", + "__strlcat_chk", + "__strlcpy_chk", + "__strlen_chk", + "__strncat_chk", + "__strncpy_chk", + "__strndup", + "__strtok_r", + "__vsnprintf_chk", + "__vsprintf_chk", + "abort", + "abs", + "access", + "acos", + "acosf", + "acosh", + "acoshf", + "acoshl", + "acosl", + "aligned_alloc", + "asin", + "asinf", + "asinh", + "asinhf", + "asinhl", + "asinl", + "atan", + "atan2", + "atan2f", + "atan2l", + "atanf", + "atanh", + "atanhf", + "atanhl", + "atanl", + "atexit", + "atof", + "atoi", + "atol", + "atoll", + "bcmp", + "bcopy", + "bzero", + "cabs", + "cabsf", + "cabsl", + "calloc", + "cbrt", + "cbrtf", + "cbrtl", + "ceil", + "ceilf", + "ceill", + "chmod", + "chown", + "clearerr", + "closedir", + "copysign", + "copysignf", + "copysignl", + "cos", + "cosf", + "cosh", + "coshf", + "coshl", + "cosl", + "ctermid", + "erf", + "erff", + "erfl", + "execl", + "execle", + "execlp", + "execv", + "execvP", + "execve", + "execvp", + "execvpe", + "exit", + "exp", + "exp10", + "exp10f", + "exp10l", + "exp2", + "exp2f", + "exp2l", + "expf", + "expl", + "expm1", + "expm1f", + "expm1l", + "fabs", + "fabsf", + "fabsl", + "fclose", + "fdim", + "fdimf", + "fdiml", + "fdopen", + "feof", + "ferror", + "fflush", + "ffs", + "ffsl", + "ffsll", + "fgetc", + "fgetc_unlocked", + "fgetpos", + "fgets", + "fgets_unlocked", + "fileno", + "fiprintf", + "flockfile", + "floor", + "floorf", + "floorl", + "fls", + "flsl", + "flsll", + "fmax", + "fmaxf", + "fmaximum_num", + "fmaximum_numf", + "fmaximum_numl", + "fmaxl", + "fmin", + "fminf", + "fminimum_num", + "fminimum_numf", + "fminimum_numl", + "fminl", + "fmod", + "fmodf", + "fmodl", + "fopen", + "fopen64", + "fork", + "fprintf", + "fputc", + "fputc_unlocked", + "fputs", + "fputs_unlocked", + "fread", + "fread_unlocked", + "free", + "frexp", + "frexpf", + "frexpl", + "fscanf", + "fseek", + "fseeko", + "fseeko64", + "fsetpos", + "fstat", + "fstat64", + "fstatvfs", + "fstatvfs64", + "ftell", + "ftello", + "ftello64", + "ftrylockfile", + "funlockfile", + "fwrite", + "fwrite_unlocked", + "getc", + "getc_unlocked", + "getchar", + "getchar_unlocked", + "getenv", + "getitimer", + "getlogin_r", + "getpwnam", + "gets", + "gettimeofday", + "htonl", + "htons", + "hypot", + "hypotf", + "hypotl", + "ilogb", + "ilogbf", + "ilogbl", + "iprintf", + "isascii", + "isdigit", + "labs", + "lchown", + "ldexp", + "ldexpf", + "ldexpl", + "llabs", + "log", + "log10", + "log10f", + "log10l", + "log1p", + "log1pf", + "log1pl", + "log2", + "log2f", + "log2l", + "logb", + "logbf", + "logbl", + "logf", + "logl", + "lstat", + "lstat64", + "malloc", + "memalign", + "memccpy", + "memchr", + "memcmp", + "memcpy", + "memmove", + "mempcpy", + "memrchr", + "memset", + "memset_pattern16", + "memset_pattern4", + "memset_pattern8", + "mkdir", + "mktime", + "modf", + "modff", + "modfl", + "nan", + "nanf", + "nanl", + "nearbyint", + "nearbyintf", + "nearbyintl", + "ntohl", + "ntohs", + "open", + "open64", + "opendir", + "pclose", + "perror", + "popen", + "posix_memalign", + "pow", + "powf", + "powl", + "pread", + "printf", + "putc", + "putc_unlocked", + "putchar", + "putchar_unlocked", + "puts", + "pvalloc", + "pwrite", + "qsort", + "read", + "readlink", + "realloc", + "reallocarray", + "reallocf", + "realpath", + "remainder", + "remainderf", + "remainderl", + "remove", + "remquo", + "remquof", + "remquol", + "rename", + "rewind", + "rint", + "rintf", + "rintl", + "rmdir", + "round", + "roundeven", + "roundevenf", + "roundevenl", + "roundf", + "roundl", + "scalbln", + "scalblnf", + "scalblnl", + "scalbn", + "scalbnf", + "scalbnl", + "scanf", + "setbuf", + "setitimer", + "setvbuf", + "sin", + "sincos", + "sincosf", + "sincosl", + "sinf", + "sinh", + "sinhf", + "sinhl", + "sinl", + "siprintf", + "snprintf", + "sprintf", + "sqrt", + "sqrtf", + "sqrtl", + "sscanf", + "stat", + "stat64", + "statvfs", + "statvfs64", + "stpcpy", + "stpncpy", + "strcasecmp", + "strcat", + "strchr", + "strcmp", + "strcoll", + "strcpy", + "strcspn", + "strdup", + "strlcat", + "strlcpy", + "strlen", + "strncasecmp", + "strncat", + "strncmp", + "strncpy", + "strndup", + "strnlen", + "strpbrk", + "strrchr", + "strspn", + "strstr", + "strtod", + "strtof", + "strtok", + "strtok_r", + "strtol", + "strtold", + "strtoll", + "strtoul", + "strtoull", + "strxfrm", + "system", + "tan", + "tanf", + "tanh", + "tanhf", + "tanhl", + "tanl", + "tgamma", + "tgammaf", + "tgammal", + "times", + "tmpfile", + "tmpfile64", + "toascii", + "trunc", + "truncf", + "truncl", + "uname", + "ungetc", + "unlink", + "unsetenv", + "utime", + "utimes", + "valloc", + "vec_calloc", + "vec_free", + "vec_malloc", + "vec_realloc", + "vfprintf", + "vfscanf", + "vprintf", + "vscanf", + "vsnprintf", + "vsprintf", + "vsscanf", + "wcslen", + "write", +] diff --git a/crates/perry-codegen/src/inprocess/gc_liveness_tests.rs b/crates/perry-codegen/src/inprocess/gc_liveness_tests.rs new file mode 100644 index 0000000000..f69f176bb4 --- /dev/null +++ b/crates/perry-codegen/src/inprocess/gc_liveness_tests.rs @@ -0,0 +1,634 @@ +//! Ground truth for the S4 liveness count: every fixture runs the shipped +//! split pipeline (prepare, count, rewrite) and asserts the prediction equals +//! the `gc.relocate`s RS4GC actually emitted, and a pinned number so a change +//! that moves both in step is still visible. + +use super::super::{default_cpu_for_triple, global_init, parse_ir_text}; +use super::*; +use inkwell::context::Context; +use inkwell::passes::PassBuilderOptions; +use inkwell::targets::{CodeModel, RelocMode, Target, TargetMachine, TargetTriple}; +use inkwell::values::AsValueRef; +use inkwell::OptimizationLevel; + +const TRIPLE: &str = "arm64-apple-darwin"; + +fn target_machine() -> TargetMachine { + global_init(&[]); + let triple = TargetTriple::create(TRIPLE); + Target::from_triple(&triple) + .expect("aarch64 target") + .create_target_machine( + &triple, + default_cpu_for_triple(TRIPLE), + "", + OptimizationLevel::None, + RelocMode::PIC, + CodeModel::Default, + ) + .expect("target machine") +} + +fn run(module: &inkwell::module::Module<'_>, tm: &TargetMachine, passes: &str) { + module + .run_passes(passes, tm, PassBuilderOptions::create()) + .unwrap_or_else(|e| panic!("`{passes}` failed: {e}")); +} + +/// The prediction for `name` on RS4GC's real input, and the `gc.relocate`s +/// RS4GC then really emitted for it. +fn census(ir: &str, name: &str) -> (FunctionLiveness, u64) { + let tm = target_machine(); + let context = Context::create(); + let module = parse_ir_text(&context, ir, "gc_liveness_fixture").expect("fixture parses"); + module.set_triple(&TargetTriple::create(TRIPLE)); + module.set_data_layout(&tm.get_target_data().get_data_layout()); + module.verify().expect("fixture verifies"); + run(&module, &tm, crate::linker::STATEPOINT_PREPARE_PASSES); + let f = module.get_function(name).expect("function survives"); + let predicted = analyze_function(f.as_value_ref()); + run(&module, &tm, crate::linker::STATEPOINT_REWRITE_ONLY_PASSES); + module.verify().expect("rewritten module verifies"); + let f = module.get_function(name).expect("function survives RS4GC"); + (predicted, count_relocates(f.as_value_ref())) +} + +/// Assert exactness against RS4GC and pin the count and the live sets. +fn assert_census(ir: &str, name: &str, relocations: u64, live: &[u32]) -> FunctionLiveness { + let (l, actual) = census(ir, name); + assert_eq!( + l.relocations, actual, + "`{name}`: predicted {} relocations, RS4GC emitted {actual}\n{l:#?}", + l.relocations + ); + assert!(l.is_exact(), "`{name}` has no derived pointers: {l:#?}"); + assert_eq!(l.relocations, relocations, "`{name}` pinned count: {l:#?}"); + let got: Vec = l.safepoints.iter().map(|s| s.live).collect(); + assert_eq!(got, live, "`{name}` per-safepoint live counts: {l:#?}"); + l +} + +const DECLS: &str = r#" +declare void @may_collect() +declare ptr addrspace(1) @alloc() +declare void @use(ptr addrspace(1)) +declare void @use2(ptr addrspace(1), ptr addrspace(1)) +declare i32 @__gxx_personality_v0(...) +"#; + +fn ir(body: &str) -> String { + format!("{DECLS}\n{body}") +} + +#[test] +fn straight_line_counts_values_live_after_each_call() { + // A call's own result is not live across it; its GC-pointer arguments + // are, even at their last use. + let l = assert_census( + &ir(r#" +define void @f() gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + %b = call ptr addrspace(1) @alloc() + call void @may_collect() + call void @use(ptr addrspace(1) %a) + call void @use(ptr addrspace(1) %b) + ret void +} +"#), + "f", + 6, + &[0, 1, 2, 2, 1], + ); + assert_eq!( + (l.statepoints(), l.invoke_statepoints(), l.max_live()), + (5, 0, 2) + ); +} + +#[test] +fn root_allocas_are_counted_as_the_ssa_values_mem2reg_makes() { + // The pre-mem2reg shape codegen emits: roots live in allocas. + assert_census( + &ir(r#" +define void @slots() gc "statepoint-example" { +entry: + %slot = alloca ptr addrspace(1) + %dead = alloca ptr addrspace(1) + %a = call ptr addrspace(1) @alloc() + store ptr addrspace(1) %a, ptr %slot + store ptr addrspace(1) null, ptr %dead + call void @may_collect() + %r = load ptr addrspace(1), ptr %slot + %d = load ptr addrspace(1), ptr %dead + call void @use2(ptr addrspace(1) %r, ptr addrspace(1) %d) + ret void +} +"#), + "slots", + // `%a` across may_collect and into `use2`; the null slot is a + // constant after mem2reg. + 2, + &[0, 1, 1], + ); +} + +#[test] +fn loops_and_phis_follow_the_cfg() { + assert_census( + &ir(r#" +define void @loop(i64 %n) gc "statepoint-example" { +entry: + %keep = call ptr addrspace(1) @alloc() + br label %head +head: + %i = phi i64 [ 0, %entry ], [ %i1, %body ] + %cur = phi ptr addrspace(1) [ %keep, %entry ], [ %next, %body ] + %c = icmp slt i64 %i, %n + br i1 %c, label %body, label %exit +body: + %next = call ptr addrspace(1) @alloc() + call void @may_collect() + %i1 = add i64 %i, 1 + br label %head +exit: + call void @use(ptr addrspace(1) %cur) + call void @use(ptr addrspace(1) %keep) + ret void +} +"#), + "loop", + // alloc in the loop: {keep}; may_collect: {keep, next}; the uses in + // exit: {cur, keep}, then {keep}. `%cur` is redefined at the header, + // so it is not live across the body's calls. + 6, + &[0, 1, 2, 2, 1], + ); +} + +const INVOKE: &str = r#" +define void @inv() gc "statepoint-example" personality ptr @__gxx_personality_v0 { +entry: + %a = call ptr addrspace(1) @alloc() + %b = call ptr addrspace(1) @alloc() + invoke void @may_collect() to label %ok unwind label %lp +ok: + call void @use(ptr addrspace(1) %a) + ret void +lp: + %l = landingpad token cleanup + call void @use(ptr addrspace(1) %b) + ret void +} +"#; + +#[test] +fn an_invoke_is_relocated_on_the_normal_and_the_unwind_edge() { + // `%a` is live into the normal destination, `%b` into the landing pad + // (Perry's retyped `landingpad token`); RS4GC relocates the union on BOTH + // edges: 2 values x 2 edges. + let l = assert_census(&ir(INVOKE), "inv", 1 + 4 + 1 + 1, &[0, 1, 2, 1, 1]); + assert_eq!(l.invoke_statepoints(), 1); + // Sabotage witness: a count that ignored the unwind edge (sum of live + // counts) would be 5, not RS4GC's 7 — this fixture tells them apart. + let naive: u64 = l.safepoints.iter().map(|s| u64::from(s.live)).sum(); + assert_eq!(naive, 5); + assert_ne!(naive, l.relocations); +} + +#[test] +fn an_invoke_of_a_nounwind_callee_becomes_a_call() { + // markAliveBlocks turns it into a call before RS4GC: one edge, and the + // landing pad becomes unreachable. + assert_census( + &ir(r#" +declare void @nothrow() nounwind +define void @nu() gc "statepoint-example" personality ptr @__gxx_personality_v0 { +entry: + %a = call ptr addrspace(1) @alloc() + %b = call ptr addrspace(1) @alloc() + invoke void @nothrow() to label %ok unwind label %lp +ok: + call void @use(ptr addrspace(1) %a) + ret void +lp: + %l = landingpad token cleanup + call void @use(ptr addrspace(1) %b) + ret void +} +"#), + "nu", + 1 + 1 + 1, + &[0, 1, 1, 1], + ); +} + +#[test] +fn leaf_calls_are_not_safepoints() { + // Call-site attribute, callee attribute, intrinsic, inline asm, and a C + // library function TargetLibraryInfo knows: none is a statepoint, so `%a` + // is live across exactly one call. + let l = assert_census( + &ir(r#" +declare void @leaf_decl() "gc-leaf-function" +declare void @plain() +declare i64 @strlen(ptr) +declare void @llvm.donothing() +define i64 @leafy(ptr %s) gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + call void @plain() "gc-leaf-function" + call void @leaf_decl() + %n = call i64 @strlen(ptr %s) + call void @llvm.donothing() + call void asm sideeffect "", ""() "gc-leaf-function" + call void @may_collect() + call void @use(ptr addrspace(1) %a) + ret i64 %n +} +"#), + "leafy", + 2, + &[0, 1, 1], + ); + assert_eq!(l.statepoints(), 3); +} + +#[test] +fn a_value_reloaded_from_its_global_after_a_safepoint_is_not_live_across_it() { + // S3's rematerialized form (after mem2reg/sccp fold its select): each + // read is a fresh load of the global below the safepoint. The control + // loads once and holds the value across both calls. + let remat = assert_census( + &ir(r#" +@G = global ptr addrspace(1) null +define void @remat() gc "statepoint-example" { +entry: + call void @may_collect() + %v1 = load ptr addrspace(1), ptr @G + call void @use(ptr addrspace(1) %v1) + call void @may_collect() + %v2 = load ptr addrspace(1), ptr @G + call void @use(ptr addrspace(1) %v2) + ret void +} +"#), + "remat", + // Only each load's own use: nothing is live across `may_collect`. + 2, + &[0, 1, 0, 1], + ); + assert_eq!(remat.statepoints(), 4); + assert_census( + &ir(r#" +@G = global ptr addrspace(1) null +define void @held() gc "statepoint-example" { +entry: + %v = load ptr addrspace(1), ptr @G + call void @may_collect() + call void @use(ptr addrspace(1) %v) + call void @may_collect() + call void @use(ptr addrspace(1) %v) + ret void +} +"#), + "held", + 4, + &[1, 1, 1, 1], + ); +} + +#[test] +fn code_after_a_noreturn_call_is_dead() { + assert_census( + &ir(r#" +declare void @die() noreturn +define void @nr(i1 %c) gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + br i1 %c, label %bad, label %good +bad: + call void @die() + call void @use(ptr addrspace(1) %a) + ret void +good: + call void @may_collect() + ret void +} +"#), + "nr", + // `die` is a safepoint, but nothing is live across it: the use after + // it is deleted as unreachable. + 0, + &[0, 0, 0], + ); +} + +#[test] +fn a_single_use_compare_is_moved_to_its_branch() { + // RS4GC sinks the icmp below the safepoint, so `%a` is live across it + // even though its last textual use is above. + assert_census( + &ir(r#" +define void @cmp() gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + %c = icmp eq ptr addrspace(1) %a, null + call void @may_collect() + br i1 %c, label %x, label %y +x: + ret void +y: + ret void +} +"#), + "cmp", + 1, + &[0, 1], + ); +} + +#[test] +fn a_single_entry_phi_is_folded_into_its_input() { + // Without the fold, `%a` and `%p` would count as two live values. + assert_census( + &ir(r#" +define void @one(i1 %c) gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + br label %next +next: + %p = phi ptr addrspace(1) [ %a, %entry ] + call void @may_collect() + call void @use2(ptr addrspace(1) %a, ptr addrspace(1) %p) + ret void +} +"#), + "one", + 2, + &[0, 1, 1], + ); +} + +#[test] +fn phi_bases_follow_rs4gc_find_base_pointer() { + // Perry's phis often merge a NaN-box tag constant with a heap pointer. + // RS4GC gives such a phi a fresh `.base` phi, live wherever it is: two + // relocations per crossing. A phi with a `null` input is its own base, + // and a phi of constants only has a constant base and is not relocated. + let l = assert_census( + &ir(r#" +declare void @sink(i64, i64, i64) +define void @tagphi(i1 %c) gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + br i1 %c, label %t, label %j +t: + br label %j +j: + %p = phi ptr addrspace(1) [ inttoptr (i64 9222246136947933185 to ptr addrspace(1)), %t ], [ %a, %entry ] + %n = phi ptr addrspace(1) [ null, %t ], [ %a, %entry ] + %k = phi ptr addrspace(1) [ inttoptr (i64 9222246136947933185 to ptr addrspace(1)), %t ], [ inttoptr (i64 9222246136947933186 to ptr addrspace(1)), %entry ] + call void @may_collect() + %x = ptrtoint ptr addrspace(1) %p to i64 + %y = ptrtoint ptr addrspace(1) %n to i64 + %z = ptrtoint ptr addrspace(1) %k to i64 + call void @sink(i64 %x, i64 %y, i64 %z) "gc-leaf-function" + ret void +} +"#), + "tagphi", + 3, + &[0, 3], + ); + assert_eq!(l.fresh_bases, 1, "{l:#?}"); +} + +#[test] +fn derived_pointers_make_the_count_a_bound() { + // A gep of a GC pointer: RS4GC relocates the derived value and adds its + // base. The count is flagged as a bound, and the bound holds. + let (l, actual) = census( + &ir(r#" +define void @der() gc "statepoint-example" { +entry: + %a = call ptr addrspace(1) @alloc() + %d = getelementptr i8, ptr addrspace(1) %a, i64 8 + call void @may_collect() + call void @use(ptr addrspace(1) %d) + ret void +} +"#), + "der", + ); + assert!(!l.is_exact(), "{l:#?}"); + assert!( + l.relocations <= actual && actual <= l.relocation_bound(), + "{l:#?} vs {actual}" + ); +} + +/// The in-process backend runs the pipeline in two halves around the count. +/// They must still spell exactly the shipped pipeline, and produce the same +/// module as the one-shot run. +#[test] +fn statepoint_pipeline_split_is_the_shipped_pipeline() { + assert_eq!( + format!( + "{},{}", + crate::linker::STATEPOINT_PREPARE_PASSES, + crate::linker::STATEPOINT_REWRITE_ONLY_PASSES + ), + crate::linker::STATEPOINT_REWRITE_PASSES + ); + let fixture = ir(INVOKE); + let tm = target_machine(); + let rewrite = |split: bool| { + let context = Context::create(); + let module = parse_ir_text(&context, &fixture, "split").expect("parses"); + module.set_triple(&TargetTriple::create(TRIPLE)); + module.set_data_layout(&tm.get_target_data().get_data_layout()); + if split { + run(&module, &tm, crate::linker::STATEPOINT_PREPARE_PASSES); + run(&module, &tm, crate::linker::STATEPOINT_REWRITE_ONLY_PASSES); + } else { + run(&module, &tm, crate::linker::STATEPOINT_REWRITE_PASSES); + } + module.print_to_string().to_string() + }; + assert_eq!(rewrite(true), rewrite(false)); +} + +/// The spill decision flips exactly at the budget, on the exact count. +#[test] +fn the_spill_decision_flips_at_the_threshold() { + let tm = target_machine(); + let context = Context::create(); + let module = parse_ir_text(&context, &ir(INVOKE), "flip").expect("parses"); + module.set_triple(&TargetTriple::create(TRIPLE)); + module.set_data_layout(&tm.get_target_data().get_data_layout()); + let rewritten = super::super::rs4gc_functions(&module); + assert!(rewritten.contains("inv")); + run(&module, &tm, crate::linker::STATEPOINT_PREPARE_PASSES); + let liveness = analyze_module(&module, &rewritten); + assert_eq!(liveness.len(), 1); + assert_eq!(liveness[0].1.relocations, 7); + + let pre = super::super::pre_rewrite_sizes(&module); + super::super::enforce_rs4gc_preflight_budget(&liveness, 7, &pre, None) + .expect("7 relocations are within a budget of 7"); + super::super::enforce_rs4gc_preflight_budget(&liveness, 0, &pre, None) + .expect("a budget of 0 disables spilling"); + let err = super::super::enforce_rs4gc_preflight_budget(&liveness, 6, &pre, None) + .expect_err("7 relocations exceed a budget of 6"); + let retry = super::super::rs4gc_budget_retry(&err).expect("the request stays typed"); + assert_eq!(retry.len(), 1); + assert_eq!(retry[0].name, "inv"); + assert_eq!(retry[0].cap, 6); + assert_eq!( + retry[0].cause, + super::super::Rs4gcBudgetCause::PreRewrite { + statepoints: 5, + invoke_statepoints: 1, + max_live: 2, + relocations: 7, + } + ); + let msg = format!("{err:#}"); + for needle in [ + "before rewrite-statepoints-for-gc", + "`inv`", + "5 statepoints", + "1 of them invokes", + "up to 2 GC values", + "would emit 7 relocations", + "budget 6", + "PERRY_ROOT_SPILL_RELOCATIONS", + "re-lower", + ] { + assert!( + msg.contains(needle), + "message must carry {needle:?}:\n{msg}" + ); + } + // A function outside RS4GC (a shadow-frame function) is never counted. + assert!(analyze_module(&module, &std::collections::HashSet::new()).is_empty()); +} + +/// #11624 follow-up: a function can sit comfortably under the relocation cap +/// and still be predicted to cross the fast-emit machine-pipeline budget — +/// the actual claude-code shape (`__25747`: 0.4 M relocations, under the +/// 1.5 Mi cap, but its rewritten body crossed the 600 k x86-64 fast-emit +/// ceiling and fell back to LLVM's O0 pipeline). The relocation cap alone +/// must not catch this; the fast-emit prediction must. +#[test] +fn a_function_under_the_relocation_cap_but_over_the_fast_emit_budget_spills() { + let tm = target_machine(); + let context = Context::create(); + let module = parse_ir_text(&context, &ir(INVOKE), "fast-emit-cliff").expect("parses"); + module.set_triple(&TargetTriple::create(TRIPLE)); + module.set_data_layout(&tm.get_target_data().get_data_layout()); + let rewritten = super::super::rs4gc_functions(&module); + run(&module, &tm, crate::linker::STATEPOINT_PREPARE_PASSES); + let liveness = analyze_module(&module, &rewritten); + assert_eq!(liveness.len(), 1); + assert_eq!(liveness[0].1.relocations, 7); + let pre = super::super::pre_rewrite_sizes(&module); + + // Comfortably under the default relocation cap, and with no fast-emit + // budget in play (`None`), the fixture must not spill. + super::super::enforce_rs4gc_preflight_budget( + &liveness, + super::super::DEFAULT_ROOT_SPILL_RELOCATIONS, + &pre, + None, + ) + .expect("7 relocations must not trip the default relocation cap"); + + // A fast-emit budget far below the fixture's predicted post-rewrite size + // (pre-rewrite instructions plus 2x its 7 relocations) must spill it, + // even though the relocation cap is untouched. + let err = super::super::enforce_rs4gc_preflight_budget( + &liveness, + super::super::DEFAULT_ROOT_SPILL_RELOCATIONS, + &pre, + Some(5), + ) + .expect_err("a predicted post-rewrite size over a 5-instruction fast-emit budget must spill"); + let retry = super::super::rs4gc_budget_retry(&err).expect("the request stays typed"); + assert_eq!(retry.len(), 1); + assert_eq!(retry[0].name, "inv"); + assert_eq!(retry[0].cap, 5); + match &retry[0].cause { + super::super::Rs4gcBudgetCause::PredictedFastEmit { + relocations, + predicted_instructions, + } => { + assert_eq!(*relocations, 7); + assert!( + *predicted_instructions > 5, + "predicted {predicted_instructions} must exceed the 5-instruction budget" + ); + } + other => panic!("expected PredictedFastEmit, got {other:?}"), + } + let msg = format!("{err:#}"); + for needle in [ + "before rewrite-statepoints-for-gc", + "`inv`", + "would emit 7 relocations", + "under the relocation cap", + "fast-emit", + "budget 5", + "PERRY_LL_FAST_EMIT_MAX_INSTRS", + ] { + assert!( + msg.contains(needle), + "message must carry {needle:?}:\n{msg}" + ); + } + + // 0 disables spilling entirely, including the fast-emit prediction. + super::super::enforce_rs4gc_preflight_budget(&liveness, 0, &pre, Some(5)) + .expect("a budget of 0 disables spilling entirely, even under a tiny fast-emit cap"); +} + +/// The relocation budget is the post-RS4GC instruction budget: every +/// relocation is one instruction of the rewritten body, so a function over it +/// would be spilled after RS4GC anyway. Change one only with the other. +#[test] +fn the_relocation_budget_is_the_post_rewrite_instruction_budget() { + assert_eq!( + super::super::DEFAULT_ROOT_SPILL_RELOCATIONS, + super::super::DEFAULT_RS4GC_MAX_INSTRS as u64 + ); +} + +/// Cyclic phis can share an existing base without any derived instruction. +/// The bound must reserve its difference array for these crossings too. +#[test] +fn cyclic_phis_with_an_existing_base_do_not_require_a_gep() { + let (l, actual) = census( + &ir(r#" +define void @cyclic(ptr addrspace(1) %seed, i1 %again) gc "statepoint-example" { +entry: + br label %head +head: + %a = phi ptr addrspace(1) [ %seed, %entry ], [ %b, %back ] + %b = phi ptr addrspace(1) [ %seed, %entry ], [ %a, %back ] + call void @may_collect() + call void @use2(ptr addrspace(1) %a, ptr addrspace(1) %b) + br i1 %again, label %back, label %exit +back: + br label %head +exit: + ret void +} +"#), + "cyclic", + ); + assert_eq!(l.derived_values, 0, "{l:#?}"); + assert!(l.derived_crossings > 0, "{l:#?}"); + assert!(!l.is_exact(), "{l:#?}"); + assert!(actual <= l.relocation_bound(), "{l:#?} vs {actual}"); +} diff --git a/crates/perry-codegen/src/inprocess/optimize_emit.rs b/crates/perry-codegen/src/inprocess/optimize_emit.rs index b27c0cda02..8c9740deea 100644 --- a/crates/perry-codegen/src/inprocess/optimize_emit.rs +++ b/crates/perry-codegen/src/inprocess/optimize_emit.rs @@ -85,9 +85,21 @@ pub(super) fn optimize_and_emit( // Sizes before the rewrite: the budget message below names them, and // the per-unit report compares them with the post-rewrite census. let budget = rs4gc_instruction_budget(); - let preflight_cap = crate::codegen::helpers::root_spill_relocation_threshold(); + let preflight_cap = root_spill_relocation_threshold(); + // Same resolved budget `fast_emit_fallbacks` uses below, including the + // `PERRY_LL_FAST_EMIT_MAX_INSTRS` override — the preflight predicts + // the same cliff that decision measures for real after RS4GC and the + // IR optimizer run (#11624 follow-up). + let fast_emit_cap = match fast_emit_budget(effective_target) { + FastEmitBudget::Off => None, + FastEmitBudget::Cap(cap) => Some(cap), + }; let rewritten_functions = rs4gc_functions(module); - let pre_sizes = if budget == RewriteBudget::Off && preflight_cap == 0 && stats.is_none() { + let pre_sizes = if budget == RewriteBudget::Off + && preflight_cap == 0 + && fast_emit_cap.is_none() + && stats.is_none() + { std::collections::HashMap::new() } else { pre_rewrite_sizes(module) @@ -100,14 +112,36 @@ pub(super) fn optimize_and_emit( .max_by_key(|(_, n)| **n) .map(|(name, n)| (name.clone(), *n)); } - // The source-level estimate is intentionally cheap but can miss - // codegen expansion (one expression becoming many collecting helper - // calls). Check the actual constructed CallBase/root shape before - // asking RS4GC to perform the potentially super-linear rewrite. - enforce_rs4gc_preflight_budget(module, preflight_cap, &pre_sizes, &rewritten_functions)?; + // S4 (RFC deferred collection): canonicalize exactly as RS4GC's own + // pipeline does, count the relocations it would emit on that input, + // and only then let it run. A function over the budget is re-lowered + // onto a shadow frame before RS4GC spends any time on it. let rewrite_started = std::time::Instant::now(); module - .run_passes(STATEPOINT_REWRITE_PASSES, &tm, PassBuilderOptions::create()) + .run_passes( + crate::linker::STATEPOINT_PREPARE_PASSES, + &tm, + PassBuilderOptions::create(), + ) + .map_err(|e| { + anyhow!( + "in-process statepoint preparation (`{}`) failed:\n{}", + crate::linker::STATEPOINT_PREPARE_PASSES, + e.to_string() + ) + })?; + let prepare_secs = rewrite_started.elapsed().as_secs_f64(); + let liveness_started = std::time::Instant::now(); + let liveness = gc_liveness::analyze_module(module, &rewritten_functions); + let liveness_secs = liveness_started.elapsed().as_secs_f64(); + enforce_rs4gc_preflight_budget(&liveness, preflight_cap, &pre_sizes, fast_emit_cap)?; + let rewrite_started = std::time::Instant::now(); + module + .run_passes( + crate::linker::STATEPOINT_REWRITE_ONLY_PASSES, + &tm, + PassBuilderOptions::create(), + ) .map_err(|e| { anyhow!( "in-process rewrite-statepoints-for-gc failed:\n{}", @@ -129,7 +163,9 @@ pub(super) fn optimize_and_emit( ) })?; if let Some(stats) = stats.as_deref_mut() { - stats.rewrite_secs = rewrite_started.elapsed().as_secs_f64(); + stats.rewrite_secs = prepare_secs + rewrite_started.elapsed().as_secs_f64(); + stats.liveness_secs = liveness_secs; + audit_liveness(module, &liveness, liveness_secs, stats); let (_, total, widest) = module_instruction_census(module); stats.post_rewrite_instructions = total; stats.post_rewrite_widest = widest; @@ -261,6 +297,60 @@ pub(super) fn optimize_and_emit( Ok(pieces) } +/// `PERRY_CODEGEN_UNIT_TIMINGS` audit of the S4 liveness count: after RS4GC, +/// compare each function's predicted relocations with the `gc.relocate`s it +/// really produced, and report one line per function that has a statepoint. +/// Mismatches are marked, so a build log answers "is the count exact on this +/// program" with a grep. +fn audit_liveness( + module: &inkwell::module::Module<'_>, + liveness: &[(String, gc_liveness::FunctionLiveness)], + liveness_secs: f64, + stats: &mut UnitCodegenStats, +) { + use inkwell::values::AsValueRef; + for (name, l) in liveness { + if l.statepoints() == 0 { + continue; + } + let actual = module + .get_function(name) + .map(|f| gc_liveness::count_relocates(f.as_value_ref())) + .unwrap_or(0); + let predicted = l.relocation_bound(); + stats.statepoints += l.statepoints(); + stats.relocations_predicted += predicted; + stats.relocations_actual += actual; + let verdict = if predicted == actual { + "exact" + } else { + stats.liveness_mismatches += 1; + "MISMATCH" + }; + eprintln!( + "[perry] gc-liveness: `{name}` {verdict}: predicted {predicted} relocations{}, rs4gc {actual}; \ + {} statepoints ({} invoke), max live {}, {} gc values, {} blocks, {} instrs, walk {}, {} us", + if l.is_exact() { "" } else { " (bound)" }, + l.statepoints(), + l.invoke_statepoints(), + l.max_live(), + l.gc_values, + l.blocks, + l.instructions, + l.work, + l.micros, + ); + } + eprintln!( + "[perry] gc-liveness: unit: {} statepoints, predicted {} relocations, rs4gc {}, {} mismatching functions, analysis {:.3}s", + stats.statepoints, + stats.relocations_predicted, + stats.relocations_actual, + stats.liveness_mismatches, + liveness_secs + ); +} + #[cfg(test)] mod tests { use super::*; @@ -366,97 +456,6 @@ mod tests { ); } - /// The source-level estimate is only a fast first line of defence. This - /// fixture pins the constructed-IR backstop: managed-root allocas count, - /// ordinary calls count, explicit GC-leaf calls and LLVM intrinsics do - /// not, and only functions which will actually enter RS4GC are governed. - #[test] - fn rs4gc_preflight_uses_constructed_roots_and_non_leaf_calls() { - let fixture = r#" -declare i64 @may_collect() -declare i64 @leaf() -declare void @llvm.donothing() - -define i64 @hot() gc "statepoint-example" { -entry: - %root = alloca ptr addrspace(1) - %plain = alloca i64 - %a = call i64 @may_collect() - %b = call i64 @may_collect() - %c = call i64 @leaf() "gc-leaf-function" - call void @llvm.donothing() - %p = load ptr addrspace(1), ptr %root - %bits = ptrtoint ptr addrspace(1) %p to i64 - %sum = add i64 %a, %b - %sum2 = add i64 %sum, %c - %out = add i64 %sum2, %bits - ret i64 %out -} - -define i64 @shadow() { -entry: - %root = alloca ptr addrspace(1) - %a = call i64 @may_collect() - ret i64 %a -} -"#; - let context = Context::create(); - let module = parse_ir_text(&context, fixture, "preflight_fixture").expect("fixture parses"); - let hot = module.get_function("hot").expect("hot"); - assert_eq!( - rs4gc_preflight_factors(hot), - (1, 2), - "plain allocas, leaf calls and intrinsics do not add RS4GC work" - ); - - // (one constructed root + two possible call-result roots) x two - // safepoints = six estimated relocations. The boundary is exclusive. - let rewritten_functions = rs4gc_functions(&module); - assert_eq!( - rs4gc_preflight_violations(&module, 5, &rewritten_functions), - vec![("hot".to_string(), 1, 2, 6)] - ); - assert!(rs4gc_preflight_violations(&module, 6, &rewritten_functions).is_empty()); - assert!(rs4gc_preflight_violations(&module, 0, &rewritten_functions).is_empty()); - - let pre = pre_rewrite_sizes(&module); - let err = enforce_rs4gc_preflight_budget(&module, 5, &pre, &rewritten_functions) - .expect_err("the constructed shape requests a spill retry"); - let retry = rs4gc_budget_retry(&err).expect("the request stays typed"); - assert_eq!(retry.len(), 1); - assert_eq!(retry[0].name, "hot"); - assert_eq!(retry[0].pre_instructions, pre.get("hot").copied()); - assert_eq!( - retry[0].cause, - Rs4gcBudgetCause::PreRewrite { - root_allocas: 1, - safepoints: 2, - estimated_relocations: 6, - } - ); - assert_eq!(retry[0].cap, 5); - let msg = format!("{err:#}"); - for needle in [ - "before rewrite-statepoints-for-gc", - "`hot`", - "1 managed-root allocas", - "2 non-leaf call sites", - "predicts 6 relocations", - "budget 5", - "PERRY_ROOT_SPILL_RELOCATIONS", - "re-lower", - ] { - assert!( - msg.contains(needle), - "message must carry {needle:?}:\n{msg}" - ); - } - - let no_rewritten_functions = std::collections::HashSet::new(); - enforce_rs4gc_preflight_budget(&module, 1, &pre, &no_rewritten_functions) - .expect("a shadow-root function is outside the preflight budget"); - } - /// Six gc values live across forty safepoints: ~60 instructions before /// `rewrite-statepoints-for-gc`, a few hundred after (each statepoint /// relocates every live value). A budget between the two is exceeded diff --git a/crates/perry-codegen/src/linker.rs b/crates/perry-codegen/src/linker.rs index 8de63cb3d5..2c7e05c2d2 100644 --- a/crates/perry-codegen/src/linker.rs +++ b/crates/perry-codegen/src/linker.rs @@ -37,6 +37,17 @@ use linker_temp::{ pub(crate) const STATEPOINT_REWRITE_PASSES: &str = "always-inline,function(mem2reg,sccp),rewrite-statepoints-for-gc"; +/// [`STATEPOINT_REWRITE_PASSES`] split where the in-process backend counts +/// RS4GC's relocations (RFC deferred collection S4): the canonicalizing half +/// produces RS4GC's exact input, `inprocess::gc_liveness` measures it, and +/// only then does the rewrite run. Running the two halves back to back is the +/// same pipeline; `statepoint_pipeline_split_is_the_shipped_pipeline` pins +/// that the halves still spell the whole. +#[cfg_attr(not(feature = "llvm-inprocess"), allow(dead_code))] +pub(crate) const STATEPOINT_PREPARE_PASSES: &str = "always-inline,function(mem2reg,sccp)"; +#[cfg_attr(not(feature = "llvm-inprocess"), allow(dead_code))] +pub(crate) const STATEPOINT_REWRITE_ONLY_PASSES: &str = "rewrite-statepoints-for-gc"; + /// Cached result of the pre-flight clang probe — evaluated once per process. /// `Some(default_triple)` if the probe succeeded, `None` if it failed. static CLANG_PROBE: OnceLock> = OnceLock::new(); diff --git a/crates/perry-codegen/src/native_emit.rs b/crates/perry-codegen/src/native_emit.rs index 031d3005f3..59eb2b9dd8 100644 --- a/crates/perry-codegen/src/native_emit.rs +++ b/crates/perry-codegen/src/native_emit.rs @@ -229,19 +229,22 @@ pub(crate) fn apply_budget_spill_retry<'a>( changed.insert(violation.name.clone()); match &violation.cause { crate::inprocess::Rs4gcBudgetCause::PreRewrite { - root_allocas, - safepoints, - estimated_relocations, + statepoints, + invoke_statepoints, + max_live, + relocations, } => eprintln!( - "perry: `{}` exceeded the pre-RS4GC relocation estimate ({} managed-root \ - allocas + {} non-leaf call-result temporaries across {} call sites = {} \ - estimated relocations; cap {}); retrying it with precise GC roots in a \ - shadow frame at the requested optimization level (#8583)", + "perry: `{}` keeps its GC roots in a shadow frame instead of statepoints: \ + rewrite-statepoints-for-gc would emit {} relocations ({} statepoints, {} of \ + them invokes relocated on both edges; up to {} GC values live across one), \ + above the budget {}. The function is still compiled at the requested \ + optimization level; only its GC-root representation changes, and its roots \ + stay precise (#8583). Override with PERRY_ROOT_SPILL_RELOCATIONS.", violation.name, - root_allocas, - safepoints, - safepoints, - estimated_relocations, + relocations, + statepoints, + invoke_statepoints, + max_live, violation.cap, ), crate::inprocess::Rs4gcBudgetCause::PostRewrite { post_instructions } => { @@ -257,6 +260,20 @@ pub(crate) fn apply_budget_spill_retry<'a>( violation.cap, ); } + crate::inprocess::Rs4gcBudgetCause::PredictedFastEmit { + relocations, + predicted_instructions, + } => eprintln!( + "perry: `{}` keeps its GC roots in a shadow frame instead of statepoints: \ + under the relocation cap ({} relocations), but rewrite-statepoints-for-gc \ + is predicted to grow it to about {} instructions, above the fast-emit \ + machine-pipeline budget {} — past that budget LLVM falls back to its O0 \ + machine pipeline instead of the optimized one (#11624). The function is \ + still compiled at the requested optimization level; only its GC-root \ + representation changes, and its roots stay precise. Override with \ + PERRY_LL_FAST_EMIT_MAX_INSTRS.", + violation.name, relocations, predicted_instructions, violation.cap, + ), } } } @@ -585,7 +602,7 @@ pub fn compile_module_units_native( 0.0 }; eprintln!( - "[perry] codegen: {module_prefix}: unit {}/{unit_total}: {} fns; pre-RS4GC {} instrs (widest {}); post-RS4GC {} instrs (x{growth:.1}; widest {}); rs4gc {:.1}s, opt {:.1}s, emit {:.1}s", + "[perry] codegen: {module_prefix}: unit {}/{unit_total}: {} fns; pre-RS4GC {} instrs (widest {}); post-RS4GC {} instrs (x{growth:.1}; widest {}); rs4gc {:.1}s (liveness {:.3}s: {} statepoints, {} relocations predicted, {} emitted, {} mismatching fns), opt {:.1}s, emit {:.1}s", i + 1, stats.functions, stats.pre_rewrite_instructions, @@ -593,6 +610,11 @@ pub fn compile_module_units_native( stats.post_rewrite_instructions, widest(&stats.post_rewrite_widest), stats.rewrite_secs, + stats.liveness_secs, + stats.statepoints, + stats.relocations_predicted, + stats.relocations_actual, + stats.liveness_mismatches, stats.optimize_secs, stats.emit_secs, ); @@ -1394,6 +1416,40 @@ mod tests { assert!(after.contains("@js_shadow_frame_pop"), "{after}"); } + /// S4 (RFC deferred collection): the pre-RS4GC spill is decided on the + /// exact relocation count of the constructed function. Under a budget the + /// fixture fits, it keeps its statepoints; under a budget of 1 it is + /// re-lowered onto a shadow frame before RS4GC runs, and still compiles. + #[test] + fn exact_relocation_budget_decides_the_spill_before_rs4gc() { + let _native = crate::codegen::helpers::NativeRootsPin::native(); + let spilled = |budget: u64| { + let mut module = precise_root_fixture(false); + let object = crate::inprocess::with_test_root_spill_threshold(budget, || { + compile_module_native(&mut module, None, "exact_budget_fixture") + }) + .expect("the fixture compiles under either budget"); + assert!(!object.is_empty()); + let function = module + .deduped_function_refs() + .into_iter() + .find(|function| function.name == "native_root_diff_fixture") + .expect("fixture function exists"); + let ir = function.to_ir(); + assert_eq!( + function.spills_roots_to_shadow_frame(), + !ir.contains("gc \"statepoint-example\""), + "{ir}" + ); + function.spills_roots_to_shadow_frame() + }; + assert!( + !spilled(1_000_000), + "a budget the fixture fits must keep statepoints" + ); + assert!(spilled(1), "a budget below the fixture's count must spill"); + } + /// The reported Claude bundle takes the split-unit worker path. Its retry /// source must stay on the producer thread (the `LlFunction` graph is not /// `Send`) while LLVM reports the typed violation from a worker. A compact diff --git a/crates/perry/tests/gc_root_spill_mixed_frames_8583.rs b/crates/perry/tests/gc_root_spill_mixed_frames_8583.rs index 97654e8bd3..277c0faab8 100644 --- a/crates/perry/tests/gc_root_spill_mixed_frames_8583.rs +++ b/crates/perry/tests/gc_root_spill_mixed_frames_8583.rs @@ -13,9 +13,10 @@ //! //! * `PERRY_ROOT_SPILL_RELOCATIONS=0` — spilling disabled, every function on //! native statepoints (the pre-#8583 lowering); -//! * `PERRY_ROOT_SPILL_RELOCATIONS=1` — spill anything with a root and a -//! call, so `run`/`make`/`main` take the shadow frame while the call-free -//! accessor `leaf` stays on statepoints — a genuinely mixed stack. +//! * `PERRY_ROOT_SPILL_RELOCATIONS=1` — spill anything RS4GC would give +//! more than one relocation (the exact count, RFC deferred collection +//! S4), so `run` takes the shadow frame while the call-free accessor +//! `leaf` stays on statepoints — a genuinely mixed stack. //! //! Both binaries run under every moving-collector configuration and must //! produce byte-identical output. If the spilled frame's roots were invisible @@ -33,8 +34,8 @@ fn perry_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_perry")) } -/// `leaf` reads a field and makes no call: at `PERRY_ROOT_SPILL_RELOCATIONS=1` -/// its estimate is `slots × 0 = 0`, so it stays on native statepoints while its +/// `leaf` reads a field and makes no call: RS4GC gives it no relocation, so at +/// `PERRY_ROOT_SPILL_RELOCATIONS=1` it stays on native statepoints while its /// callers spill. `run` holds `a`/`b`/`keep` live across allocating calls, so a /// minor that fires inside `make` must find those roots in `run`'s shadow frame /// and the `leaf` argument in `leaf`'s statepoint frame on the same stack.