Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions changelog.d/11788-10511-bitwise.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
perf: bitwise ops on destructured or annotated locals run native inside a Number-versioned loop clone (#10511)

Two changes:

1. Unary `~` on an operand the compiler cannot prove is a Number now takes the
same inline guard as the binary bitwise operators (`|v| < 2^63`, with the
exact `js_dynamic_bitnot` on the cold arm), instead of calling the helper
on every evaluation. BigInt, `valueOf` and throwing coercions keep their
exact semantics on the cold arm.

2. New loop tier `stmt/number_local_loop.rs`. In a receiver-free loop whose
bitwise operators read locals with no function-scope Number proof (noble's
`let { A, B, C, D } = this` SHA-2 round), a local is Number at every read
inside the loop when (1) it holds a Number at entry, which one inline test
checks, (2) every write that can execute in the loop, in the body AND the
condition AND the update clause, produces a Number given the admitted
locals (a greatest fixed point), and (3) no closure, `await`, `yield`,
boxed cell or module global can write it. The loop then runs in a clone
whose 5L Number scope holds those locals, each in a plain F64 alloca that
is written back on exit. When the entry test fails, the ordinary loop runs.

Result (instr/iter): `destructure` 212 to 57 and `annotated` 213 to 57 (node
41). `prop` (78) and `bigmask` (186) are dominated by `any`-receiver property
reads that are not hoisted, not by bitwise ops, and are unchanged. Tests: IR
tests in `stmt/number_local_loop_tests.rs` and
`expr/unary_bitnot_tests.rs`. Sabotaging the condition/update walk or the fixed
point turns the clause and concatenation witnesses red. The gap test
`test_gap_10511_number_local_loop.ts` covers ToInt32 edge values, non-Number
entries, clause writes, break/return/throw exits, nested loops, and BigInt
TypeErrors.
28 changes: 28 additions & 0 deletions changelog.d/11788-10718-array-rmw.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
perf: array read-modify-write in a multi-statement loop body runs on the dense range tier (#10718)

`for (let i = 0; i < N; i++) { a[i] = a[i] + 1; s += a[i]; }` fell off every
loop tier. The classic range mode admits only one statement, because a side exit
after a store would replay the store. The dense mode, which has no side exits,
admitted only masked `a[e & K]` stores. The loop therefore ran on the generic
path and re-validated the receiver for each of its three accesses: 206
instructions per element against node's 39.

The dense mode now admits counter-offset stores (`a[i] = …`, `a[i ± c] = …`)
under the same rule as its masked stores. The entry guard validates the whole
counter window (in bounds, hole-free, raw-f64, plain, integrity-clean), and the
stored value must be a statically genuine double: a literal, the counter, a
guarded element read, or `+ - * /` and negation over those. So the store needs
no value check and no side exit, and an iteration still runs entirely in one
copy. The compound-assignment alias fold (#10743) now also runs over
multi-statement bodies, so `a[i] += 1; s += a[i]` takes the same path. The
matcher verifies accumulators with the same array set the lowering uses, which
lifts the old single-array restriction (`a[i] = a[i] + b[i]; s += a[i]`).

Result: `arrRmw` goes from 206 to 14.6 instructions per element (node 39.2), and
the single-statement form goes from 24.5 to 20.5. Tests: IR tests in
`stmt/range_loop_dense_store_tests.rs`. Removing the genuine-RHS rule turns
`declines_a_store_whose_value_is_not_provably_a_double` red (verified by
sabotage). The gap test `test_gap_10718_array_rmw_dense.ts` covers holes,
deleted elements, out-of-bounds windows, numeric strings, BigInt, `valueOf`,
element-kind changes between loop entries, frozen arrays, aliasing, and
-0/NaN/Infinity.
9 changes: 8 additions & 1 deletion crates/perry-codegen/src/expr/binary.rs
Original file line number Diff line number Diff line change
Expand Up @@ -415,14 +415,21 @@ fn lower_guarded_numeric_arith(
/// exponent/mantissa tower. The Numbers it turns away (NaN, ±Infinity,
/// |v| >= 2^63) are the ones ToInt32 has to special-case, and the cold arm's
/// helper already does.
fn emit_is_int64_exact_number(ctx: &mut FnCtx<'_>, value: &str) -> String {
pub(super) fn emit_is_int64_exact_number(ctx: &mut FnCtx<'_>, value: &str) -> String {
const TWO_POW_63: &str = "0x43E0000000000000";
let magnitude = ctx
.block()
.call(DOUBLE, "llvm.fabs.f64", &[(DOUBLE, value)]);
ctx.block().fcmp("olt", &magnitude, TWO_POW_63)
}

/// Whether an unproven bitwise operand takes the inline guard (#10418,
/// #10511): both A/B knobs must be on. Unary `~` shares the binary
/// operators' gate so one switch reverts every bitwise guard together.
pub(super) fn guarded_bitwise_enabled() -> bool {
guarded_arith_enabled() && inline_nonbigint_bitwise_enabled()
}

/// `PERRY_GUARDED_ARITH=0` restores the unconditional dynamic helper for
/// `-`, `*`, `/` and the bitwise operators.
fn guarded_arith_enabled() -> bool {
Expand Down
7 changes: 2 additions & 5 deletions crates/perry-codegen/src/expr/index_get.rs
Original file line number Diff line number Diff line change
Expand Up @@ -44,12 +44,9 @@ mod foreign_counter;
pub(super) mod guarded_array;
pub(crate) use foreign_counter::{
affine_counter_occurrences, affine_index_fits_i64, emit_affine_index_i64_with,
packed_f64_loop_index_parts,
};
use foreign_counter::{
affine_packed_loop_read, emit_affine_index_i64, foreign_packed_loop_read,
packed_f64_loop_offset_read,
packed_f64_loop_index_parts, packed_f64_loop_offset_read,
};
use foreign_counter::{affine_packed_loop_read, emit_affine_index_i64, foreign_packed_loop_read};
pub(crate) use guarded_array::emit_array_region_guard;
mod inline_dyn_typed_array;

Expand Down
30 changes: 28 additions & 2 deletions crates/perry-codegen/src/expr/index_set.rs
Original file line number Diff line number Diff line change
Expand Up @@ -219,7 +219,12 @@ fn packed_f64_loop_fact_for_index(
) -> Option<(PackedF64LoopFact, u32, i32)> {
let (idx_id, offset) = super::packed_f64_loop_index_parts(index)?;
let fact = packed_f64_loop_fact(ctx, arr_id, idx_id)?;
if offset != 0 && !fact.allow_holes {
// A dense range guard validated the whole constant-offset window too
// (`window_validated` without `allow_holes`), so its offsets are proven
// exactly like the hole-tolerant guard's (#10718). Only the length-bound
// guard of the versioned matcher leaves an offset store unproven; an
// affine fact proves the receiver only.
if offset != 0 && !fact.allow_holes && (!fact.window_validated || fact.affine_indices) {
return None;
}
Some((fact, idx_id, offset))
Expand Down Expand Up @@ -1017,7 +1022,28 @@ pub(crate) fn lower(
// than abort codegen, let U32 facts fall through to the
// generic/bounded array-store path below (correct, just
// not the packed fast path).
if !matches!(fact.array_kind, PackedNumericLoopKind::U32) {
// A dense range copy has no side exit to take: an
// iteration must run entirely in one copy, because a
// multi-statement body may already have stored. The
// matcher admits a store there only with a statically
// genuine RHS (`dense_masked_store_rhs_is_admissible`,
// this predicate's match-time twin); a store that
// reaches here without that proof is matcher/lowering
// drift and must not get a side-exiting store.
let dense_copy =
!fact.allow_holes && fact.window_validated && !fact.affine_indices;
let side_exit_forbidden_but_needed = dense_copy
&& !super::masked_window::masked_store_rhs_is_genuine_f64(
ctx,
value.as_ref(),
);
debug_assert!(
!side_exit_forbidden_but_needed,
"dense range store without a genuine-f64 RHS proof"
);
if !matches!(fact.array_kind, PackedNumericLoopKind::U32)
&& !side_exit_forbidden_but_needed
{
if let Some(i32_slot) = ctx.i32_counter_slots.get(&idx_id).cloned() {
let idx_i32 = load_packed_loop_index_i32(ctx, &i32_slot, offset);
return lower_packed_numeric_loop_index_set(
Expand Down
19 changes: 19 additions & 0 deletions crates/perry-codegen/src/expr/masked_window.rs
Original file line number Diff line number Diff line change
Expand Up @@ -256,6 +256,7 @@ pub(crate) fn masked_store_rhs_is_genuine_f64(ctx: &FnCtx<'_>, expr: &Expr) -> b
Expr::IndexGet { object, index } => match object.as_ref() {
Expr::LocalGet(arr_id) => {
masked_window_fact_for_index(ctx, *arr_id, index.as_ref()).is_some()
|| counter_read_is_genuine_f64(ctx, *arr_id, index.as_ref())
}
_ => false,
},
Expand All @@ -277,6 +278,24 @@ pub(crate) fn masked_store_rhs_is_genuine_f64(ctx: &FnCtx<'_>, expr: &Expr) -> b
}
}

/// #10718: a counter-offset read (`a[i]`, `a[i ± c]`) served by an active
/// raw-f64 `PackedF64LoopFact` materializes a genuine double. Every lowering
/// of such a read either loads a slot the entry guard proved holds a raw-f64
/// number (canonical by the store-side invariant) or leaves the iteration
/// first: the hole-tolerant range copy hole-checks and side-exits, and an
/// offset the length-bound guard does not cover bounds-checks and
/// side-exits — both BEFORE the value exists. Affine (receiver-only) facts
/// and the i32/u32 kinds are excluded; they are not this argument.
fn counter_read_is_genuine_f64(ctx: &FnCtx<'_>, arr_id: u32, index: &Expr) -> bool {
crate::expr::index_get::packed_f64_loop_offset_read(ctx, arr_id, index).is_some_and(
|(fact, idx_id, _, _)| {
matches!(fact.array_kind, crate::expr::PackedNumericLoopKind::F64)
&& !fact.affine_indices
&& ctx.i32_counter_slots.contains_key(&idx_id)
},
)
}

/// Emit the in-window element STORE for a store-admitting fact: the dense
/// entry guard proved a plain, integrity-clean raw-f64 numeric array with the
/// whole static window in bounds and hole-free, and the caller proved the
Expand Down
52 changes: 52 additions & 0 deletions crates/perry-codegen/src/expr/unary.rs
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,56 @@ use super::{is_known_i32_range, lower_expr, FnCtx};
/// `lower_bitwise_operand_i32` only produces a value for an operand it can
/// prove is a Number, and a BigInt-typed operand declines up front, so a
/// BigInt `~x` (a BigInt result) still reaches the dynamic helper in [`lower`].
/// `~v` for an operand the compiler cannot prove is a Number (#10511): one
/// inline test, the numeric arm inline, the dynamic helper cold.
///
/// This is the unary twin of the binary bitwise guard
/// (`binary::lower_guarded_numeric_arith`) and uses the same test,
/// `|v| < 2^63`. Every NaN-boxed tag is a NaN bit pattern, so a string,
/// object, boolean, `undefined`, an int32 box or a BigInt fails the ordered
/// compare and reaches `js_dynamic_bitnot`, which keeps ToNumeric's exact
/// semantics: a BigInt stays a BigInt (`~1n === -2n`), `valueOf` runs, and
/// a throwing coercion throws. The Numbers the test turns away — NaN,
/// ±Infinity, |v| >= 2^63 — are the ones whose ToInt32 needs the helper's
/// special cases; for every Number that passes, truncating to i64 and keeping
/// the low 32 bits IS ToInt32, so the numeric arm is one `fptosi`.
///
/// `v` is the already-lowered operand: nothing runs between its evaluation
/// and either arm, so no rooting window opens that the old unconditional
/// helper call did not already have.
fn lower_guarded_bitnot(ctx: &mut FnCtx<'_>, v: &str) -> String {
let is_num = super::binary::emit_is_int64_exact_number(ctx, v);
let fast_idx = ctx.new_block("guarded_bitnot.numeric");
let slow_idx = ctx.new_block("guarded_bitnot.dynamic");
let merge_idx = ctx.new_block("guarded_bitnot.merge");
let fast_label = ctx.block_label(fast_idx);
let slow_label = ctx.block_label(slow_idx);
let merge_label = ctx.block_label(merge_idx);
ctx.block().cond_br(&is_num, &fast_label, &slow_label);

ctx.current_block = fast_idx;
let fast_val = {
let blk = ctx.block();
let i = blk.toint32_fast(v);
let flipped = blk.xor(I32, &i, "-1");
blk.sitofp(I32, &flipped, DOUBLE)
};
let fast_end = ctx.block().label.clone();
ctx.block().br(&merge_label);

ctx.current_block = slow_idx;
super::emit_versioned_loop_callback_deopt(ctx);
let slow_val = ctx
.block()
.call(DOUBLE, "js_dynamic_bitnot", &[(DOUBLE, v)]);
let slow_end = ctx.block().label.clone();
ctx.block().br(&merge_label);

ctx.current_block = merge_idx;
ctx.block()
.phi(DOUBLE, &[(&fast_val, &fast_end), (&slow_val, &slow_end)])
}

pub(crate) fn lower_bitnot_value(
ctx: &mut FnCtx<'_>,
operand: &Expr,
Expand Down Expand Up @@ -140,6 +190,8 @@ pub(crate) fn lower(ctx: &mut FnCtx<'_>, expr: &Expr) -> Result<String> {
let blk = ctx.block();
let flipped = blk.xor(I32, &i, "-1");
Ok(blk.sitofp(I32, &flipped, DOUBLE))
} else if super::binary::guarded_bitwise_enabled() {
Ok(lower_guarded_bitnot(ctx, &v))
} else {
Ok(blk.call(DOUBLE, "js_dynamic_bitnot", &[(DOUBLE, &v)]))
}
Expand Down
20 changes: 20 additions & 0 deletions crates/perry-codegen/src/expr/unary_bitnot_tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,26 @@ fn erased_direct_bitnot_retains_bigint_dispatch() {
);
}

/// #10511: the erased operand keeps the helper, but only on the COLD arm of
/// one inline test (the binary bitwise guard, `|v| < 2^63`), so a Number
/// operand never reaches the call. Reverting `~` to the unconditional helper
/// keeps the call but loses the guard blocks, and this goes red.
#[test]
fn erased_direct_bitnot_takes_the_inline_number_guard() {
let ir = ir_for(
"guarded_erased_bitnot",
vec![erased(X, "x"), result(bitnot(Expr::LocalGet(X)))],
);
assert!(
ir.contains("guarded_bitnot.numeric") && ir.contains("guarded_bitnot.dynamic"),
"an erased ~x must be one inline Number test with a cold helper arm:\n{ir}"
);
assert!(
ir.contains("0x43E0000000000000"),
"the guard is the binary bitwise operators' |v| < 2^63 test:\n{ir}"
);
}

#[test]
fn potentially_bigint_binary_result_retains_bitnot_dispatch() {
let unknown_value = |id| Expr::PropertyGet {
Expand Down
Loading
Loading