Skip to content

perf: stores to name/type/value/length/url/E… are 47–680× slower than Node (key literal not flagged interned, so the write IC never primes) #10500

Description

@proggeramlug

Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. The audit reported an "IC cliff at ≥ 5 subclass fields" in
@noble/hashes; reduction shows the cliff is not the field count but the field NAME: SHA-256's fifth state word is
called E. A static-key store o.<k> = v on a receiver whose layout is not statically known overwrites an existing
slot through the write PIC, but for keys the runtime has already interned (E, PI, now, NaN, name, type,
value, length, message, url, method, headers, body, …) the PIC never primes: every store costs ≈ 5,000
instructions more than the same store to an unaffected name.

Reproduction

bench.ts (20 lines):

// Static-key store `o.<name> = v` on a receiver whose class is not statically known
const variant = process.argv[2] || "sha_fields_E"; const N = Number(process.argv[3] || "2000000");
abstract class ShaE { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract E: number;
  protected set(A: number, B: number, C: number, D: number, E: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.E = E; }
  step(i: number): number { const { A, B, C, D, E } = this; this.set((E + i) | 0, A, B, C, D); return this.A; } }
class Sha256E extends ShaE { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected E = 5; constructor() { super(); } }
abstract class ShaX { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract X: number;
  protected set(A: number, B: number, C: number, D: number, X: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.X = X; }
  step(i: number): number { const { A, B, C, D, X } = this; this.set((X + i) | 0, A, B, C, D); return this.A; } }
class Sha256X extends ShaX { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected X = 5; constructor() { super(); } }
const sE = new Sha256E(), sX = new Sha256X();
const objs: any[] = [{ name: 0, type: 0, nam: 0, typ: 0 }, { name: 1, type: 1, nam: 1, typ: 1 }];
const V: Record<string, (n: number) => number> = {
  sha_fields_E(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sE.step(i)) % 1000000007; return a; },
  sha_fields_X(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sX.step(i)) % 1000000007; return a; },
  store_name_type(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.name = i; o.type = i; a += o.name & 1; } return a; },
  store_nam_typ(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.nam = i; o.typ = i; a += o.nam & 1; } return a; },
};
V[variant](N / 5 | 0); const t0 = performance.now(); const cs = V[variant](N);
console.log(`variant=${variant} checksum=${cs} ms=${(performance.now() - t0).toFixed(2)}`);
PERRY_NO_AUTO_OPTIMIZE=1 perry compile bench.ts -o bench
for v in sha_fields_X sha_fields_E store_nam_typ store_name_type; do node bench.ts $v 2000000; ./bench $v 2000000; done

Measurements

Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 2,000,000; Perry
instructions per iteration = (whole-process instructions:u − 38 M startup) / 2.4 M.

variant Node loop ms Perry loop ms ratio Perry instructions (per iter) Node wall ms Perry wall ms
sha_fields_X (control: 5th state word named X) 37.7 202.9 5.4× 2.61 G (1,070) 164 299
sha_fields_E (noble shape: 5th state word named E) 37.6 1,757.2 47× 14.73 G (6,120) 185 2,181
store_nam_typ (control: o.nam = i; o.typ = i) 2.9 59.0 20× 1.03 G (410) 86 94
store_name_type (o.name = i; o.type = i) 3.0 2,042.7 680× 24.00 G (9,980) 116 2,622

Checksums identical. The two SHA variants differ only in one identifier; the one extra slow store is ≈ 5,050
instructions. Renaming name/type to nam/typ removes ≈ 4,800 instructions per store.

Name sweep (one function o.<k> = i per name on an object literal that already has every key, N = 300,000;
instructions per call including ≈ 350 of loop/call overhead):

  • slow (≈ 5,100–5,500): length, name, value, type, message, key, next, size, headers, url,
    method, body, start, flags, text, source, buffer, offset, E, PI, LN2, SQRT2, NaN,
    Infinity, now.
  • fast (≈ 350): x, done, data, id, status, code, error, result, index, count, end, pos,
    kind, parent, input, callback, options, ttl, state, target, A–D, F–H, e, MAX_VALUE,
    EPSILON, UTC.
  • Reads are not affected (a read of this.E from the base method was as fast as this.A), and neither are stores where the receiver's class is statically known (a method on
    the class that declares the field lowers to a direct slot store).

Impact

  • @noble/hashes 2.2.0 SHA-256 (sha2.js:43: SHA2_32B.set(A, …, H) writes this.E = E | 0 on the _SHA256
    subclass instance, once per compression block; legacy.js:40 has the same store): the numeric-group report (v0.5.1587) attributed ≈ 27 % of sha256's
    Perry CPU to this property bucket (set() 14.7 % inclusive) and reproduced the gap as "4 fields 2.7×, 5 fields 28×".
  • Every package that stores to one of the slow names on a dynamically-typed receiver: err.message = …,
    this.name = …, config.headers = …, config.url/method/body = …, node.flags = …, this.type/value/length = …
    (axios, pg, mysql2, typescript all do this in hot paths; not separately attributed by the audit).

Mechanism

  • js_put_value_set_ic_miss (crates/perry-runtime/src/proxy/put_value.rs:374) runs the store and then declines to
    prime unless the key string's GcHeader has GC_FLAG_INTERNED set and GC_FLAG_FORWARDED clear
    (put_value.rs:444-445) (verified). Without a prime, the codegen PIC (crates/perry-codegen/src/expr/proxy_reflect.rs:465,
    identical IR for this.E and this.X, checked with --trace llvm) misses on every execution.
  • gdb at js_put_value_set_ic_miss (verified): the key pointer passed for E and now has gc_flags = 0x02; for x
    and UTC it has gc_flags = 0x12 (GC_FLAG_INTERNED = 0x10, crates/perry-runtime/src/gc/types.rs:1145).
  • Codegen materializes every string-pool literal with js_string_from_bytes
    (crates/perry-codegen/src/codegen/string_pool.rs:442-447), which does not intern (verified). The runtime interns
    the names it spells itself through canonical_key / intern_ascii_literal → intern_dispatch_bytes
    (crates/perry-runtime/src/string/mod.rs:210, mod.rs:446, crates/perry-runtime/src/string/intern.rs), which
    flags the runtime's own copy (verified). The slow set matches names the runtime installs or dispatches on — builtin
    property names, Math constants, now (inferred).
  • (inferred) When the program's pooled copy of such a name is later canonicalized, the intern table already holds the
    runtime's copy, so the lookup returns that pointer and never flags the pooled one (js_string_intern,
    string/intern.rs:58-99, sets GC_FLAG_INTERNED only on a table miss). The PIC call site keeps passing the
    pooled, unflagged copy, so the prime check fails forever. Names the runtime has not pre-interned are inserted and
    flagged on first use, which is why x/nam/X hit.

What fast looks like

The write PIC should prime on any key whose content is canonical, not on the identity of the pooled copy: either
intern string-pool literals at module init and store the canonical pointer in the handle global, or have the miss
handler prime with (and codegen compare against) the canonical pointer. Targets: sha_fields_E equal to
sha_fields_X (≈ 1,070 instructions per iteration, from 6,120); store_name_type equal to store_nam_typ; every name
in the sweep at ≈ 350 instructions per call. A regression test can assert the prime for a store to name.

Notes

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    package-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions