Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. The audit reported an "IC cliff at ≥ 5 subclass fields" in
@noble/hashes; reduction shows the cliff is not the field count but the field NAME: SHA-256's fifth state word is
called E. A static-key store o.<k> = v on a receiver whose layout is not statically known overwrites an existing
slot through the write PIC, but for keys the runtime has already interned (E, PI, now, NaN, name, type,
value, length, message, url, method, headers, body, …) the PIC never primes: every store costs ≈ 5,000
instructions more than the same store to an unaffected name.
Reproduction
bench.ts (20 lines):
// Static-key store `o.<name> = v` on a receiver whose class is not statically known
const variant = process.argv[2] || "sha_fields_E"; const N = Number(process.argv[3] || "2000000");
abstract class ShaE { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract E: number;
protected set(A: number, B: number, C: number, D: number, E: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.E = E; }
step(i: number): number { const { A, B, C, D, E } = this; this.set((E + i) | 0, A, B, C, D); return this.A; } }
class Sha256E extends ShaE { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected E = 5; constructor() { super(); } }
abstract class ShaX { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract X: number;
protected set(A: number, B: number, C: number, D: number, X: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.X = X; }
step(i: number): number { const { A, B, C, D, X } = this; this.set((X + i) | 0, A, B, C, D); return this.A; } }
class Sha256X extends ShaX { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected X = 5; constructor() { super(); } }
const sE = new Sha256E(), sX = new Sha256X();
const objs: any[] = [{ name: 0, type: 0, nam: 0, typ: 0 }, { name: 1, type: 1, nam: 1, typ: 1 }];
const V: Record<string, (n: number) => number> = {
sha_fields_E(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sE.step(i)) % 1000000007; return a; },
sha_fields_X(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sX.step(i)) % 1000000007; return a; },
store_name_type(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.name = i; o.type = i; a += o.name & 1; } return a; },
store_nam_typ(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.nam = i; o.typ = i; a += o.nam & 1; } return a; },
};
V[variant](N / 5 | 0); const t0 = performance.now(); const cs = V[variant](N);
console.log(`variant=${variant} checksum=${cs} ms=${(performance.now() - t0).toFixed(2)}`);
PERRY_NO_AUTO_OPTIMIZE=1 perry compile bench.ts -o bench
for v in sha_fields_X sha_fields_E store_nam_typ store_name_type; do node bench.ts $v 2000000; ./bench $v 2000000; done
Measurements
Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 2,000,000; Perry
instructions per iteration = (whole-process instructions:u − 38 M startup) / 2.4 M.
| variant |
Node loop ms |
Perry loop ms |
ratio |
Perry instructions (per iter) |
Node wall ms |
Perry wall ms |
sha_fields_X (control: 5th state word named X) |
37.7 |
202.9 |
5.4× |
2.61 G (1,070) |
164 |
299 |
sha_fields_E (noble shape: 5th state word named E) |
37.6 |
1,757.2 |
47× |
14.73 G (6,120) |
185 |
2,181 |
store_nam_typ (control: o.nam = i; o.typ = i) |
2.9 |
59.0 |
20× |
1.03 G (410) |
86 |
94 |
store_name_type (o.name = i; o.type = i) |
3.0 |
2,042.7 |
680× |
24.00 G (9,980) |
116 |
2,622 |
Checksums identical. The two SHA variants differ only in one identifier; the one extra slow store is ≈ 5,050
instructions. Renaming name/type to nam/typ removes ≈ 4,800 instructions per store.
Name sweep (one function o.<k> = i per name on an object literal that already has every key, N = 300,000;
instructions per call including ≈ 350 of loop/call overhead):
- slow (≈ 5,100–5,500):
length, name, value, type, message, key, next, size, headers, url,
method, body, start, flags, text, source, buffer, offset, E, PI, LN2, SQRT2, NaN,
Infinity, now.
- fast (≈ 350):
x, done, data, id, status, code, error, result, index, count, end, pos,
kind, parent, input, callback, options, ttl, state, target, A–D, F–H, e, MAX_VALUE,
EPSILON, UTC.
- Reads are not affected (a read of
this.E from the base method was as fast as this.A), and neither are stores where the receiver's class is statically known (a method on
the class that declares the field lowers to a direct slot store).
Impact
- @noble/hashes 2.2.0 SHA-256 (
sha2.js:43: SHA2_32B.set(A, …, H) writes this.E = E | 0 on the _SHA256
subclass instance, once per compression block; legacy.js:40 has the same store): the numeric-group report (v0.5.1587) attributed ≈ 27 % of sha256's
Perry CPU to this property bucket (set() 14.7 % inclusive) and reproduced the gap as "4 fields 2.7×, 5 fields 28×".
- Every package that stores to one of the slow names on a dynamically-typed receiver:
err.message = …,
this.name = …, config.headers = …, config.url/method/body = …, node.flags = …, this.type/value/length = …
(axios, pg, mysql2, typescript all do this in hot paths; not separately attributed by the audit).
Mechanism
js_put_value_set_ic_miss (crates/perry-runtime/src/proxy/put_value.rs:374) runs the store and then declines to
prime unless the key string's GcHeader has GC_FLAG_INTERNED set and GC_FLAG_FORWARDED clear
(put_value.rs:444-445) (verified). Without a prime, the codegen PIC (crates/perry-codegen/src/expr/proxy_reflect.rs:465,
identical IR for this.E and this.X, checked with --trace llvm) misses on every execution.
- gdb at
js_put_value_set_ic_miss (verified): the key pointer passed for E and now has gc_flags = 0x02; for x
and UTC it has gc_flags = 0x12 (GC_FLAG_INTERNED = 0x10, crates/perry-runtime/src/gc/types.rs:1145).
- Codegen materializes every string-pool literal with
js_string_from_bytes
(crates/perry-codegen/src/codegen/string_pool.rs:442-447), which does not intern (verified). The runtime interns
the names it spells itself through canonical_key / intern_ascii_literal → intern_dispatch_bytes
(crates/perry-runtime/src/string/mod.rs:210, mod.rs:446, crates/perry-runtime/src/string/intern.rs), which
flags the runtime's own copy (verified). The slow set matches names the runtime installs or dispatches on — builtin
property names, Math constants, now (inferred).
- (inferred) When the program's pooled copy of such a name is later canonicalized, the intern table already holds the
runtime's copy, so the lookup returns that pointer and never flags the pooled one (js_string_intern,
string/intern.rs:58-99, sets GC_FLAG_INTERNED only on a table miss). The PIC call site keeps passing the
pooled, unflagged copy, so the prime check fails forever. Names the runtime has not pre-interned are inserted and
flagged on first use, which is why x/nam/X hit.
What fast looks like
The write PIC should prime on any key whose content is canonical, not on the identity of the pooled copy: either
intern string-pool literals at module init and store the canonical pointer in the handle global, or have the miss
handler prime with (and codegen compare against) the canonical pointer. Targets: sha_fields_E equal to
sha_fields_X (≈ 1,070 instructions per iteration, from 6,120); store_name_type equal to store_nam_typ; every name
in the sweep at ≈ 350 instructions per call. A regression test can assert the prime for a store to name.
Notes
Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. The audit reported an "IC cliff at ≥ 5 subclass fields" in
@noble/hashes; reduction shows the cliff is not the field count but the field NAME: SHA-256's fifth state word is
called
E. A static-key storeo.<k> = von a receiver whose layout is not statically known overwrites an existingslot through the write PIC, but for keys the runtime has already interned (
E,PI,now,NaN,name,type,value,length,message,url,method,headers,body, …) the PIC never primes: every store costs ≈ 5,000instructions more than the same store to an unaffected name.
Reproduction
bench.ts(20 lines):Measurements
Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 2,000,000; Perry
instructions per iteration = (whole-process
instructions:u− 38 M startup) / 2.4 M.sha_fields_X(control: 5th state word namedX)sha_fields_E(noble shape: 5th state word namedE)store_nam_typ(control:o.nam = i; o.typ = i)store_name_type(o.name = i; o.type = i)Checksums identical. The two SHA variants differ only in one identifier; the one extra slow store is ≈ 5,050
instructions. Renaming
name/typetonam/typremoves ≈ 4,800 instructions per store.Name sweep (one function
o.<k> = iper name on an object literal that already has every key, N = 300,000;instructions per call including ≈ 350 of loop/call overhead):
length,name,value,type,message,key,next,size,headers,url,method,body,start,flags,text,source,buffer,offset,E,PI,LN2,SQRT2,NaN,Infinity,now.x,done,data,id,status,code,error,result,index,count,end,pos,kind,parent,input,callback,options,ttl,state,target,A–D,F–H,e,MAX_VALUE,EPSILON,UTC.this.Efrom the base method was as fast asthis.A), and neither are stores where the receiver's class is statically known (a method onthe class that declares the field lowers to a direct slot store).
Impact
sha2.js:43:SHA2_32B.set(A, …, H)writesthis.E = E | 0on the_SHA256subclass instance, once per compression block;
legacy.js:40has the same store): the numeric-group report (v0.5.1587) attributed ≈ 27 % of sha256'sPerry CPU to this property bucket (
set()14.7 % inclusive) and reproduced the gap as "4 fields 2.7×, 5 fields 28×".err.message = …,this.name = …,config.headers = …,config.url/method/body = …,node.flags = …,this.type/value/length = …(axios, pg, mysql2, typescript all do this in hot paths; not separately attributed by the audit).
Mechanism
js_put_value_set_ic_miss(crates/perry-runtime/src/proxy/put_value.rs:374) runs the store and then declines toprime unless the key string's
GcHeaderhasGC_FLAG_INTERNEDset andGC_FLAG_FORWARDEDclear(
put_value.rs:444-445) (verified). Without a prime, the codegen PIC (crates/perry-codegen/src/expr/proxy_reflect.rs:465,identical IR for
this.Eandthis.X, checked with--trace llvm) misses on every execution.js_put_value_set_ic_miss(verified): the key pointer passed forEandnowhasgc_flags = 0x02; forxand
UTCit hasgc_flags = 0x12(GC_FLAG_INTERNED = 0x10,crates/perry-runtime/src/gc/types.rs:1145).js_string_from_bytes(
crates/perry-codegen/src/codegen/string_pool.rs:442-447), which does not intern (verified). The runtime internsthe names it spells itself through
canonical_key/intern_ascii_literal→intern_dispatch_bytes(
crates/perry-runtime/src/string/mod.rs:210,mod.rs:446,crates/perry-runtime/src/string/intern.rs), whichflags the runtime's own copy (verified). The slow set matches names the runtime installs or dispatches on — builtin
property names,
Mathconstants,now(inferred).runtime's copy, so the lookup returns that pointer and never flags the pooled one (
js_string_intern,string/intern.rs:58-99, setsGC_FLAG_INTERNEDonly on a table miss). The PIC call site keeps passing thepooled, unflagged copy, so the prime check fails forever. Names the runtime has not pre-interned are inserted and
flagged on first use, which is why
x/nam/Xhit.What fast looks like
The write PIC should prime on any key whose content is canonical, not on the identity of the pooled copy: either
intern string-pool literals at module init and store the canonical pointer in the handle global, or have the miss
handler prime with (and codegen compare against) the canonical pointer. Targets:
sha_fields_Eequal tosha_fields_X(≈ 1,070 instructions per iteration, from 6,120);store_name_typeequal tostore_nam_typ; every namein the sweep at ≈ 350 instructions per call. A regression test can assert the prime for a store to
name.Notes
A..H); a 6-field subclass writingFis fast, a 5-field subclass writing
Qis fast, and a 6-field subclass writingEat index 5 is slow.o.k = vis ~100× slower than Node (static-key write IC primes only overwrites; no add-transition cache) #10496).