Summary
/// A box is simply a heap-allocated JSValue bit slot.
#[repr(C)]
pub struct Box { pub value: u64 }
A box has no header — eight bytes of payload and nothing else. It therefore cannot answer any question about itself, so every question is answered by a thread-local side table keyed by its address:
BOX_REGISTRY (PtrHashSet<usize>) — "is this pointer a box?", because the object carries no identity;
BOX_CAPTURE_COUNTS (RefCell<PtrHashMap<usize, usize>>) — the capture-edge count, because the object carries no count field.
Giving Box one header word would let both live in the object. On the profile below that is ~9.8% of a real run, and — because a hash entry costs more than the word it would replace — it also looks like a net memory win.
Measured
tsc --noEmit demo.ts (two-line input) compiled with Perry 0.5.1596 + #10656/#10685, sample leaf histogram, 4,917 attributed samples:
| symbol |
samples |
share |
what it is |
is_registered_box_ptr |
254 |
5.2% |
BOX_REGISTRY probe (+ plausibility heuristic + 8-slot cache) |
js_closure_set_box_capture_ptr |
171 |
3.5% |
capture-slot publish |
increment_cell_capture_count |
107 |
2.2% |
BOX_CAPTURE_COUNTS hash |
decrement_cell_capture_count |
75 |
1.5% |
BOX_CAPTURE_COUNTS hash |
publish_box_cell |
44 |
0.9% |
|
is_registered_box_ptr is the single largest symbol in the whole profile, which is otherwise flat (next is 4.0%).
Both counting paths are the same shape:
fn increment_cell_capture_count(cell: usize, amount: usize) {
BOX_CAPTURE_COUNTS.with(|counts| { // thread-local resolution
let mut counts = counts.borrow_mut(); // RefCell borrow
let record = counts.entry(cell).or_default(); // hash lookup + insert
...
Why the current design cannot be optimised further in place
box.rs's own comment explains the ceiling, and it is correct:
A negative cache would NOT be sound — an address that is not a box today can be minted as one tomorrow — so a miss always falls through to the hash set, and only a confirmed positive is recorded.
The same comment records that the positive cache already took this family from 8.2% + 5.9% + 5.5% down to today's numbers. That was the available win. What remains is structural: a slot holding a raw box pointer is indistinguishable from a slot holding a NaN-boxed value, so the runtime must ask an external authority, and negatives — the overwhelming majority — can never be cached because the question is about an address rather than about a value.
A header answers the question at the object, where "not a box today, box tomorrow" cannot arise: the bytes either are a box or are not, and the check is a load and a compare.
Proposal: give Box a header word
#[repr(C)]
pub struct Box { pub header: u64, pub value: u64 } // 8 -> 16 bytes
What that removes:
BOX_REGISTRY entirely — is_registered_box_ptr becomes a tag compare on a header the caller is about to touch anyway;
- the three 8-slot
BoxPtrCachees in HotTls and their eviction/quarantine interplay;
is_plausible_box_ptr's heuristic, and with it the perry#4898 false-positive class (a read-only __TEXT.__cstring address that passes every structural check) — a tagged header cannot be forged by an unrelated allocation;
BOX_CAPTURE_COUNTS — the count becomes a field, so increment/decrement is a load/add/store instead of TLS + RefCell + hash.
Memory (estimate, labelled as such). Per live box today: 8 B of Box, plus a PtrHashSet entry (~10 B with control bytes and load factor), plus — when captured — a PtrHashMap<usize, usize> entry (~19 B). That is up to ~37 B. With a header: 16 B and no tables. So this should reduce footprint as well as work, which is unusual for a "make it faster" change and is worth measuring before and after. It also fits the GC census finding of 147 MB of side tables against a 23 MB live heap.
Scope and risks
Box, I32Box and BoolBox are all headerless (I32Box and BoolBox are align(8) wrappers around a scalar), so all three would change shape. Codegen emits box allocations and capture-slot reads/writes directly, so the field offset change reaches perry-codegen, not just the runtime.
- The GC traces boxes via the registry today; a header makes tracing more precise rather than less, but the tracing path must be moved over in the same change.
- Worth checking whether the header can be narrower than a word (a tag byte plus a 32-bit count would fit in one
u32 pair) before committing to +8 B.
Relationship to the wider pattern
Predicates answering "what kind of pointer is this?" are 18.5% of this profile — is_registered_box_ptr 5.2%, is_registered_buffer_slow 2.8%, classify_heap_generation_uncached 1.3%, resolve_strategy_slow 1.2%, get_parent_class_id 1.2%, is_class_object_ptr 1.1%, keys_find_slot_by_bytes 1.0%, plus is_registered_map and classify_arena. Add the capture-count tables and it is roughly 27% of the run spent on metadata that is stored beside objects instead of in them.
Arrays already do this the right way and it is measurably cheaper: receiver_may_be_registered_exotic reads the GC header's obj_type and is one warm byte plus a compare — see #10694, where wiring the indexing path to it halved 79.7 M probes. Boxes have no equivalent because they have no header to read.
Related: #10694, #10688.
Summary
A box has no header — eight bytes of payload and nothing else. It therefore cannot answer any question about itself, so every question is answered by a thread-local side table keyed by its address:
BOX_REGISTRY(PtrHashSet<usize>) — "is this pointer a box?", because the object carries no identity;BOX_CAPTURE_COUNTS(RefCell<PtrHashMap<usize, usize>>) — the capture-edge count, because the object carries no count field.Giving
Boxone header word would let both live in the object. On the profile below that is ~9.8% of a real run, and — because a hash entry costs more than the word it would replace — it also looks like a net memory win.Measured
tsc --noEmit demo.ts(two-line input) compiled with Perry 0.5.1596 + #10656/#10685,sampleleaf histogram, 4,917 attributed samples:is_registered_box_ptrBOX_REGISTRYprobe (+ plausibility heuristic + 8-slot cache)js_closure_set_box_capture_ptrincrement_cell_capture_countBOX_CAPTURE_COUNTShashdecrement_cell_capture_countBOX_CAPTURE_COUNTShashpublish_box_cellis_registered_box_ptris the single largest symbol in the whole profile, which is otherwise flat (next is 4.0%).Both counting paths are the same shape:
Why the current design cannot be optimised further in place
box.rs's own comment explains the ceiling, and it is correct:The same comment records that the positive cache already took this family from 8.2% + 5.9% + 5.5% down to today's numbers. That was the available win. What remains is structural: a slot holding a raw box pointer is indistinguishable from a slot holding a NaN-boxed value, so the runtime must ask an external authority, and negatives — the overwhelming majority — can never be cached because the question is about an address rather than about a value.
A header answers the question at the object, where "not a box today, box tomorrow" cannot arise: the bytes either are a box or are not, and the check is a load and a compare.
Proposal: give
Boxa header wordWhat that removes:
BOX_REGISTRYentirely —is_registered_box_ptrbecomes a tag compare on a header the caller is about to touch anyway;BoxPtrCachees inHotTlsand their eviction/quarantine interplay;is_plausible_box_ptr's heuristic, and with it the perry#4898 false-positive class (a read-only__TEXT.__cstringaddress that passes every structural check) — a tagged header cannot be forged by an unrelated allocation;BOX_CAPTURE_COUNTS— the count becomes a field, so increment/decrement is a load/add/store instead of TLS +RefCell+ hash.Memory (estimate, labelled as such). Per live box today: 8 B of
Box, plus aPtrHashSetentry (~10 B with control bytes and load factor), plus — when captured — aPtrHashMap<usize, usize>entry (~19 B). That is up to ~37 B. With a header: 16 B and no tables. So this should reduce footprint as well as work, which is unusual for a "make it faster" change and is worth measuring before and after. It also fits the GC census finding of 147 MB of side tables against a 23 MB live heap.Scope and risks
Box,I32BoxandBoolBoxare all headerless (I32BoxandBoolBoxarealign(8)wrappers around a scalar), so all three would change shape. Codegen emits box allocations and capture-slot reads/writes directly, so the field offset change reachesperry-codegen, not just the runtime.u32pair) before committing to +8 B.Relationship to the wider pattern
Predicates answering "what kind of pointer is this?" are 18.5% of this profile —
is_registered_box_ptr5.2%,is_registered_buffer_slow2.8%,classify_heap_generation_uncached1.3%,resolve_strategy_slow1.2%,get_parent_class_id1.2%,is_class_object_ptr1.1%,keys_find_slot_by_bytes1.0%, plusis_registered_mapandclassify_arena. Add the capture-count tables and it is roughly 27% of the run spent on metadata that is stored beside objects instead of in them.Arrays already do this the right way and it is measurably cheaper:
receiver_may_be_registered_exoticreads the GC header'sobj_typeand is one warm byte plus a compare — see #10694, where wiring the indexing path to it halved 79.7 M probes. Boxes have no equivalent because they have no header to read.Related: #10694, #10688.