Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. String.prototype.replace/match/split with an untouched
RegExp spend most of each call reading rx.flags (whose getter performs 8 more observable Gets) and, for split,
rx.constructor[@@species], each through js_reflect_get inside a setjmp trap frame. On a 3-character subject this makes
replace 181×, match 148–212× and split 190× Node, while the matching itself is ~7 % of the time.
Reproduction
bench.ts:
// RegExp-taking String methods on a pristine RegExp, tiny subject (engine work negligible)
const variant = process.argv[2] || "replace"; const N = Number(process.argv[3] || "300000");
const RE_G: any = /\+/g, RE: any = /\+/, s = "a+b", b64 = "ndKK8/0z8Qw+OyczXXftx3rfjSEDizNo8Sb5awzd7Fw=";
const ops: Record<string, () => number> = {
flags: () => RE_G.flags.length, // the getter alone (8 inner Gets)
global: () => (RE_G.global ? 1 : 0), // control: one flag getter
replace: () => s.replace(RE_G, " ").length, // reads rx.flags, then matches
replace_empty: () => s.replace(RE_G, "").length, // control: canonical removal path skips the flags Get
replaceAll_str: () => s.replaceAll("+", " ").length, // control: no RegExp
match_g: () => s.match(RE_G)!.length, // reads rx.flags
match: () => s.match(RE)!.length, // reads rx.flags (non-global)
test: () => (RE.test(s) ? 1 : 0), // control: exec Get skipped (#10166)
split: () => s.split(RE).length, // constructor + @@species + flags + new sticky RegExp
split_str: () => s.split("+").length, // control: string separator
jwt_chain: () => b64.replace(/=/g, "").replace(/\+/g, "-").replace(/\//g, "_").length,
};
const op = ops[variant];
function run(n: number): number { let acc = 0; for (let i = 0; i < n; i++) acc += op(); return acc; }
run(N / 5 | 0); const t0 = performance.now(); const cs = run(N);
console.log(`variant=${variant} checksum=${cs} ms=${(performance.now() - t0).toFixed(2)}`);
PERRY_NO_AUTO_OPTIMIZE=1 perry compile bench.ts -o bench
node bench.ts replace 200000; ./bench replace 200000 # repeat per variant
Measurements
N = 200,000. Median of 3, shared host (load average ~100+ on 64 threads; instruction counts are the load-independent
figure). "Instructions/iter" = whole-process instructions:u ÷ (1.2·N timed + warm-up iterations).
| variant |
Node loop ms |
Perry loop ms |
ratio |
Perry instructions/iter |
Node wall ms |
Perry wall ms |
re.flags |
5.2 |
2,201.5 |
420× |
72.5 k |
132 |
2,639 |
re.global (control) |
1.7 |
202.6 |
119× |
7.0 k |
177 |
282 |
s.replace(/\+/g, " ") |
19.2 |
3,473.2 |
181× |
117.6 k |
148 |
4,059 |
s.replace(/\+/g, "") (control) |
17.3 |
587.0 |
34× |
20.1 k |
166 |
744 |
s.replaceAll("+", " ") (control) |
19.4 |
662.3 |
34× |
26.9 k |
163 |
892 |
s.match(/\+/g) |
20.1 |
2,984.6 |
148× |
94.7 k |
185 |
3,577 |
s.match(/\+/) |
12.7 |
2,696.1 |
212× |
90.3 k |
157 |
3,262 |
/\+/.test(s) (control) |
7.2 |
220.4 |
30× |
6.3 k |
110 |
307 |
s.split(/\+/) |
25.9 |
4,922.8 |
190× |
155.3 k |
175 |
5,964 |
s.split("+") (control) |
12.1 |
308.9 |
26× |
12.3 k |
193 |
408 |
| jws base64url chain |
98.9 |
8,615.5 |
87× |
286.6 k |
287 |
10,369 |
Checksums identical. The same replace with an empty replacement (which already takes a proven-builtin path that skips
the Gets) costs 20 k instructions instead of 118 k; non-global match costs 90 k against 6 k for test on the same RegExp.
Profile (perf record -g, inclusive): replace — flags getter 55 %, get_symbol(@@replace) 5.8 %, engine
execute_output 7 %; match(/g) — flags getter 67 %, get_symbol(@@match) 7.9 %, engine 7.5 %; split — flags getter
40 %, species 22 %, perex_construct::construct 6.8 %. Top self symbols are the generic property path:
try_data_get_bytes, from_utf8, js_reflect_get, get_accessor_descriptor, js_object_get_field_by_name,
exception::try_push_with_kind.
Impact
From the audit profiles (v0.5.1587, strings and objects group reports); "RegExp protocol Gets" was a separate bucket
from regex-engine time:
- date-fns 4.x
format: 38.8 % of the mode, ≈34.5 % of the date-fns loop (formatStr.match(longFormattingTokensRegExp) and
.match(formattingTokensRegExp), both global).
- qs 6.x parse: 29 % (
utils.decode: str.replace(/\+/g, ' ') per key and value).
- validator 13.x: isEmail 18.7 %, isURL 10.8 % (
isByteLength: encodeURI(str).split(/%..|./)).
- dayjs
format: 9.8 % (str.replace(REGEX_FORMAT, fn)).
- jsonwebtoken 9.0.3 sign:
js_string_replace_js 40 % of sign, verify 13 % (jwa/jws base64url
.replace(/=/g,'').replace(/\+/g,'-').replace(/\//g,'_'); includes engine time).
Group mean across the four string packages: ≈16.7 % of Perry time.
Mechanism
All citations read at 7661bc0 (verified) unless marked.
crates/perry-runtime/src/regex/perex_replace.rs:69 — RegExp.prototype[@@replace] does dispatch::get(&receiver, b"flags");
:239 String.prototype.replace first looks up @@replace via dispatch::get_symbol. The only bypass is :218
perex_remove::try_remove, which admits an empty replacement string only (perex_remove.rs:17-35).
crates/perry-runtime/src/regex/perex_match_search.rs:59-67 match_flags → dispatch::get(receiver, b"flags"), used for
global and non-global match (:187); :274/:300 look up @@match first.
crates/perry-runtime/src/regex/perex_split.rs:126-127 — match_all::species then get(flags); :155-165 constructs a new
splitter RegExp (js_regexp_construct) on every call; :335 looks up @@split.
crates/perry-runtime/src/regex/match_all.rs:86-101 species: Get(rx, "constructor") + Get(C, @@species).
crates/perry-runtime/src/regex/perex_match_search.rs:29-51 — the flags getter performs 8 dispatch::gets (hasIndices,
global, ignoreCase, multiline, dotAll, unicode, unicodeSets, sticky).
crates/perry-runtime/src/regex/perex_dispatch.rs:52-63 get: a RuntimeHandleScope, key interning
(canonical_key), an api::caught setjmp trap frame, and js_reflect_get; each flag getter then enters
object/regex_proto_thunks.rs:755 regexp_get_property inside a second catch_js_throw frame.
- The proof needed to skip all of this already exists:
crates/perry-runtime/src/object/regex_canonical.rs:45-84 computes a
per-prototype-ShapeId Proof { flags: bool, … } (all nine flag accessors are the native thunks), :89-123 exec() admits
an untouched RegExp, and :127-155 replace() adds the @@replace check — but only perex_remove and
perex_dispatch.rs:118 consult it; flags, @@match, @@split and @@species are never answered from it.
What fast looks like
When regex_canonical::exec(rx) and its flags proof hold, read global/unicode/sticky straight from the
RegExpHeader (as perex_remove already does), answer @@replace/@@match/@@split and the species constructor from the
same shape proof, and let split reuse the cached non-sticky program without allocating a splitter object. Targets on
this microbenchmark: s.replace(/\+/g, " ") ≤ the empty-replacement control (≈20 k instructions/call), s.match(re) and
s.split(re) within 1.5× of the test / string-split controls, re.flags ≤ 2 k instructions. Getting to ≤2× Node
additionally needs the per-call engine cost tracked in #10166.
Notes
Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64.
String.prototype.replace/match/splitwith an untouchedRegExp spend most of each call reading
rx.flags(whose getter performs 8 more observable Gets) and, forsplit,rx.constructor[@@species], each throughjs_reflect_getinside a setjmp trap frame. On a 3-character subject this makesreplace181×,match148–212× andsplit190× Node, while the matching itself is ~7 % of the time.Reproduction
bench.ts:Measurements
N = 200,000. Median of 3, shared host (load average ~100+ on 64 threads; instruction counts are the load-independent
figure). "Instructions/iter" = whole-process
instructions:u÷ (1.2·N timed + warm-up iterations).re.flagsre.global(control)s.replace(/\+/g, " ")s.replace(/\+/g, "")(control)s.replaceAll("+", " ")(control)s.match(/\+/g)s.match(/\+/)/\+/.test(s)(control)s.split(/\+/)s.split("+")(control)Checksums identical. The same
replacewith an empty replacement (which already takes a proven-builtin path that skipsthe Gets) costs 20 k instructions instead of 118 k; non-global
matchcosts 90 k against 6 k forteston the same RegExp.Profile (
perf record -g, inclusive):replace— flags getter 55 %,get_symbol(@@replace)5.8 %, engineexecute_output7 %;match(/g)— flags getter 67 %,get_symbol(@@match)7.9 %, engine 7.5 %;split— flags getter40 %,
species22 %,perex_construct::construct6.8 %. Top self symbols are the generic property path:try_data_get_bytes,from_utf8,js_reflect_get,get_accessor_descriptor,js_object_get_field_by_name,exception::try_push_with_kind.Impact
From the audit profiles (v0.5.1587, strings and objects group reports); "RegExp protocol Gets" was a separate bucket
from regex-engine time:
format: 38.8 % of the mode, ≈34.5 % of the date-fns loop (formatStr.match(longFormattingTokensRegExp)and.match(formattingTokensRegExp), both global).utils.decode:str.replace(/\+/g, ' ')per key and value).isByteLength:encodeURI(str).split(/%..|./)).format: 9.8 % (str.replace(REGEX_FORMAT, fn)).js_string_replace_js40 % of sign, verify 13 % (jwa/jws base64url.replace(/=/g,'').replace(/\+/g,'-').replace(/\//g,'_'); includes engine time).Group mean across the four string packages: ≈16.7 % of Perry time.
Mechanism
All citations read at 7661bc0 (verified) unless marked.
crates/perry-runtime/src/regex/perex_replace.rs:69— RegExp.prototype[@@replace] doesdispatch::get(&receiver, b"flags");:239String.prototype.replace first looks up@@replaceviadispatch::get_symbol. The only bypass is:218perex_remove::try_remove, which admits an empty replacement string only (perex_remove.rs:17-35).crates/perry-runtime/src/regex/perex_match_search.rs:59-67match_flags→dispatch::get(receiver, b"flags"), used forglobal and non-global
match(:187);:274/:300look up@@matchfirst.crates/perry-runtime/src/regex/perex_split.rs:126-127—match_all::speciesthenget(flags);:155-165constructs a newsplitter RegExp (
js_regexp_construct) on every call;:335looks up@@split.crates/perry-runtime/src/regex/match_all.rs:86-101species:Get(rx, "constructor")+Get(C, @@species).crates/perry-runtime/src/regex/perex_match_search.rs:29-51— theflagsgetter performs 8dispatch::gets (hasIndices,global, ignoreCase, multiline, dotAll, unicode, unicodeSets, sticky).
crates/perry-runtime/src/regex/perex_dispatch.rs:52-63get: aRuntimeHandleScope, key interning(
canonical_key), anapi::caughtsetjmp trap frame, andjs_reflect_get; each flag getter then entersobject/regex_proto_thunks.rs:755regexp_get_propertyinside a secondcatch_js_throwframe.crates/perry-runtime/src/object/regex_canonical.rs:45-84computes aper-prototype-ShapeId
Proof { flags: bool, … }(all nine flag accessors are the native thunks),:89-123exec()admitsan untouched RegExp, and
:127-155replace()adds the@@replacecheck — but onlyperex_removeandperex_dispatch.rs:118consult it;flags,@@match,@@splitand@@speciesare never answered from it.What fast looks like
When
regex_canonical::exec(rx)and itsflagsproof hold, readglobal/unicode/stickystraight from theRegExpHeader(asperex_removealready does), answer@@replace/@@match/@@splitand the species constructor from thesame shape proof, and let
splitreuse the cached non-sticky program without allocating a splitter object. Targets onthis microbenchmark:
s.replace(/\+/g, " ")≤ the empty-replacement control (≈20 k instructions/call),s.match(re)ands.split(re)within 1.5× of thetest/ string-split controls,re.flags≤ 2 k instructions. Getting to ≤2× Nodeadditionally needs the per-call engine cost tracked in #10166.
Notes
test; its fix skipped onlythe
execGet) and perf(regex): replace with a string template holds ~1 KB per output piece — 545 MB RSS on a 550 KB subject, 4.5x Node #10411 (template replacement memory). Those issues do not cover theflags/@@species/symbol Gets.flags/execproperty, a patchedRegExp.prototypegetter orSymbol.speciesmust still be observed; the ShapeId + symbol-epoch proof inregex_canonical.rsis designed for exactlythat.
splitcontrol is itself 26× here; long string splits are covered separately in perf:str.split(".")with a string separator is 345× slower than Node on a 179-char JWT (Perex KMP reads the subject one UTF-16 unit per bounded-reader call) #10519.