feat(bench,physics): finalize reference benchmarks and connected-body collision control - #732
Conversation
WebKit's WebGPU adapter substitution reports `apple apple apple apple`. `dropLeadingNoise` keeps a last part rather than reducing to the empty string, so that survived as the bare vendor word `apple` - which the category-word check accepted, because `apple` is a vendor rather than a category. The adapter preference then chose it over the WebGL2 stamp's `Apple GPU`, and one machine's runs published as `apple-macos-27-beta-webkit.json` instead of falling back to the CPU model and overwriting `m3-max-macos-27-beta-webkit.json`. That is how the same Mac came to hold two profiles. A part made only of vendor and product-line words now identifies no machine, for the same reason a category word does not. The superseded physics-only profile for that machine is removed with it.
A push carrying only a published benchmark profile rebuilt `dist` and ran all four gate groups - nine typechecks, ten lints, the sync and site groups - to validate a JSON file none of them opens. The one check that does read it, `verify:bench-results`, sits inside the lint group and was reached last. `scripts/ci/push-scope.ts` classifies the pushed range instead: a range that changes no tracked file verifies nothing, one carrying published profiles alone runs only their own gate, and anything else - including a range whose scope cannot be determined - takes the full path unchanged. The hook stays a convenience rather than an authority, so narrowing here is safe: CI re-runs its own selection on every push regardless.
The exclusion rested on a 45 % pixel difference against Pixi and on ExoJS issuing twenty draw calls to Pixi's three. The draw count was fixed upstream; the picture difference turned out to be two mistakes in this arm, not in either filter. `BlurFilter.strength` is the spacing between taps, not a standard deviation - the weights are a fixed table that already is a sigma-2 kernel - so handing it the archetype's sigma blurred the Pixi arm about twice as far as the contract asks. And the two filters derive their reach from different multiples of the blur they were given, which left a band a few pixels deep at the filtered region's border. With both corrected the arms agree to within 18 of 255 on the worst channel of the worst pixel at every rung, over 98 % of pixels inside 8. Evidence, including the numbers before each correction, is in the private workspace. Phaser stays out: its blur is an iterative step filter on a camera, with no kernel to configure to the shared contract.
The per-cell summary was `key=value` runs padded by hand, so two arms of one scenario could not be compared without parsing a line first - and they were printed arm-major, which put them nowhere near each other. Now one table, grouped by scenario and load, in both domains. The README was 748 lines of contract with the commands at the top. The contract moves to docs/harness.md unchanged; the README becomes what someone needs to run a benchmark, read what it printed, and contribute their machine's numbers - and says how to do that last part, which nothing pointed at before. `results/README.md` still documented a three-run recipe per domain; the harness has measured both domains in one `bench:reference` invocation for a while.
…ed table `coversArchetype` existed twice - once in each adapter, once in the driver's capability list - and only the driver's copy decides which cells are built. A clause present in an adapter alone therefore changed nothing: the cell was created anyway and `buildScene` fell through to the ordinary sprite scene. Two rows were being published that way. The Excalibur arm implements neither hit testing nor a particle system, and was measured on both: `interaction-picking` and the particle scenarios were sprite counts under a name that promised something else. Its tilemap clause had drifted the same way, as had the culled Pixi arm's. The predicates move to `coverage.ts` and both sides import them. The summary table is bordered now, with a rule between scenarios so the arms of one row read as a group.
…row publish Three pooled `bench:reference` runs on the RTX 5070 Ti, Chromium. The profile gains the seven rendering scenarios it never carried, plus `fx-blur` and `ui-layout-update`. `ui-layout-update` did not reach the table on the first attempt. The mechanism rule reads zero draw calls on both arms as a cell that failed to render its scene, which was true of every archetype until this one: it measures a box-tree solve and its leaves have no painted surface, so drawing nothing is the correct behaviour and the zeros are the evidence. A drawless archetype now keeps its counters and says so in its mechanism sentence. `fx-blur` publishes a loss: Pixi leads it by 1.3x on WebGPU and 1.4x on WebGL2.
The page falls back to the raw archetype id when it has no title, so the new row would have shipped as `ui-layout-update` beside `Hit testing` and `Blur effect`. The fallback stays - the page has to render a profile from a newer harness than it knows - and a test now keeps the repository's own profiles away from it.
…scenario The archetype had a catalog entry, a ladder and no arm, so every reference run resolved 23 of 24 scenarios and nothing said so. The suite-plan test skipped any catalogued scenario without an archetype ladder, which is exactly the shape of that gap; it now asserts the ids instead. ExoJS lays the tree out through `Stack`, Pixi through `@pixi/layout` over Yoga. Phaser 4 and Excalibur 0.32 ship no layout engine at all and sit the archetype out, so this is a published two-arm comparison. The harness checks the resolved rectangles against a closed-form model of the shared definition rather than against the other arm, and the structural probe's "a non-empty scene must draw" self-check is exempted here: this is the one scene that submits nothing by design. Found on the way: a nested `Stack` was pinned to an explicit box the first time its parent laid it out, because the parent wrote back a size the child already had. It then stopped sizing itself to its content for the rest of its life.
Not everyone installs from npm, and nothing on the landing page pointed at the release bundles. The link targets `/releases/latest` rather than a tag built from the package version: GitHub resolves it to whatever is newest, while the version in `package.json` is the one being developed and is ahead of the last release for most of a cycle.
…able ZIP name A squash commit's body is the pull request description, and the changelog put all of it under the bullet. Fifteen of those end to end made the v0.17.0 release page several thousand words of unbroken prose - the entries stopped being scannable, which is the one thing release notes are for. An entry now leads with its opening paragraph and keeps the rest in a `<details>` fold: nothing is lost, neither on the release page nor in `CHANGELOG.md`. The release also uploads a second copy of the Full ZIP under a version-less name. GitHub serves the newest release's assets at `/releases/latest/download/<asset>`, which is a permanent direct-download URL only for a name that does not move with the version - so the site's download button can link the archive itself rather than the release page.
…nd the solver's own softness
Bundle ReportChanges will increase total bundle size by 19.56MB (60.55%) ⬆️
Affected Assets, Files, and Routes:view changes for bundle: exo-esm-esmAssets Changed:
view changes for bundle: exojs-physics-esmAssets Changed:
view changes for bundle: exo-esm-modules-esmAssets Changed:
view changes for bundle: site-server-esmAssets Changed:
App Routes Affected:
view changes for bundle: exo-iife-min-Exo-iifeAssets Changed:
view changes for bundle: exo-full-iife-Exo-iifeAssets Changed:
view changes for bundle: exo-iife-Exo-iifeAssets Changed:
view changes for bundle: exo-full-iife-min-Exo-iifeAssets Changed:
|
Exoridus
left a comment
There was a problem hiding this comment.
Visual review: release-approved from my side. The benchmark overview now reads cleanly on desktop and mobile in both themes; editorial card pairing, fastest-first comparable rows, wide mobile bars, aligned amber frame-budget values, and the Overview / All measurements / Methodology hierarchy are coherent. I found no visual release blocker in the final screenshot set.
The PR CI run is green, including the unit lane (pnpm test + allocation + physics perf), so the one unexplained local 1/13 Physics-suite failure did not reproduce in the authoritative PR run.
GitHub does not allow the PR author to formally approve their own PR, so this is recorded as a review comment rather than an APPROVE review.
Finalizes the reference benchmarks and their page, and lands the physics work the page's own numbers turned up: a
collideConnectedoption on every joint, and a fix for a revolute-chain solver defect the benchmark's seam contacts had been masking. One branch, one PR, as agreed; nothing in it is split out.1. Benchmark and reference architecture
Adds the Nape-JS arm, the
ui-layout-updateandfx-blurarms, one definition of which archetypes an arm covers, the reference-run scope with its push gates, a bordered run summary, and a README that reads as a guide. Two reference profiles are published:rtx-5070-ti-windows-11-chromiumandm3-max-macos-27-beta-webkit, each pooled from three runs, both re-measured on the current arms.2. Benchmark site: redesign and final polish
One page, rendering and physics together, built at build time from the profiles under
packages/exojs-bench/results/and from nothing else. Six rendering and four physics scenarios open each section, fixed inbench-cards.tsbefore any run happens - the rendering six in three editorial pairs (a scene that never changes beside one where everything does, text beside tiles, an effect beside clipping) so the top of the page can never become a selection of whatever ExoJS won. The rows of a comparable load read quickest first, because the section says lower is better; ExoJS is found by its colour. A withheld row keeps the canonical order, since it publishes no ranking.The axis is the browser, not the machine.
Chromium | WebKit, Chromium first as an editorial choice. Switching machines changes CPU, GPU, OS, driver and engine at once, which is not an A/B of anything a reader can name. It is still not a browser benchmark while each browser is measured on its own machine: the machine is printed beside every set of figures as provenance, nothing is averaged across profiles, and one browser with several machines links to the measurements page instead of growing a second row of buttons.Three depths, one data source.
/benchmarks/is the overview;/benchmarks/full/("All measurements") is every row with spread, p95, mechanism and omissions;/benchmarks/methodology/is how the numbers are made. The old page carried all three stacked.On a phone the scenarios are sections of one card per domain: label and time on one line, the bar full width beneath, one inset separator, no outer contour. Same markup, same data, same per-scenario detail. The light theme's card surfaces sit a step above the canvas instead of white on near-white - scoped to this page rather than retuned in the theme.
Each card opens its own detail, rendered from that load's comparisons, so the load, backend and profile a reader sees are the ones the numbers were read out of by construction. A figure past the 16.7 ms frame budget is drawn as the same figure in amber - same box, same alignment, no glyph - and is the control that opens the note explaining it; one legend on the page says what amber means and that the measurement is valid. Fixed arm naming (
Nape-JS), proportional bars with no minimum width, a shared time column per card, selection marked by weight and ground as well as hue, the global.panelrule kept out of the card.3. Measurement and timer correctness
An unmeasured comparison was published as
0.00 ms. A physics arm the clock could not separate stores a zero, and the card drew it with a bar - the fastest figure on the page for the arm that produced nothing, while the same cell's detail saidnot comparable.publishedMswithholds a figure the cell never established; a bar is a duration, so a pair the clock could not separate keeps its times and loses only the factor and the winner, stated once per row and read off the actual outcome.formatMsprints<0.001rather than rounding a positive value to zero, and drops the trailing zeros the significant-figure rule reached for.The table showed a different load than the card. Rows are now keyed on archetype and load, with a load column; the global "measured at N nodes" note that printed
0 nodesis gone.A clock that was never observed read as a perfect one.
physicsClock()returnedresolutionMs: 0, which the batch sizing and the coarse-timing note read as the finest clock on record. It isnullnow and every consumer says "not observed". The note no longer printsInfinity% quantisationfor a sample the clock returned as zero, and names the batch size instead of asserting why the target was missed.The GPU line claimed something the profile does not record - the frame bracket, hardware timer or present cadence, with no mark saying which - and is no longer printed under that name.
4. Connected-body collision: API, lifecycle, CCD
JointOptions.collideConnected, shared by Distance, Revolute, Weld, Prismatic and Wheel, public and readonly. Defaulttrue- the behaviour every joint has had - as a compatibility choice; the reason it had to staytrueis gone (section 5) and the default is free to be revisited.falsetakes the pair out of collision before the narrow phase through a pair-keyed, reference-counted map on the world: not a category/mask filter, which cannot say "this pair and no other". No hot-path allocation; the key is numeric and the lookup sits behind a size check. CCD applies the same map, so a swept bullet is not stopped by the neighbour the discrete path may not collide it with.enabledstays independent: disabling a joint suspends its constraint and does not hand a ragdoll's limbs to the contact solver.Lifecycle: a body's teardown removes its incident joints and releases their pair claims (a joint left behind is prepared, warm-started and solved against a destroyed body every step, and holds a claim nothing can release);
addJointrefuses a joint on one body or a destroyed one immediately and, at the deferred point where membership is settled, one whose body belongs to another world - while the private unattached anchor aMouseJointstands in for its second body stays legal. Duplicate add and duplicate remove leave the count where it was.Every rule has a test, and the CCD one was falsified: with the sweep's filter removed it is the one test that fails.
5. Revolute-chain stability fix
A hanging revolute chain of seven or more equal links, left to itself, gained speed without bound along its own axis: 8 links reached ~1760 px/s inside 100 steps, 12 diverged, while 2 to 6 settled and slept. The figures were identical at 1, 10, 100 and 1125 chains, so it was one chain's behaviour and not a scaling effect, and
|ω|maxstayed at zero throughout - the links stretched and over-corrected, they did not swing. WithcollideConnected: truethe same chains held still: the seam contacts were damping the defect, not causing it.Two causes, each load-bearing. The point constraint solved every sub-step against the anchor error measured at frame start, so each sub-step re-corrected an error the previous one had already taken out; it now rotates its arms by the body's accumulated sub-step rotation and re-derives the anchor error from the accumulated deltas, the way the contact solver always has, and its effective mass follows the live arms. And the rigid joint path applied that correction through an unscaled Baumgarte bias (
0.2 / h, mass scale 1, impulse scale 0) that relaxed none of the impulse it accumulated; the rigid path now takes the solver's own soft factors, derived from twice the contact stiffness, which is the formulation the contacts already use. Reverting either half alone fails a test: the first breaks an existing soft-joint test, the second the new chain tests.After the fix every length from 2 to 16 links and the benchmark's full 1125-chain scene settle to rest and sleep, hanging within a few pixels of where they were built. The regression test runs 7-, 8- and 12-link chains for 600 steps and asserts invariants rather than numbers: finite, within 8 px of the rest position, peak speed below the settling motion, at rest at the end, and asleep. Distance, Weld and Wheel take the same rigid softness; the whole physics suite, the allocation gate and the perf gate are green.
6. Benchmark definition and reference profiles
The joint scene's connected-body semantics is now set explicitly on every arm (
collideConnected=falseon ExoJS and Planck,setContactsEnabled(false)on Rapier,Constraint.ignore=trueon Nape-JS; Matter.js has no per-constraint switch and is documented as running at its defaults, where it produces none). A small diagnostic run on this machine has all five arms at zero contacts.The published profiles predate that alignment and were not regenerated. The
jointsrow therefore staysNot comparableon the page - every arm's time visible, no bar, no factor, no winner - with the reason stated on the card, in its details, in the measurements table and in the methodology: the arms were not doing the same work when the profiles were measured, the harness now configures them to, and the comparison returns with the next reference measurement. No other benchmark definition changed, and no profile was touched for a UI fix.Verification
Physics: 557 tests including the chain regression, allocation gate, perf gate - one run of thirteen reported a single failure that did not name itself and did not reproduce in the twelve others, all of which passed clean. Site: 160 tests including the invalid-value rules against the published profiles,
astro check, full build. Bench: typecheck and lint. Root typecheck.eslint --max-warnings=0everywhere. Driven in a browser: markers open their note by pointer and by Enter, the only state chip the profiles produce isTiming-limited, switching profiles leaves exactly one stamp and one panel visible, a card switched to another load opens a detail naming that load. Screenshots at 1440x1000 and 390x844, dark and light, all three routes.