Skip to content

Add a performance baseline harness for the data layer (#407) - #411

Draft
alex-rawlings-yyc wants to merge 1 commit into
mainfrom
perf/407-performance-baseline
Draft

alex-rawlings-yyc wants to merge 1 commit into
mainfrom
perf/407-performance-baseline

Conversation

@alex-rawlings-yyc

@alex-rawlings-yyc alex-rawlings-yyc commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Part of #407. This builds the datasets, instrumentation and harnesses; the baseline numbers come from a run on the reference machine.

Datasets. npm run perf:capture captures the WEB sample's USJ from a launched app into perf/.cache (gitignored). npm run perf:datasets generates seeded drafts from it: small (Philippians), medium (Psalms), large (whole Bible at 50%) and large-loadable (whole Bible at 30%). Each draft has stale, competing, phrase, free-translation and morpheme records, payloads shared by content, and saved projects for the picker. Generation fails unless the result has no invariant violations and re-anchoring leaves it unchanged, so loading one never writes. Small uses WEB's Philippians rather than test-data/pt9-projects, so every tier shares one text and one generator.

Instrumentation. src/utils/perf-marks.ts records ilz: performance entries at the phase boundaries: USJ fetch, tokenize, re-anchor, store mount, book render, dispatch to render, autosave, Save, project listing and concordance. It records them only while the interlinearizer.perfMarks localStorage key is true; otherwise every call is a no-op. mergeAnalyses moved beside splitAnalysisByBook so Node can time it.

Harnesses. npm run perf:bench times the pure functions and measures heap in Node. npm run perf:app drives the app over CDP: opening a fresh view, switching books, gloss edits, the heap with a full undo history, the picker, Save and the concordance. It builds the extension for production through global setup's new E2E_BUILD_SCRIPT. It also runs over empty extension storage and restores the developer's storage on teardown. Both harnesses write JSON to perf/results/, and e2e-tests/README.md has the procedure.

Limits. Baseline numbers need a packaged core, because a dev core serves the WebView development React. Reusing a packaged app launched with --extensions and --remote-debugging-port=9223 is untested. The large draft (119 MB) exceeds core's 100 MB websocket message cap, so only the Node bench measures it. Opening a view measures a freshly opened WebView, not the first load after app start.

First pass on a busy Linux machine against a dev core, so not a baseline:

  • Dispatch to render for one gloss edit: 15 ms small, 167 ms medium, 924 ms large-loadable; about 3 s for large in Node.
  • A full undo history adds about 20 MB of heap on small, 217 MB on medium and 1.4 GB on large-loadable.
  • Save on large-loadable sends 75 MB and takes 10 s.
  • Listing a source's projects reads every project in storage, whatever its source.

All four tiers ran end to end in the app on Linux, with large skipped as designed.


This change is Reviewable

@alex-rawlings-yyc alex-rawlings-yyc self-assigned this Oct 9, 2026
@coderabbitai

coderabbitai Bot commented Oct 9, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@alex-rawlings-yyc

alex-rawlings-yyc commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

[Generated by Claude, not yet human-reviewed]

Baseline: 2026-10-09 @ 7a19ef7

Raw results: perf/results/bench-2026-10-09-7a19ef71.json, perf/results/app-2026-10-09-7a19ef71.json, and perf/results/app-2026-10-09-7a19ef71-medium-rerun.json (see Caveats).

Machine: i7-12700KF (20 threads), 32 GB RAM, Windows 11 Pro (10.0.26200), Node v25.8.1. paranext-core 0d06e2f.
Protocol: extension built with build:production; app runs: 5 per scenario, bench runs: 7 (+1 warmup). Values are median (p95).

Tiers

Tier Books Coverage Word tokens Approved links Draft JSON Saved projects
small PHP 100% 2,219 2,149 0.70 MB 5
medium PSA 100% 41,324 40,110 11.3 MB 5
large-loadable all 82 30% 905,971 263,100 74.8 MB 1
large all 82 50% 905,971 438,790 118.9 MB 5

large is 133.4 MB on the wire, which is over paranext-core's 100 MB per-message websocket cap. The app can't load it, so it has bench numbers only.

1. Book-load latency (in app)

"Open" means a newly opened WebView that fetches the draft and the text from scratch. "Switch" means returning to the view book in a WebView that already holds the draft.

Phase small medium large-loadable
Open → interactive 634 ms (776 ms) 1,260 ms (1,349 ms) 3,499 ms (3,638 ms)
USJ fetch 205 ms (277 ms) 574 ms (623 ms) 2,383 ms (2,524 ms)
Draft fetch 78 ms (83 ms) 317 ms (332 ms) 2,083 ms (2,103 ms)
Draft parse 1.1 ms (1.2 ms) 16 ms (31 ms) 155 ms (163 ms)
Tokenize 1.9 ms (6.2 ms) 23 ms (31 ms) 21 ms (24 ms)
Reanchor 11 ms (14 ms) 150 ms (184 ms) 382 ms (417 ms)
Store mount 6.3 ms (7.4 ms) 19 ms (19 ms) 90 ms (93 ms)
Book render 342 ms (442 ms) 803 ms (852 ms) 2,803 ms (2,920 ms)
Switch: book render 75 ms (79 ms) 218 ms (227 ms) 254 ms (297 ms)
Switch: USJ fetch 16 ms (21 ms) 109 ms (117 ms) 110 ms (147 ms)
Switch: tokenize 1.3 ms (1.7 ms) 20 ms (25 ms) 21 ms (43 ms)
Switch: reanchor 6.8 ms (9.8 ms) 141 ms (143 ms) 329 ms (442 ms)

The bench times the same steps in Node, without IPC or rendering:

Step small medium large-loadable large
tokenize 1.7 ms (2.6 ms) 31 ms (34 ms) 35 ms (39 ms) 33 ms (36 ms)
parse draft 1.5 ms (1.8 ms) 22 ms (23 ms) 165 ms (174 ms) 281 ms (301 ms)
reanchor 8.1 ms (9.7 ms) 189 ms (199 ms) 303 ms (317 ms) 457 ms (476 ms)
store mount 4.1 ms (12.6 ms) 68 ms (75 ms) 468 ms (526 ms) 1,090 ms (1,891 ms)
read shards 1.6 ms (1.8 ms) 25 ms (32 ms) 235 ms (247 ms) 391 ms (404 ms)

2. Edit latency

Step small medium large-loadable large (bench only)
App: dispatch → render 5.4 ms (26 ms) 71 ms (97 ms) 444 ms (554 ms) —
App: autosave serialize 1.7 ms (1.8 ms) 30 ms (33 ms) 183 ms (193 ms) —
App: autosave round-trip (no debounce) 16 ms (30 ms) 253 ms (270 ms) 1,457 ms (1,494 ms) —
Bench: dispatch + selector rebuild 6.6 ms (10.6 ms) 119 ms (144 ms) 608 ms (712 ms) 1,186 ms (1,412 ms)
Bench: serialize draft 1.9 ms (1.9 ms) 31 ms (32 ms) 180 ms (183 ms) 292 ms (345 ms)
Bench: write shards 1.7 ms (1.8 ms) 33 ms (33 ms) 152 ms (164 ms) 309 ms (361 ms)

3. Memory

These are WebView JS heap deltas after a forced GC, measured against the heap before the view opened. "Full undo" means after 100 further edits.

small medium large-loadable large (bench only)
App: view loaded 56 MB 169 MB 352 MB —
App: with full undo history 70 MB 351 MB 1,877 MB —
Bench: parsed draft retained 1.1 MB 18 MB 119 MB 189 MB
Bench: undo history retained 3.0 MB 42 MB 340 MB 522 MB

4. Bytes on the wire

small medium large-loadable large
Per autosave (saveDraft) 0.79 MB 12.7 MB 83.8 MB 133.4 MB (over cap)
Per Save 0.70 MB 11.3 MB 74.7 MB —
getProjectsForSource response 3.9 MB (5 projects) 63.5 MB (5 projects) 83.8 MB (1 project) 667 MB (5 projects)

5. Project picker and Save (in app)

small medium large-loadable
Picker: getProjectsForSource 79 ms (139 ms) 294 ms (306 ms) 457 ms (495 ms)
Save 98 ms (145 ms) 517 ms (531 ms) 3,523 ms (3,608 ms)

6. Concordance (every book present)

small medium large-loadable large
App: read all books 1,249 ms (1,304 ms) 1,186 ms (1,311 ms) 1,382 ms (1,423 ms) —
App: build entries 46 ms (50 ms) 44 ms (49 ms) 47 ms (48 ms) —
Bench: full build 539 ms (564 ms) 512 ms (614 ms) 779 ms (803 ms) 1,057 ms (1,102 ms)

Observations

  • Memory is the sharpest problem. At large-loadable, a full undo history takes the WebView heap to about 1.9 GB above baseline. That's 5.3× the loaded view, and much more than the 340 MB the bench estimates, so the in-app snapshots retain far more than the Node model accounts for.
  • Every autosave re-sends the whole draft. At large-loadable that's 84 MB per edit and 1.5 s round-trip, and Save takes 3.5 s. This is the clearest target for the migration.
  • Opening the view is dominated by I/O, not computation. At large-loadable, USJ fetch (2.4 s) and draft fetch (2.1 s) dwarf tokenize, reanchor and store mount (about 0.5 s together).
  • Edit latency grows with the size of the whole draft, not the book being viewed. Medium and large-loadable both view PSA, yet dispatch → render goes from 71 to 444 ms. The selector index rebuild is linear in the analysis.
  • Concordance cost is fixed, about 1.2–1.4 s, almost all of it reading every book's text. It barely changes across tiers.

Caveats / gaps against the protocol

  • No true cold runs. "Open" is a fresh WebView in an app that was already running, not the first load after app start. "Switch" is the warm case. A first-load-after-launch scenario still needs adding.
  • The host was the dev renderer, not a packaged app. The extension was a production build.
  • Medium had a flaky failure. In the full run, the medium tier failed at the Save step because the Save menu item was detached from the DOM mid-click, and the spec's error handling dropped all of that tier's results. A medium-only re-run passed, and its numbers are what appear above (…-medium-rerun.json).
  • The working tree was dirty (dirty: true). It held uncommitted changes to e2e-tests/tests/perf/* (polling for the source project id). No data-layer code was changed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant