Repository navigation
Add a performance baseline harness for the data layer (#407) - #411
alex-rawlings-yyc wants to merge 1 commit into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
[Generated by Claude, not yet human-reviewed] Baseline: 2026-10-09 @
|
| Tier | Books | Coverage | Word tokens | Approved links | Draft JSON | Saved projects |
|---|---|---|---|---|---|---|
| small | PHP | 100% | 2,219 | 2,149 | 0.70 MB | 5 |
| medium | PSA | 100% | 41,324 | 40,110 | 11.3 MB | 5 |
| large-loadable | all 82 | 30% | 905,971 | 263,100 | 74.8 MB | 1 |
| large | all 82 | 50% | 905,971 | 438,790 | 118.9 MB | 5 |
large is 133.4 MB on the wire, which is over paranext-core's 100 MB per-message websocket cap. The app can't load it, so it has bench numbers only.
1. Book-load latency (in app)
"Open" means a newly opened WebView that fetches the draft and the text from scratch. "Switch" means returning to the view book in a WebView that already holds the draft.
| Phase | small | medium | large-loadable |
|---|---|---|---|
| Open → interactive | 634 ms (776 ms) | 1,260 ms (1,349 ms) | 3,499 ms (3,638 ms) |
| USJ fetch | 205 ms (277 ms) | 574 ms (623 ms) | 2,383 ms (2,524 ms) |
| Draft fetch | 78 ms (83 ms) | 317 ms (332 ms) | 2,083 ms (2,103 ms) |
| Draft parse | 1.1 ms (1.2 ms) | 16 ms (31 ms) | 155 ms (163 ms) |
| Tokenize | 1.9 ms (6.2 ms) | 23 ms (31 ms) | 21 ms (24 ms) |
| Reanchor | 11 ms (14 ms) | 150 ms (184 ms) | 382 ms (417 ms) |
| Store mount | 6.3 ms (7.4 ms) | 19 ms (19 ms) | 90 ms (93 ms) |
| Book render | 342 ms (442 ms) | 803 ms (852 ms) | 2,803 ms (2,920 ms) |
| Switch: book render | 75 ms (79 ms) | 218 ms (227 ms) | 254 ms (297 ms) |
| Switch: USJ fetch | 16 ms (21 ms) | 109 ms (117 ms) | 110 ms (147 ms) |
| Switch: tokenize | 1.3 ms (1.7 ms) | 20 ms (25 ms) | 21 ms (43 ms) |
| Switch: reanchor | 6.8 ms (9.8 ms) | 141 ms (143 ms) | 329 ms (442 ms) |
The bench times the same steps in Node, without IPC or rendering:
| Step | small | medium | large-loadable | large |
|---|---|---|---|---|
| tokenize | 1.7 ms (2.6 ms) | 31 ms (34 ms) | 35 ms (39 ms) | 33 ms (36 ms) |
| parse draft | 1.5 ms (1.8 ms) | 22 ms (23 ms) | 165 ms (174 ms) | 281 ms (301 ms) |
| reanchor | 8.1 ms (9.7 ms) | 189 ms (199 ms) | 303 ms (317 ms) | 457 ms (476 ms) |
| store mount | 4.1 ms (12.6 ms) | 68 ms (75 ms) | 468 ms (526 ms) | 1,090 ms (1,891 ms) |
| read shards | 1.6 ms (1.8 ms) | 25 ms (32 ms) | 235 ms (247 ms) | 391 ms (404 ms) |
2. Edit latency
| Step | small | medium | large-loadable | large (bench only) |
|---|---|---|---|---|
| App: dispatch → render | 5.4 ms (26 ms) | 71 ms (97 ms) | 444 ms (554 ms) | — |
| App: autosave serialize | 1.7 ms (1.8 ms) | 30 ms (33 ms) | 183 ms (193 ms) | — |
| App: autosave round-trip (no debounce) | 16 ms (30 ms) | 253 ms (270 ms) | 1,457 ms (1,494 ms) | — |
| Bench: dispatch + selector rebuild | 6.6 ms (10.6 ms) | 119 ms (144 ms) | 608 ms (712 ms) | 1,186 ms (1,412 ms) |
| Bench: serialize draft | 1.9 ms (1.9 ms) | 31 ms (32 ms) | 180 ms (183 ms) | 292 ms (345 ms) |
| Bench: write shards | 1.7 ms (1.8 ms) | 33 ms (33 ms) | 152 ms (164 ms) | 309 ms (361 ms) |
3. Memory
These are WebView JS heap deltas after a forced GC, measured against the heap before the view opened. "Full undo" means after 100 further edits.
| small | medium | large-loadable | large (bench only) | |
|---|---|---|---|---|
| App: view loaded | 56 MB | 169 MB | 352 MB | — |
| App: with full undo history | 70 MB | 351 MB | 1,877 MB | — |
| Bench: parsed draft retained | 1.1 MB | 18 MB | 119 MB | 189 MB |
| Bench: undo history retained | 3.0 MB | 42 MB | 340 MB | 522 MB |
4. Bytes on the wire
| small | medium | large-loadable | large | |
|---|---|---|---|---|
Per autosave (saveDraft) |
0.79 MB | 12.7 MB | 83.8 MB | 133.4 MB (over cap) |
| Per Save | 0.70 MB | 11.3 MB | 74.7 MB | — |
getProjectsForSource response |
3.9 MB (5 projects) | 63.5 MB (5 projects) | 83.8 MB (1 project) | 667 MB (5 projects) |
5. Project picker and Save (in app)
| small | medium | large-loadable | |
|---|---|---|---|
Picker: getProjectsForSource |
79 ms (139 ms) | 294 ms (306 ms) | 457 ms (495 ms) |
| Save | 98 ms (145 ms) | 517 ms (531 ms) | 3,523 ms (3,608 ms) |
6. Concordance (every book present)
| small | medium | large-loadable | large | |
|---|---|---|---|---|
| App: read all books | 1,249 ms (1,304 ms) | 1,186 ms (1,311 ms) | 1,382 ms (1,423 ms) | — |
| App: build entries | 46 ms (50 ms) | 44 ms (49 ms) | 47 ms (48 ms) | — |
| Bench: full build | 539 ms (564 ms) | 512 ms (614 ms) | 779 ms (803 ms) | 1,057 ms (1,102 ms) |
Observations
- Memory is the sharpest problem. At large-loadable, a full undo history takes the WebView heap to about 1.9 GB above baseline. That's 5.3× the loaded view, and much more than the 340 MB the bench estimates, so the in-app snapshots retain far more than the Node model accounts for.
- Every autosave re-sends the whole draft. At large-loadable that's 84 MB per edit and 1.5 s round-trip, and Save takes 3.5 s. This is the clearest target for the migration.
- Opening the view is dominated by I/O, not computation. At large-loadable, USJ fetch (2.4 s) and draft fetch (2.1 s) dwarf tokenize, reanchor and store mount (about 0.5 s together).
- Edit latency grows with the size of the whole draft, not the book being viewed. Medium and large-loadable both view PSA, yet dispatch → render goes from 71 to 444 ms. The selector index rebuild is linear in the analysis.
- Concordance cost is fixed, about 1.2–1.4 s, almost all of it reading every book's text. It barely changes across tiers.
Caveats / gaps against the protocol
- No true cold runs. "Open" is a fresh WebView in an app that was already running, not the first load after app start. "Switch" is the warm case. A first-load-after-launch scenario still needs adding.
- The host was the dev renderer, not a packaged app. The extension was a production build.
- Medium had a flaky failure. In the full run, the medium tier failed at the Save step because the
Savemenu item was detached from the DOM mid-click, and the spec's error handling dropped all of that tier's results. A medium-only re-run passed, and its numbers are what appear above (…-medium-rerun.json). - The working tree was dirty (
dirty: true). It held uncommitted changes toe2e-tests/tests/perf/*(polling for the source project id). No data-layer code was changed.
Part of #407. This builds the datasets, instrumentation and harnesses; the baseline numbers come from a run on the reference machine.
Datasets.
npm run perf:capturecaptures the WEB sample's USJ from a launched app intoperf/.cache(gitignored).npm run perf:datasetsgenerates seeded drafts from it: small (Philippians), medium (Psalms), large (whole Bible at 50%) and large-loadable (whole Bible at 30%). Each draft has stale, competing, phrase, free-translation and morpheme records, payloads shared by content, and saved projects for the picker. Generation fails unless the result has no invariant violations and re-anchoring leaves it unchanged, so loading one never writes. Small uses WEB's Philippians rather thantest-data/pt9-projects, so every tier shares one text and one generator.Instrumentation.
src/utils/perf-marks.tsrecordsilz:performance entries at the phase boundaries: USJ fetch, tokenize, re-anchor, store mount, book render, dispatch to render, autosave, Save, project listing and concordance. It records them only while theinterlinearizer.perfMarkslocalStorage key istrue; otherwise every call is a no-op.mergeAnalysesmoved besidesplitAnalysisByBookso Node can time it.Harnesses.
npm run perf:benchtimes the pure functions and measures heap in Node.npm run perf:appdrives the app over CDP: opening a fresh view, switching books, gloss edits, the heap with a full undo history, the picker, Save and the concordance. It builds the extension for production through global setup's newE2E_BUILD_SCRIPT. It also runs over empty extension storage and restores the developer's storage on teardown. Both harnesses write JSON toperf/results/, ande2e-tests/README.mdhas the procedure.Limits. Baseline numbers need a packaged core, because a dev core serves the WebView development React. Reusing a packaged app launched with
--extensionsand--remote-debugging-port=9223is untested. The large draft (119 MB) exceeds core's 100 MB websocket message cap, so only the Node bench measures it. Opening a view measures a freshly opened WebView, not the first load after app start.First pass on a busy Linux machine against a dev core, so not a baseline:
All four tiers ran end to end in the app on Linux, with large skipped as designed.
This change is