Repository navigation
Conversation
yeahjack
marked this pull request as ready for review
October 5, 2026 14:34
yeahjack
force-pushed
the
perf/mver-geometry-cache-20261005
branch
from
October 5, 2026 19:39
c1316ba to
bc04087
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cache exact-input Mver tablet geometry
Independent branch:
perf/mver-geometry-cache-20261005, based on666999650f8afb405fa34dbf1ee0b98cd7145117.Scope and preservation
six float inputs and handedness; check independently before and after keys
no-op. An all-mode prototype was narrowed after measurable mouse-miss overhead
scale, vertex generation, GL state changes and ordered draw calls
per instance in its existing allocation; no new heap allocation calls
Verification
preamble symbol remapping
changing color/scale/reference size/offsets, signed zero, subnormals, infinities,
distinct NaN payloads, deterministic arbitrary float patterns, failed geometry,
load/reload/clear and valid-cache-to-failed-load transitions covered
Other core/archive dependencies remain Release-built. LeakSanitizer cannot run
under this sandbox's ptrace setup; leak detection is not claimed
Final measurements
Intel Xeon Platinum 8573C, Linux x86-64, GCC 14.2.0, Release -Os + IPO for
baseline/candidate/core; deliberately non-IPO opaque GL/full-geometry sinks.
Pinned to allowed CPU 0. Three runs, each with 3 warm-up and 15 measured rounds.
Each variant performs 2,000 operations per trial in alternating 100-operation
AB/BA blocks. Project builds/tests paused during measurements. Shared-host noise
remains; all final raw data and analysis code are committed.
CPU timings, microseconds per geometry call or complete two-phase CPU draw:
Wall-clock medians (upstream → candidate, µs): stationary tablet 46.378 → 6.727;
changing tablet 45.578 → 25.779; mid-frame-changing tablet 45.792 → 45.544;
stationary mouse 27.235 → 26.587; changing mouse 26.233 → 26.001;
stationary geometry 18.939 → 0.135; changing geometry 19.345 → 19.224.
Full wall/CPU p95 and paired statistics are in the committed summary CSV.
Conservative interpretation: ~85% less CPU preparation for repeated tablet
inputs and ~42% less when tablet coordinates change once per frame. No claimed
benefit for mouse controls or two-miss tablet frames; their differences are
noise-scale. The original all-mode prototype produced +1.17%, +3.31%, +2.56%
paired moving-mouse CPU overhead in the improved protocol, so that path was
excluded. Exploratory runs 1–5 remain in the shared study results, not this PR.
These are component costs, not end-to-end input latency. p95 is the percentile
of batch averages, not input-event tail latency. No actual native window, GPU
rendering, Cubism SDK, OS input-capture, compositor or display latency was tested.
Identical ordered GL commands/payloads support unchanged visual output under
identical GPU state; actual native-window screenshots are not claimed.
Reproduction
See
tests/performance/MVER.mdandtests/performance/results/mver-20261005/README.md.Final raw files:
run-6.csv,run-7.csv,run-8.csv; aggregate:summary-final.csv. Local study also retains all eight raw runs and build/testlogs under the shared results directory.
Re-review status (2026-10-05)
The production diff, frozen upstream reference, regression coverage and raw benchmark statistics were independently rechecked. No actionable issue was found in the reviewed scope; the production change was kept minimal without additional refactoring. The published tree
af19f1ba1c4c899eeaede44a7bf0ded7b95dcacematches the locally tested tree.The full Linux Release diagnostic build and all 7 CTests passed again. Native Windows/macOS, licensed Cubism rendering and actual GPU/window behavior remain unverified.
CI has not passed: the upstream workflow run for head
c1316baf5a02f38d57c83cf14fe5ce41c50c264cisaction_requiredwith zero executed jobs and zero check runs. A repository maintainer must approve the workflow before cross-platform CI can execute. Ready for review does not mean ready to merge.