Skip to content

perf: reuse exact-input tablet pointer geometry - #102

Open
yeahjack wants to merge 5 commits into
vladelaina:mainfrom
yeahjack:perf/mver-geometry-cache-20261005
Open

yeahjack wants to merge 5 commits into
vladelaina:mainfrom
yeahjack:perf/mver-geometry-cache-20261005

Conversation

@yeahjack

@yeahjack yeahjack commented Oct 5, 2026 •

Copy link
Copy Markdown

Cache exact-input Mver tablet geometry

Independent branch: perf/mver-geometry-cache-20261005, based on
666999650f8afb405fa34dbf1ee0b98cd7145117.

Scope and preservation

  • Cache the last successful tablet geometry using exact representations of all
    six float inputs and handedness; check independently before and after keys
  • Keep the original mouse geometry computation path and original mouse after-key
    no-op. An all-mode prototype was narrowed after measurable mouse-miss overhead
  • Preserve all geometry arithmetic, textures, device/button selection, color,
    scale, vertex generation, GL state changes and ordered draw calls
  • Do not cache failures; invalidate on clear/load; detect changes between phases
  • Private overlay size on this Linux x86-64 build: 360 → 600 bytes, +240 bytes
    per instance in its existing allocation; no new heap allocation calls

Verification

  • Full diagnostic Release build, GCC 14.2.0, warnings as errors: pass
  • All 7 CTests: pass; source-line, localization, runtime safety policies: pass
  • 41,066 byte-exact geometry comparisons, 20,618 counted computations: zero failures
  • 8,308 draw-phase comparisons, 44,765,568 exact GL trace/payload bytes: zero differences
  • Frozen renderer body verified byte-for-byte against upstream, apart from test
    preamble symbol remapping
  • Dense input grids, handedness, all buttons, mouse/tablet, missing textures,
    changing color/scale/reference size/offsets, signed zero, subnormals, infinities,
    distinct NaN payloads, deterministic arbitrary float patterns, failed geometry,
    load/reload/clear and valid-cache-to-failed-load transitions covered
  • Focused ASan/UBSan checks of changed code and geometry calculation pass.
    Other core/archive dependencies remain Release-built. LeakSanitizer cannot run
    under this sandbox's ptrace setup; leak detection is not claimed
  • Independent read-only review found no production equivalence issue

Final measurements

Intel Xeon Platinum 8573C, Linux x86-64, GCC 14.2.0, Release -Os + IPO for
baseline/candidate/core; deliberately non-IPO opaque GL/full-geometry sinks.
Pinned to allowed CPU 0. Three runs, each with 3 warm-up and 15 measured rounds.
Each variant performs 2,000 operations per trial in alternating 100-operation
AB/BA blocks. Project builds/tests paused during measurements. Shared-host noise
remains; all final raw data and analysis code are committed.

CPU timings, microseconds per geometry call or complete two-phase CPU draw:

Scenario Upstream median Candidate median Upstream p95 batch mean Candidate p95 batch mean Paired median delta Paired change
Tablet, stationary 46.390 6.722 51.682 7.992 -39.247 -85.10%
Tablet, changing each frame 45.592 25.781 50.259 31.347 -19.138 -42.02%
Tablet, changed again between phases 45.787 45.558 54.095 51.564 -0.029 -0.07%
Mouse, stationary control 27.226 26.533 31.635 31.225 +0.090 +0.37%
Mouse, changing control 26.232 25.948 29.826 29.954 +0.146 +0.52%
Geometry helper, stationary 18.948 0.138 22.624 0.249 -18.711 -99.26%
Geometry helper, changing 19.354 19.233 23.003 23.536 -0.180 -0.95%

Wall-clock medians (upstream → candidate, µs): stationary tablet 46.378 → 6.727;
changing tablet 45.578 → 25.779; mid-frame-changing tablet 45.792 → 45.544;
stationary mouse 27.235 → 26.587; changing mouse 26.233 → 26.001;
stationary geometry 18.939 → 0.135; changing geometry 19.345 → 19.224.
Full wall/CPU p95 and paired statistics are in the committed summary CSV.

Conservative interpretation: ~85% less CPU preparation for repeated tablet
inputs and ~42% less when tablet coordinates change once per frame. No claimed
benefit for mouse controls or two-miss tablet frames; their differences are
noise-scale. The original all-mode prototype produced +1.17%, +3.31%, +2.56%
paired moving-mouse CPU overhead in the improved protocol, so that path was
excluded. Exploratory runs 1–5 remain in the shared study results, not this PR.

These are component costs, not end-to-end input latency. p95 is the percentile
of batch averages, not input-event tail latency. No actual native window, GPU
rendering, Cubism SDK, OS input-capture, compositor or display latency was tested.
Identical ordered GL commands/payloads support unchanged visual output under
identical GPU state; actual native-window screenshots are not claimed.

Reproduction

See tests/performance/MVER.md and
tests/performance/results/mver-20261005/README.md.
Final raw files: run-6.csv, run-7.csv, run-8.csv; aggregate:
summary-final.csv. Local study also retains all eight raw runs and build/test
logs under the shared results directory.

Re-review status (2026-10-05)

The production diff, frozen upstream reference, regression coverage and raw benchmark statistics were independently rechecked. No actionable issue was found in the reviewed scope; the production change was kept minimal without additional refactoring. The published tree af19f1ba1c4c899eeaede44a7bf0ded7b95dcace matches the locally tested tree.

The full Linux Release diagnostic build and all 7 CTests passed again. Native Windows/macOS, licensed Cubism rendering and actual GPU/window behavior remain unverified.

CI has not passed: the upstream workflow run for head c1316baf5a02f38d57c83cf14fe5ce41c50c264c is action_required with zero executed jobs and zero check runs. A repository maintainer must approve the workflow before cross-platform CI can execute. Ready for review does not mean ready to merge.

@yeahjack
yeahjack marked this pull request as ready for review October 5, 2026 14:34
@yeahjack
yeahjack force-pushed the perf/mver-geometry-cache-20261005 branch from c1316ba to bc04087 Compare October 5, 2026 19:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant