Repository navigation
Conversation
yeahjack
marked this pull request as ready for review
October 5, 2026 14:33
yeahjack
force-pushed
the
perf/shortcut-filter-20261005
branch
from
October 5, 2026 16:28
8301d93 to
758fa12
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Prefilter behavior shortcuts with the existing pure press/release matchers when the active catalog has at least 32 behaviors. Unrelated inputs avoid scanning behavior IDs. Smaller catalogs retain ID-first matching because parsing stale bindings can cost more than a short scan.
The production optimization is now only 8 added lines. The original inner ID-first matching/action body and following code remain verbatim; actual large-catalog hits intentionally recheck the pure matchers. Binding/catalog order, global priority, repeated downs, releases, expression selection and Alt+1…0 fallback are unchanged. No cache, allocation or persistent state is added.
Independent base:
4283de1599c7da914138f82a99405d67c2c861ec. This refresh contains only this optimization, its validation and the small build cleanups below; it does not restore functionality removed upstream or depend on another optimization PR.Verification
Fresh measurements on this base
GCC 14.2, Release
-Osplus IPO/LTO, Xeon Platinum 8573C shared Linux cloud host. Old/new implementations use the same executable, compiler settings and validator-checked fixtures. Three runs each contain three warmup pairs and 31 alternating measured pairs per scenario. Other project builds/tests were paused; no exclusive CPU or affinity isolation is claimed.Medians of the three run-level wall medians:
The first cached-match version caused an avoidable stale-31 regression. GCC/LTO inspection showed loop indices spilling to memory while match booleans occupied registers. The simpler guard restores both loop indices to registers. In three contemporaneous cached/guard comparisons, stale-31 changed from 0.884–0.943× to 1.094–1.687×, without tuning the threshold or fixtures. This supports removing the structural cost; noisy timings do not establish its precise causal magnitude.
Residual costs remain: one behavior with 64 stale bindings costs 6.60–37.01 ns/event more (0.9–7.2%) in these runs. The guard adds a decision per binding and the measured dispatcher stack reservation is 72 bytes versus baseline 56 (rejected cached version: 88); these are compiler-specific benchmark observations, not application memory claims. The threshold is an empirical tradeoff, not a universal speedup guarantee. Run 2 has substantial shared-host variation, so upper speedup ranges should not be treated as stable device performance.
All 78 scenario/run summaries and 2,418 final guard-only raw paired samples, wall/process-CPU measurements and reproduction commands are in
tests/runtime/README-shortcuts.mdand its linked CSV files. These are dispatch component costs with deterministic external-action doubles. Batch-average p95 values do not establish individual-event tails, whole-app CPU, FPS or input-to-frame latency. Bounded calibration and shared-host noise limit small-difference conclusions. No previous-base measurements are reused.Build compatibility and scope
Two small fixes for warnings inherited from the current baseline are kept separate from the optimization: split two independent audit
ifstatements onto separate lines, and compileremove_receiptunder the sameBONGO_CAT_HAS_CUBISMcondition as its existing caller. The helper body and call path are unchanged. No warning exemptions are required.Linux diagnostic-backend coverage does not establish native Windows/macOS behavior, licensed Cubism rendering, GPU/window correctness or input-to-photon latency. LeakSanitizer is unavailable under this executor's tracing; leak checking is not claimed.
Current-head review and CI status
Published head:
758fa1294b2b169c8b8455ada637d45e86c802e2; tested tree:63c6bb13e6de74ea1236d12ffd974c6d10202c69. The tree was verified against the tested local source after a protected, expected-old-head ref update. GitHub reports no merge conflict and the PR is Ready for review. No merge, auto-merge or deployment was performed.CI has not passed: the current-head workflow run completed as
action_required, with zero executed jobs and zero check runs. A repository maintainer must approve the workflow before cross-platform jobs can run. The fork has no run for this head. Local strict builds and focused tests are the established evidence; Ready for review is not a ready-to-merge claim.