Skip to content

perf: howl_detector per-stage bench; the guard's cost on Linux GCC - #87

Merged
tap merged 7 commits into
mainfrom
perf/guard-linux-cost
Oct 9, 2026
Merged

tap merged 7 commits into
mainfrom
perf/guard-linux-cost

Conversation

@tap

@tap tap commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

HowlGuardHost.CostPerBlock recorded the guard at 8575 / 12794 ns per block (float, 1 / 2 mics) on the Linux GCC runner (job 110465810942) against 2180 / 4983 ns on the Intel Mac, while the canceller ran ~2x faster there. This PR finds and removes the excess.

What changed

  1. bench/bench_howl_detector.cpp: a per-stage breakdown, float and double, through the library's own code: howl_bank (resonators + envelopes: a detector whose tick never comes), howl_tick (howl_detail::decide on envelope snapshots), howl_tick_logs (its 66 std::log10 calls), howl_fit, howl_readouts, bench_guard/policy_m1 (the guard's own step, no detector) and bench_guard/m1, m2 (the whole guard, as CostPerBlock times it). The existing b64_f64 / b64_f32 cases are unchanged.
  2. include/mutap/howl_detector.h: the bank loop moves, unchanged, into howl_detail::run_bank with __restrict parameters; the member forwards to it.
  3. Docs: bench/README.md (the table below and the cause), the detector's COST paragraph, the CostPerBlock measured comment.

Temporary diagnostic commits (run_bank replicas, GCC's vectorizer report, the stages on Linux Clang and macOS runners, an identity dump, a CostPerBlock print) were reverted in 3a0610c; their numbers are below.

The breakdown (Benchmark smoke job, GCC 13.3, medians of 5 x 0.3 s, ns per block, float / double)

stage before (run 37921088083) after (run 37922694658)
howl_bank 10291 / 10182 1288 / 2434
howl_tick 730 / 945 725 / 945
howl_tick_logs 541 / 759 517 / 767
howl_fit 34.9 / 36.7 34.5 / 36.4
howl_readouts 241 / 293 248 / 294
bench_guard/policy_m1 30.9 / 48.1 29.9 / 48.8
howl_detector/b64 10885 / 11135 2006 / 3337
bench_guard/m1 10919 / 11250 2059 / 3407
bench_guard/m2 18204 / 18475 4113 / 6754

Other hosts (TEMP job, same stages; before and after are different runner instances, so compare within a column only): Linux Clang before: bank 1073 / 2061, detector 1615 / 2746, guard m1 1633 / 2780; after: 1320 / 2445, 2000 / 3363, 2061 / 3432. macOS arm64 CI before: bank 1271 / 1217, detector 1516 / 1833, guard m1 1379 / 1524; after: 1203 / 1368, 1734 / 1708, 1623 / 1694.

HowlGuardHost.CostPerBlock after (TEMP print, run 37922694658): Linux GCC float 1998 / 3986 ns (0.62 / 1.24 % of that runner's 321 us canceller), double 3371 / 6793 ns (1.03 / 2.07 %); Linux Clang float 2002 / 3990, double 3296 / 6561. The Intel Mac's recorded figure is 2180 / 4983 float.

The cause

All of the excess was the bank. GCC 13.3 left the band loop scalar (-fopt-info-vec: "couldn't vectorize loop", "not vectorized: no vectype for stmt" on the coefficient loads) and compiled the envelope's attack / release pick as comiss / jbe, a branch white noise mispredicts; Clang vectorized the loop. Replicas of the loop on the GCC runner (float / double): verbatim 10103 / 10383; __restrict on the local pointers 10603 / 10379 (no change); an indexed, branch-free pick 4364 / 4380 (scalar); the arrays as __restrict parameters 1164 / 2273 with the original ternary, the same with a bit-mask pick 1182 / 2311; the resonators alone 2545 / 2689, with restrict parameters 653 / 1283. GCC honours restrict on parameters, not on these locals.

Proof of identical output

The loop's arithmetic is unchanged, operation for operation. A dump of every readout (tripped, verdict, trigger, confidence, peak / line Hz, prominence, growth per pass and per s, rise, harmonic and subharmonic ratios, level, power, every band level) after every process_block call, over 4 configs x 6 signals that drive every trigger path (none / growth / ceiling / level = 65614 / 6824 / 19446 / 456 ticks) with irregular partitions, 3,440,844 values, compared main's header against this one with cmp: bit-identical on Linux GCC 13.3, Linux Clang and macOS arm64 (CI, run 37922694658) and AppleClang 17 x86_64 (local). Nothing the tests print moves: the 186 detector / guard / chain / two-mic / reverb tests (--gtest_filter='*Howl*:*howl*:*Afc*:*afc*:*TwoMic*:*two_mic*:*Reverb*:*reverb*') run on main's header and the full ctest on this branch (Intel Mac, AppleClang 17, Release, -DMUTAP_WERROR=ON; 420 / 420 passed) print the same 288 lines in the 74 tests that print anything, once CostPerBlock's two wall-clock lines are set aside. The detector and the guard are in neither the fingerprint harness nor the icount scenarios; the ratchet and the fingerprint legs pass unchanged.

On the Intel Mac (AppleClang already vectorized the loop) nothing moves beyond noise: four interleaved rounds at load ~7, howl_bank float main 1395-1467 ns vs this branch 1410-1476 ns, the detector 2133-2330 vs 2117-2359 ns.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A

tap and others added 7 commits October 9, 2026 06:00
The guard costs ~4x on the Linux GCC runner what it costs on the Intel
Mac (HowlGuardHost.CostPerBlock's comment, from PR #80's CI) while the
canceller runs ~2x faster there. To localize that in one CI run the
detector bench now times one block's stages separately, float and
double, through the library's own code:

  howl_bank       resonators + envelopes (a detector whose tick never
                  comes, so process_block runs run_bank alone)
  howl_tick       howl_detail::decide on envelope snapshots
  howl_tick_logs  decide()'s 66 std::log10 calls alone
  howl_fit        one band's least-squares fit
  howl_readouts   harmonic / subharmonic readouts + 32 band_level_db
  guard policy    howl_guard update() + apply(), the detector not run
  guard m1 / m2   the whole guard per block, as CostPerBlock times it

The existing b64_f64 / b64_f32 cases are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
Diagnosis only, reverted before the PR is ready: replicas of
howl_detector::run_bank (verbatim, __restrict, an indexed select,
resonators alone, envelopes alone, bands-outer), the howl stages at
5 x 0.3 s in bench-smoke (GCC), GCC's -fopt-info-vec report and
disassembly of run_bank, and the same stages on Linux Clang and macOS.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
…ity harness

Round 1 on the Linux GCC runner: run_bank is scalar there with a
comiss/jbe branch on the envelope's attack/release select (10.2 us per
block, float and double alike); __restrict on LOCAL pointers changed
nothing, an indexed select took it to 4.3 us. Round 2 tries __restrict
PARAMETERS, both arms computed before the select, and the coefficient by
bit mask; dumps GCC's vectorizer notes for the replicas; and adds a
bit-for-bit identity dump of every detector readout (main's header vs
the branch's) for the fix to come.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
…orizes it

The guard cost ~4x on the Linux GCC runner what it costs on the Intel
Mac. The per-stage bench put all of it in run_bank, the 32 resonators and
their envelopes: GCC 13.3 left the band loop scalar, with a comiss/jbe
branch on the envelope's y > env pick that noise mispredicts, while Clang
vectorized it (10103 vs 1060 ns per block, float, Linux GCC vs Linux
Clang; double 10383 vs 2039). The tick (decide()) costs the same order
on every host (730 ns float on GCC, 546 on Clang).

On GCC, __restrict on the loop's LOCAL pointers changed nothing (10603
ns); an indexed, branch-free pick took the scalar loop to 4364 ns; the
same loop with its arrays as __restrict PARAMETERS vectorized (1164 ns
float, 2273 double), with the original ternary. So the loop moves,
unchanged, into howl_detail::run_bank with restrict parameters; the
member forwards to it.

Each band's arithmetic is the old loop's, operation for operation. A
dump of every readout after every process_block call (4 configs, 6
signals that drive every trigger path, irregular partitions; 3440844
values) is bit-identical between main's header and this one on
AppleClang x86_64 (local); the PR runs the same dump on Linux GCC,
Linux Clang and macOS CI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
… macOS legs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
The run_bank replicas, the identity dump and the TEMP CI steps (GCC
vectorizer report, howl stages on Linux Clang and macOS, the identity
comparison, the CostPerBlock print) did their job; their measurements are
in bench/README.md and #87. ci.yml is main's again and
bench_howl_detector.cpp is the per-stage bench of the first commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
bench/README.md records the per-stage breakdown on the Linux GCC runner
before and after the bank's __restrict parameters (the bank 10291 ->
1288 ns float, 10182 -> 2434 double; the guard per block 10919 -> 2059 /
11250 -> 3407 at one mic), the cause, and the bit-identity check.
howl_detector.h's COST paragraph and HowlGuardHost.CostPerBlock's
measured comment gain the Linux GCC figures (CostPerBlock on job
113795845819: float 1998 / 3986 ns, double 3371 / 6793 ns at 1 / 2 mics).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
@tap
tap marked this pull request as ready for review October 9, 2026 14:28
@tap
tap merged commit 00590e5 into main Oct 9, 2026
36 checks passed
tap added a commit that referenced this pull request Oct 9, 2026
The run_bank replicas, the identity dump and the TEMP CI steps (GCC
vectorizer report, howl stages on Linux Clang and macOS, the identity
comparison, the CostPerBlock print) did their job; their measurements are
in bench/README.md and #87. ci.yml is main's again and
bench_howl_detector.cpp is the per-stage bench of the first commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
@tap
tap deleted the perf/guard-linux-cost branch October 9, 2026 14:31
tap added a commit that referenced this pull request Oct 9, 2026
bench_aec.cpp's -Wsign-conversion under -DMUTAP_WERROR=ON on AppleClang
(CI's bench-smoke builds without -Werror), and run_bank's bit-identity
not compared under the Cortex-M / Hexagon toolchains.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
tap added a commit that referenced this pull request Oct 9, 2026
bench_aec.cpp's -Wsign-conversion under -DMUTAP_WERROR=ON on AppleClang
(CI's bench-smoke builds without -Werror), and run_bank's bit-identity
not compared under the Cortex-M / Hexagon toolchains.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant