Skip to content

howl_guard: re-arm hold, soundcheck calibration, mandatory cap (phase 2 item 6b, PR C) - #80

Merged
tap merged 2 commits into
mainfrom
feat/howl-guard-policy
Oct 2, 2026
Merged

tap merged 2 commits into
mainfrom
feat/howl-guard-policy

Conversation

@tap

@tap tap commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

PR C of the safety layer (phase 2 item 6b), stacked on #78 (feat/howl-guard, 55a62fd). Spec: the safety design note §4c, the decisions of 2026-10-01: (1) the re-arm hold, (2) a soundcheck calibration step, (3) cap_db mandatory for opening without a declaration.

What lands

  • Re-arm hold (howl_guard.h). After a timer re-arm, the verdict trigger (LOST) arms only once the verdict has been continuously ok for release_hold_s (1.5 s, the same hold as a release). Before, the first ok block re-armed it. Walk releases and every other path are unchanged.
  • Soundcheck sampler. calibrate_begin() samples each mic's D and A′ once per block for calibrate_s (30 s). calibrate_end(apply) returns a std::span<const guard_calibration>, one entry per mic: blocks, median / p95 / max of D and A′, and the suggested d_db / a_db = median + cal_d_margin_db (4 dB) / cal_a_margin_db (3 dB). With apply each mic runs on its own thresholds. set_thresholds(mic, d, a) restores a stored soundcheck; clear_calibration() returns to the policy's; threshold_d_db() / threshold_a_db() read them. Applied thresholds re-arm LOST (an ok block under them first), survive reset() and are not touched by set_policy().
    • Sampler: a fixed histogram, not a reservoir or a running quantile. 0.1 dB bins; D over [−60, +20) dB, A′ over [−100, +20) dB; 2000 uint32 counts (8 KB) per mic, allocated by the constructor. O(1) per block, exact to the bin width for any quantile, identical on every platform, no random source. A reservoir needs a random source and an O(n) selection; P² has no error bound on the bimodal window a cold start produces.
  • Cap rule. PR B's ARMING already behaved this way without a cap: on timeout it stayed ARMING with unprotected raised and no OPEN. This PR documents the rule as mandatory and closes one hole: a mic that latched before any declaration (its floor was the cap) now re-arms when the cap is removed. Before, it kept the old cap as its floor.
  • Tests. State machine (float and double, both emulated selections): the re-arm hold (one block short of the hold does not arm LOST; the full hold does; a walk release still does), the sampler (exact bins, percentiles, the window cut-off, apply / reset / clear / out-of-range), the no-cap rule and the latched-undeclared re-arm, noexcept on the new entry points. Host:
    • three new gated rows: ColdStartWithoutACapStaysArmed, SoundcheckCalibrationOnStableMaterial and SoundcheckCalibrationSeesAWalk;
    • LouderCouplingRearms and ColdStartWithoutBackingTrack now gate the new numbers;
    • two new MUTAP_SLOW sweeps, CalibrationMargins (shadow guards over a 9 × 8 margin grid) and CalibrationApplied;
    • LouderCoupling is split out of WalksAndLouderCoupling, which is now Walks.
  • docs/howl-guard.md (new sections, every number from this PR's runs), HANDOFF item 12, the header comments, a README line.

Measured (macOS 15.7 x86_64, i9-8950HK, AppleClang 17, Release, double; every run from reset)

1. Re-arm hold: F → 2F before → after

"Before" is this PR's test harness on PR B's guard (55a62fd). Same rows, same host.

Runs Re-armed Re-arm after the duck, median [min, max] (s) LOST-ducks after the first re-arm (runs) All LOST-ducks after the change Howl blocks after the change / after the re-arm
Gated (cabin, mt5, studio, hall × seeds 1, 21), before 8 6 5.00 [5.00, 10.01] 16 (6) 28 0 / 0
Gated, after 8 6 5.00 [5.00, 10.01] 0 (0) 12 0 / 0
Sweep (six rooms × five seed sets), before 30 26 5.00 [5.00, 16.73] 47 (20) 87 0 / 0
Sweep, after 30 26 5.00 [5.00, 16.73] 1 (1) 41 0 / 0

Time ducked or releasing after the change, median per room, before → after: gated mt5 0.96 → 0.51, studio 0.94 → 0.26, hall 0.93 → 0.26, cabin 0.56 → 0.56; sweep mt9 0.95 → 0.77.

  • Gated as a rate with margin: at most 4 LOST-ducks after a re-arm over the 8 runs (0.5 a run, a quarter of the old 2.0), and at least 4 of 8 re-armed. Not a bare zero.
  • Cabin still cycles, by another route. Its verdict holds ok while ducked, so it releases on the walk path, not by a timer re-arm. The verdict is then lost again at full gain: 7 LOST-ducks in 2 gated runs, 18 in 5 sweep runs, unchanged. The sweep's one remaining LOST-duck after a re-arm is also cabin's. A LOST within probation counted as a strike would reach this case; it was not measured.
  • No regression: the walk rows (gated and sweep) and every cold-start row that PR B printed read the same to the last printed digit.
    • Walk at exact − 6: release − reconvergence median 1.59 s [1.38] gated, 1.60 s [1.38] sweep; 0 early.
    • Walk at the limit − 6: 0 of 12 / 0 of 30 ducked.
    • Declaration medians: 1.77–2.07 s gated, 1.76–2.11 s sweep.
    • Gap, two-mic, audible-cost and bus-stage gated tables: identical as well.

2. Soundcheck calibration: the margin sweep

HowlGuardSweep.CalibrationMargins: 210 runs at the canceller's limit − 6 with the backing track, a 30 s soundcheck from reset.

  • Stable runs (180): 30 s more of the same song; the audible-cost grid (six rooms × five seed sets × voiced+aux / speech+aux × 0 / 2 / 5 Hz).
  • Walk runs (30): the six walks × five seed sets, the walk at 40 s.
  • Method: the live guard keeps the factory thresholds. Beside it, 72 shadow guards (one per margin pair) take the same residual and statistics and apply their soundcheck at 30 s. A shadow's first duck is exact until its gain leaves the live guard's; 0 inexact here.

Stable runs with a duck after the soundcheck, of 180:

D margin \ A′ margin + 0 dB + 1 to + 10 dB, or A′ off
+ 0 174 140
+ 0.5 135 36
+ 1 131 11
+ 1.5 131 7
+ 2 131 3
+ 3 131 0
+ 4 (default) 131 0
+ 6 131 0
D off 131 0

Walks seen, of 30 (a duck within 5 s of the change); in brackets, runs that ducked before the walk:

D margin \ A′ margin + 0 + 1 + 2 + 3 (default) + 4 to + 10, or off
+ 0 0 [30] 2 [28] 2 [28] 2 [28] 2 [28]
+ 0.5 0 [30] 13 [17] 13 [17] 13 [17] 13 [17]
+ 1 0 [30] 21 [9] 21 [9] 21 [9] 21 [9]
+ 1.5 0 [30] 24 [6] 24 [6] 24 [6] 24 [6]
+ 2 0 [30] 28 [2] 28 [2] 28 [2] 28 [2]
+ 3 0 [30] 30 [0] 30 [0] 30 [0] 30 [0]
+ 4 (default) 0 [30] 30 [0] 30 [0] 30 [0] 30 [0]
+ 6 0 [30] 28 [0] 18 [0] 16 [0] 16 [0]
D off 0 [30] 23 [0] 1 [0] 1 [0] 1 [0]

2079.35 s on 4 threads.

The margins.

  • D: the smallest margin with 0 stable ducks is + 3 dB (+ 2 dB: 3 runs ducked). The default is + 4 dB: one step of headroom, 2 dB over the last margin that ducked.
  • A′: + 0 always ducks, since half the blocks sit above the median. + 1 dB already ducked none. The default is + 3 dB, a factor of 2.
  • Walks at the defaults: 30 of 30 seen, 0.31 s after the change. The factory thresholds saw 0 of 30.

The soundcheck medians over the 210 runs:

  • D: −12.15 to −6.25 dB, 5 to 11 dB under the factory −1.235.
  • A′: −30.55 to −25.85 dB.

What the soundcheck prints (gated row; two of its eight lines):

  cabin voiced+aux seed 1            soundcheck 22500 blocks: D median   -8.15 dB (p95   -5.55, max   +0.15); A' median  -27.75 dB (p95  -27.15, max   -0.05) -> d_db   -4.15, a_db  -24.75 (applied)
  mt5 speech+aux seed 1              soundcheck 22500 blocks: D median   -8.35 dB (p95   -7.35, max   +0.45); A' median  -29.15 dB (p95  -25.45, max   -0.05) -> d_db   -4.35, a_db  -26.15 (applied)

Applied (the live guard on its soundcheck thresholds at the default margins):

Stable runs with a duck Walks ducked (by LOST) Duck after the walk, median (s) Release − misalignment-oracle reconvergence, median [min] (s) Early Howl blocks
Gated (cabin, mt5; seeds 1, 21) 0 of 8 12 of 12 (11) 0.31 1.74 [1.63] 0 0
Sweep (CalibrationApplied, 1227 s) 0 of 180 30 of 30 (29) 0.31 1.74 [1.63] 0 0

Applied d_db across the sweep: −8.15 to −2.25 dB; a_db: −27.55 to −22.85 dB. The duck not by LOST was a detector TRIP (mt5 → mt105).

Gates, with margin:

  • stable: at most 2 of 8 runs with a duck (measured 0);
  • walks: at least 8 of 12 ducked (measured 12), the release median over 0.5 s (1.74), at most 2 early (0);
  • 0 howl blocks;
  • the shadow at the live margins tracks the live guard exactly (structural).

3. Cap rule: track-off rows

20 s from reset at the limit − 6 (gated: cabin and mt5, seeds 1, 21, 41).

Row Cap Runs Declared OPEN_CAPPED Time to OPEN_CAPPED, median [max] (s) unprotected from (s) Time in ARMING, median (s) Ducks after OPEN_CAPPED (runs) Howl blocks
held note dry limit − 6 6 0 6 10.00 [10.00] 10.00 10.00 2 (2) 0
music dry limit − 6 6 0 6 10.00 [10.00] 10.00 10.00 0 0
speech dry limit − 6 6 0 6 10.00 [10.00] 10.00 10.00 0 0
held note none 6 0 0 — 10.00 20.00 (all) 0 0
music none 6 0 0 — 10.00 20.00 (all) 0 0
speech none 6 0 0 — 10.00 20.00 (all) 0 0

"10.00" is block 7499, 9.9987 s.

Sweep (six rooms × five seed sets, 30 runs a row):

  • Under the cap: every track-off run reached OPEN_CAPPED at 10.00 s, with 0 howl blocks. The held note ducked after OPEN_CAPPED in 11 of 30 runs at S1 (9 of 20 at exact + 3, 4 of 20 at exact + 6), music in 1 of 20 at exact + 3, speech never. These are detector TRIPs (LOST is disarmed before a declaration), not classified against the burst oracle.
  • Without a cap (held note, music, speech): 90 of 90 runs stayed in ARMING for all 20 s, 0 howl blocks.
  • The guarded cold-start sweep now totals 690 runs with 0 howl blocks.

Defaults

None of PR B's defaults moved. New fields:

  • calibrate_s = 30: the protocol's window.
  • cal_d_margin_db = 4 and cal_a_margin_db = 3: measured above.

What did not separate, or could not be measured

  • F → 2F through the walk path (cabin) still cycles; the LOST-in-probation-as-a-strike alternative was not measured.
  • The margins are one loop's and one song's. Not measured:
    • a soundcheck on one song with the show on another;
    • a soundcheck at a different gain from the show;
    • a soundcheck without the backing track: the verdict never declares there, and the sampler reads ARMING's statistics.
  • The 30 s window includes about 2 s of the cold start before the declaration.
  • Under the cap, the held note's post-OPEN_CAPPED ducks were not classified against the burst oracle.
  • arm64 / Linux numbers: only CI pass / fail. The new gates leave room for platform drift, but no second host has printed these tables.
  • HowlGuardSweep.OperatingPoints was not re-run: the guard is not in it.

Verification

  • cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DMUTAP_WERROR=ON && cmake --build build: clean.
  • ctest --test-dir build --output-on-failure: 374 of 374 passed (21 skipped: the 20 MUTAP_SLOW sweeps and PortableRandom.MatchesLibstdcxxBitForBit), 3143.71 s. The 16 gated guard rows took 53.09 to 88.65 s each (a first full run failed ColdStartWithoutBackingTrack on my own bound, ≥ 10.0 s against block 7499 = 9.9987 s; fixed to 10 − 1.5 blocks, then the full run above).
  • Gated rows alone (--gtest_filter='HowlGuardHost.*'): 777.37 s.
  • Sweeps, 4 threads (MUTAP_SLOW=1 MUTAP_SLOW_THREADS=4): CalibrationMargins 2079.35 s, CalibrationApplied 1227 s, ColdStart 1890 s, Walks 245 s, LouderCoupling 88 s, TwoMicsAndAudibleCost 567 s (its two-mic and audible-cost tables match PR B's).
  • afc_chain.h is unchanged (bit-identical without a guard by construction). mutap_fingerprint passed. No icount or fingerprint workload includes afc_chain.h or howl_guard.h.
  • pre-commit run --files <changed>: passed. scripts/tidy.sh on test_howl_guard.cpp, test_howl_guard_host.cpp and test_afc_chain.cpp: clean. A direct clang-tidy-18 -p build-tidy count on each: 0 warnings.
  • GCC safety: the new trace loop in the host test walks clamped_range iterators, as 55a62fd does. The calibrate_end / quantile loops walk iterators over each mic's histogram slice.

Follow-up commit 8c2e368 (CI)

  • HowlGuardHost.CostPerBlock: Linux GCC CI (job 110465810942) read the guard at 4.61 / 6.86 % (float, 1 / 2 mics) and 4.59 / 7.15 % (double) of a canceller, against 0.57 / 1.09 / 0.84 / 1.73 % on the Intel Mac, failing the 4.5 / 7 % bounds by 0.001. The ratios are now printed and RecordProperty'd; the gate is < 25 % (a gross regression only).
  • HowlGuardHost.ColdStartWithoutBackingTrack: macOS arm64 CI ducked after OPEN_CAPPED in 5 of 6 held-note runs (Intel 2 of 6) and failed duck_runs <= 4. The ducks are now classified against the burst oracle: on Intel both are within 0.5 s of a +20 dB loop-born burst (0 off a burst). Gated: 0 howl blocks, OPEN_CAPPED at the timeout (>= 10 s − 1.5 blocks, median < 11 s), runs with a duck off any burst <= 6 of 18 (Intel 0). The duck count is reported, not gated.
  • Re-run locally after the change: both rows pass (52.17 s, 17.69 s); pre-commit and scripts/tidy.sh clean, a direct clang-tidy-18 count 0.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A

@tap tap mentioned this pull request Oct 1, 2026
@tap
tap force-pushed the feat/howl-guard-policy branch from 5555843 to 5eaa678 Compare October 1, 2026 16:12
@tap
tap changed the base branch from feat/howl-guard to feat/howl-guard-land October 1, 2026 16:12
tap and others added 2 commits October 1, 2026 17:16
… 2 item 6b, PR C)

- After a timer re-arm, LOST arms only once the verdict has been ok for
  release_hold_s (not on the first ok block). F -> 2F, LOST-ducks after
  the re-arm: gated 16 in 6 of 8 runs -> 0; sweep 47 in 20 of 30 -> 1.
  Walk and cold-start rows unchanged to the last printed digit.
- Soundcheck sampler: calibrate_begin() / calibrate_end(apply) over a
  fixed 0.1 dB histogram per mic (8 KB, allocated by the constructor);
  per-mic thresholds = 30 s median + cal_d_margin_db (4) /
  cal_a_margin_db (3), measured with shadow guards over 180 stable and
  30 walk runs (D + 3 the smallest with 0 stable ducks). Applied: 0 of
  180 stable runs ducked, 30 of 30 walks at the limit - 6 seen and
  released after the misalignment oracle (median 1.74 s, min 1.63).
- cap_db is mandatory for opening without a declaration; removing the
  cap now also re-arms a mic latched before any declaration. Without a
  cap 90 of 90 track-off sweep runs stay in ARMING for all 20 s.
- New gated rows ColdStartWithoutACapStaysArmed,
  SoundcheckCalibrationOnStableMaterial, SoundcheckCalibrationSeesAWalk;
  sweeps CalibrationMargins, CalibrationApplied; LouderCoupling split out
  of Walks. docs/howl-guard.md, HANDOFF item 12, README.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
…gates ducks off a burst

- HowlGuardHost.CostPerBlock: the wall-clock ratio moves by 4-8x between
  hosts (Intel 0.57 / 1.09 % float, Linux GCC CI 4.61 / 6.86 % float and
  4.59 / 7.15 % double, failing the 4.5 / 7 % bounds by 0.001). Print and
  RecordProperty the ratios; gate < 25 % only.
- HowlGuardHost.ColdStartWithoutBackingTrack: the held note's ducks after
  OPEN_CAPPED are classified against the burst oracle (Intel: 2 of 6 runs,
  both ducks on a loop-born burst, 0 off; macOS arm64 CI: 5 of 6 runs,
  which failed the <= 4 bound). The duck count is reported; gated: howl
  blocks, the OPEN_CAPPED time, runs with a duck off any burst <= 6 of 18.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
@tap
tap force-pushed the feat/howl-guard-policy branch from 8c2e368 to 1798fce Compare October 1, 2026 22:16
@tap
tap changed the base branch from feat/howl-guard-land to main October 1, 2026 22:16
tap added a commit that referenced this pull request Oct 2, 2026
…nstead of apt gcc-arm-none-eabi

Ubuntu's gcc-arm-none-eabi pulls the 463 MB libstdc++-arm-none-eabi-newlib
package from the runner's Azure mirror. That fetch stalled past the
15-minute step limit three times on 2026-10-01 (#81 twice,
#80 once) after eating two 45-minute legs on #67. Both
Cortex legs now restore Arm's arm-gnu-toolchain-13.2.rel1-x86_64-arm-none-eabi
from the Actions cache, or download it (179 MB from Arm's CDN, curl with
retries, verified against Arm's published sha256) on a miss, and put its
bin on PATH; only qemu-system-arm still comes from apt. It is the same
upstream release Ubuntu noble packages (15:13.2.rel1-2), so the
fingerprint and icount gates on this PR's own run are the check that the
numerics did not move.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
@tap
tap merged commit 449cf6c into main Oct 2, 2026
36 checks passed
tap added a commit that referenced this pull request Oct 9, 2026
The guard costs ~4x on the Linux GCC runner what it costs on the Intel
Mac (HowlGuardHost.CostPerBlock's comment, from PR #80's CI) while the
canceller runs ~2x faster there. To localize that in one CI run the
detector bench now times one block's stages separately, float and
double, through the library's own code:

  howl_bank       resonators + envelopes (a detector whose tick never
                  comes, so process_block runs run_bank alone)
  howl_tick       howl_detail::decide on envelope snapshots
  howl_tick_logs  decide()'s 66 std::log10 calls alone
  howl_fit        one band's least-squares fit
  howl_readouts   harmonic / subharmonic readouts + 32 band_level_db
  guard policy    howl_guard update() + apply(), the detector not run
  guard m1 / m2   the whole guard per block, as CostPerBlock times it

The existing b64_f64 / b64_f32 cases are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDhgJqxUsdmWPKh6pKrV5A
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant