Monorepo migration: the tap::sr family (async + bridge) - #49
Merged
Merged
Conversation
PLAN.md is the authoritative plan: charter (44.1<->48 only, synchronous, speed-first), DspTap-substrate architecture, API surface, profiles, three-leg test strategy, milestones M0-M7, and the two M0 extraction PR outlines (DspTap gains the shared FIR substrate; SampleRateTap adopts it via submodule). HANDOFF.md preserves the original SampleRateTap synchronous-engine design brief as provenance, with a status preamble recording the seven decisions revised during planning (separate repo via DspTap, pinned-eps cross-validation, construction-time design, no unified factory API, speed-first scope, inherited channel kernels, family profile vocabulary). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Milestone M1 of PLAN.md — everything except the converter: - CMake: header-only tap::ratio INTERFACE target (include/tap/ratio/, C++20) over the DspTap substrate, pinned as submodules/dsptap at the FIR-substrate merge and linked as tap::dsp; TAP_RATIO_WERROR warnings target; GoogleTest test harness. - include/tap/ratio/ratio.h: umbrella header carrying version constants, the charter docstring, and the identity constants (L = 160 up, L = 147 down). - tests/test_skeleton.cpp: pins the identity constants (coprime phase counts) and proves the substrate end to end at this library's own geometry — design a prototype at L = 147, quantize a branch row-sum- exactly to Q15, dot it against DC through the shared kernel. - TapHouse adoption: canonical .clang-format / .clang-tidy / .pre-commit-config.yaml / STYLE.md / scripts/tidy.sh, plus the SessionStart hook for Claude Code web sessions (submodule init + pre-commit install). - CI: host build+test matrix (Linux/macOS/Windows, -Werror) and an ASan+UBSan job, all with recursive submodule checkout; the Tap House Style workflow (taphouse drift check @v5 + clang-tidy over project TUs with the submodule excluded). - README carrying the charter (identity boundaries, family position, status pointer to PLAN.md), CLAUDE.md, MIT LICENSE. Verified locally: GCC and clang -Werror builds clean, 2/2 tests green under both, scripts/tidy.sh clean, pre-commit clang-format clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The numbers-first half of the converter (PLAN.md milestone M2):
- notebooks/design_spike.ipynb (executed, committed): re-derives the four
prototype designs in numpy from the same published math as the shipping
C++, proves each pinned taps-per-phase count minimal (two fewer fails
the >=1 dB-margin criterion), and cross-checks the economy 48->44.1
conversion end-to-end through scipy's upfirdn polyphase engine: the
audible band measures -100 dBFS, all alias products confined above
20 kHz. Pinned geometries: down 147x78 / 147x184, up 160x44 / 160x96
(economy/transparent), the ~2x direction asymmetry as predicted.
- include/tap/ratio/design.h: compile-time direction enum + ratio_traits
(L/M/rates/stopband edges), profile presets carrying the pinned taps
(economy 70 dB/19 kHz default, transparent 120 dB/20 kHz), prototype
design via the shared tap::dsp kaiser math plus per-branch DC
normalization -- every polyphase branch sums to exactly 1.0, killing
the fs_out/L-harmonic spurs a raw design's branch-sum spread would
inject from DC/LF energy, and letting fixed-point row-sum quantization
land on format unity exactly.
- include/tap/ratio/schedule.h: the constexpr (phase, advance) superblock
table -- phase(n) = nM mod L, advances {0,1} up / {1,2} down summing to
exactly M -- plus frames_needed(pos, out), the deterministic input-need
arithmetic the pull-style composition depends on.
- include/tap/ratio/phase_table.h: basic_phase_table<S, D> -- exactly L
phase-major tap-reversed rows (no interpolation, no wrap row, no
power-of-two rounding), quantized per row with the shared utility.
- Contract tests (22 new; suite 24): DFT spec sweeps for all four designs
with measured prints, minimality-adjacent asymmetry pin, exhaustive
schedule verification from every superblock position, and per-phase
row-sum + DC guarantees for every phase of all four tables across
float/Q15/Q31.
- PLAN.md: (spike) placeholders in section 4 and section 8 replaced with the
pinned numbers (alias floors, storage, latency figures).
Verified: GCC and clang -Werror clean, 24/24 green under both, tidy and
clang-format clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The measured worst-stopband figures quoted in design.h's profile table predated the per-branch DC normalization; the shipping designs measure -72.1/-72.8 dB (economy) and -121.7 dB (transparent), as test_design.cpp prints. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The M2 commit staged files before the format hook rewrote them, so the committed copies drifted from the formatter by alignment-only diffs; this restores the canonical formatting the CI gate checks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The float engine (PLAN.md milestone M3), one table-row dot per output over the compile-time schedule — no interpolation, no fractional-delay state, no servo: - include/tap/ratio/converter.h: basic_converter<S, D> with both call shapes — process() (push-transform: consume all input, emit what becomes ready) and pull() (produce exactly N, draw input through a noexcept PopFn; dry sources short-return and resume losslessly) — plus exact frames_needed()/outputs_for() accounting, flush() (drains the taps()-frame tail), reset(), latency_input_frames(), runtime channel count over planar per-channel delay lines (coefficient row shared per frame: inter-channel phase coherence exact). Alignment contract: zero-primed and causal, y[n] = sum_k x[floor(nM/L)-k] * h[phase(n)+kL] — scipy upfirdn's streaming prefix exactly. Float aliases converter_to_48k / converter_to_44k1 (Q15/Q31 in M4). - tools/reference/make_reference_vectors.py + committed tests/reference/reference_vectors.h: the independent golden leg — scipy.signal.upfirdn in float64 over the same designs, 1000 frames of deterministic noise, all four direction x profile cases. - tests/test_converter.cpp (18 new; suite 42): impulse response reproduces the coefficient table bit-for-bit through the whole engine (every produced sample IS one stored coefficient); scipy vectors matched sample-for-sample from n=0 within the float32 coefficient floor; pull==process bit-exact under a dribbling source; accounting verified exhaustively from every superblock position against walking consumption (and outputs_for as the exact inverse of the need table); alias acceptance on real converted audio; stereo channels bit- identical to mono runs (crosstalk exactly zero); flush/reset lifecycle. Measurement finding folded into PLAN.md section 8: economy's in-band floor is upsampling IMAGE leakage (e.g. a 23 kHz tone's image at 25 kHz folds to 19.1 kHz at -85 dBFS), bounded by the stopband — the arithmetic-confinement claim covers decimation aliases only, and the earlier "nothing measurable below 20 kHz" phrasing holds only at the transparent tier. Deepening these images is the k*fs image-zeros lever (M7). Verified: GCC and clang -Werror clean, 42/42 green under both, tidy and clang-format clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
A scripted edit inserted target_include_directories twice; keep one, with a comment saying why it exists (the committed reference vectors resolve relative to tests/). Also re-runs CI: the previous round's only red was a clang-tidy job on the push-event twin run that hung and hit the 30-minute timeout — its pull_request twin passed the identical commit in 2m37s, so the failure was a stalled runner, not the code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Q15/Q31 aliases for both directions (the engine was format-generic already; this milestone proves the numeric contracts through the whole streaming path) plus 11 tests (suite 53): - Q31 tracks the float golden model within 5e-8 per sample (-147 dB) on the reference noise — transferring the committed scipy leg to fixed point — and measures 146 dB at 997 Hz through transparent, EXCEEDING float, whose float32 I/O is its own bound. - Q15 is format-limited: 76.1 dB economy / 72.8 dB transparent at half scale. Transparent is measurably WORSE at Q15 — Q1.14 coefficient noise stacks with tap count (184 vs 78) while the deeper filter buys nothing 16 bits can express — so economy is the recommended Q15 pairing (cheaper AND quieter), pinned by test and stated in PLAN section 8 and the README. - Full-scale (99%) drive saturates without wrapping (second-difference bound); full-scale DC emerges within one LSB at every phase of the superblock end to end; pull() stays bit-identical to process() for integer samples under a dribbling source. Thresholds sit ~4 dB under measured, printed [ measured ] per the family convention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The handoff doc's central idea, landed where it actually works (PLAN section 6.3): SampleRateTap's fractional_resampler — the async engine's mu-interpolated datapath — driven at PINNED eps = L/M - 1 with no servo, over the identical input and the identical plain-Kaiser prototype, must agree with this library's exact rational machine on every phase. Measured: worst disagreement 3.5e-6 down (-109 dB) / 1.2e-5 up (-99 dB), all 147 and all 160 phases covered, unchanged between async L=512 and L=1024 — the async table's mu-interpolation residual sits below the one deliberate filter difference between the machines (RatioTap's per-branch DC normalization, a ~5e-6 perturbation), so the exact and interpolated engines agree to the last systematic difference we chose to introduce. Two alignment subtleties the test documents and handles: - The resampler consumes its advance BEFORE each dot, so its output n is the converter's output n+1; and priming it with T-1 zeros + signal reproduces the converter's zero-primed first window exactly. - A length-LT linear-phase prototype delays (LT-1)/(2L) = T/2 - 1/(2L) input samples, which depends on L: the exact (L=147/160) and async (512/1024) tables center the same continuous filter 1/(2L)-1/(2L') apart (~0.0024 samples, ~2e-3 signal error uncompensated). One eps-folded phase advance on the resampler's first step cancels it. Mechanism: SampleRateTap arrives as a TEST-ONLY submodule consumed headers-only (its own CMake would add_subdirectory a second dsptap and collide; both repos pin the identical dsptap tree, so our tap::dsp serves its includes). Never linked into the shipped target. Suite 57. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The final v0.1 milestone (PLAN.md M6): the composition example and the family-convention verification layer. - examples/bluetooth_bridge.cpp: the documented answer to "44.1<->48 across independent clocks" — RatioTap converts the NUMBER (exact rational, clock-agnostic), SampleRateTap absorbs the CLOCK (near-unity servo). Deterministic +200 ppm two-clock simulation of the receive path; measured: servo locks at +200.1 ppm, tone recovered at exactly 997.000 Hz at amplitude 0.5000, SNR 73 dB (economy tier), total latency 2.00 ms (0.50 RatioTap + 1.50 ASRC). Exits nonzero if any of that regresses, so it doubles as an integration check. - tools/capi/: minimal C ABI over the float converters (create/process/ flush/frames_needed/outputs_for/latency), standalone-buildable for the ctypes bridge; notebooks/ratiotap_py.py builds build_capi/ on first import (family pattern). - notebooks/ratio_demo.ipynb (executed, committed): drives the SHIPPING C++ through the ABI — the exact-accounting demo (including the honest pre-advance subtlety: a fresh superblock costs 159, steady-state exactly 160), the economy spectral contract on hostile program material (worst audible-band product -95.9 dBFS, bound asserted), and the measured passband (+/-0.002 dB to the 19 kHz edge, roll-off where designed). - CMake: TAP_RATIO_BUILD_EXAMPLES (top-level ON) / TAP_RATIO_BUILD_CAPI (OFF) options; srt_headers dev-only target hoisted to the top level, shared by the cross-validation test and the bridge example. - README: v0.1 status + quick start; CLAUDE.md current-state refresh; PLAN marks M0-M6 complete, v0.1 shipped. Verified: GCC and clang -Werror clean, 57/57 green under both, the bridge example passes under both, tidy sweep and clang-format clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The ctypes bridge's bytecode cache slipped into the M6 commit; remove it and ignore the pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
MSVC /W4 with warnings-as-errors fires benign C4324 (structure padded due to alignment specifier) on the ASRC ring's deliberate cache-line alignas when the bluetooth_bridge example instantiates it — failing the Windows CI leg on both duplicate runs. Dependency headers are not this repo's warning surface: SYSTEM includes (/external:W0 on MSVC, -isystem elsewhere) exempt them while our own code keeps the full gate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The measurement harness the M7 optimization campaign is gated on, landing before the first lever per PLAN.md section 7. Ported from SampleRateTap's machinery (the family's embedded story lives there first), renamed for this repo: - bench/icount/: eight fixed workloads, one binary per scenario (direction x float/Q15/Q31 on economy, plus both transparent float legs; 2 s of stereo virtual audio each, precomputed input so libm stays out of the measured loop, checksum sink defeats DCE). - tools/qemu_insn_plugin/ + scripts/icount.py: deterministic QEMU instruction counting, gated two-sided (+/-3%) against bench/baselines.json — a regression fails, and an improvement beyond tolerance fails too so the gate stays tight. - bench/baselines.json: all 24 baselines (3 targets x 8 scenarios) measured with the CI-pinned QEMU 8.2.2; re-measurement is bit-identical, and the doctored-baseline drill confirms both failure modes fire. - cmake/ toolchains + platform/ startup/linker scripts for Cortex-M55 (QEMU mps3-an547), Cortex-M33 (mps2-an505, Pico-2 class) and Hexagon (static musl, qemu-user); tests/bare_metal_main.cpp runs the emulation-sized suite one-shot (no argv on bare metal). - CI: three cross test jobs + the icount-ratchet job, toolchain/QEMU sources SHA256-pinned; hexagon ctest runs -j 4 (independent qemu-user processes; the per-test soft-double table construction dominates). Verified locally: host 57/57; M55 and M33 bare-metal suites green under QEMU; Hexagon 55/55 (EXPECT_THROW excluded: static musl terminates instead of unwinding, same known debt as SampleRateTap's leg). The recorded baselines already rank the levers per target: M33 Q15 is ~6x cheaper than the float path (245 M vs 1.53 B insns, up/economy), Hexagon fixed point ~5x cheaper than float, while on M55 float currently beats Q15/Q31 (97 M vs 119 M) — the fixed-point dot kernels are a named M7 target on that core. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Lever 1 of the optimization campaign (PLAN section 7), measured by the M7a ratchet and re-recorded in bench/baselines.json: - process() dispatches to a superblock walk: outputs_for() settles the trip count before the first sample moves (no exhaustion checks in the loop), the schedule cursor / history end / pending gap live in locals the compiler keeps in registers, and input/output advance by pointer bumps instead of per-frame index multiplies. - Mono and stereo are stamped out as specializations (CH = 1 / 2) with the per-channel loops dissolved and the planar history pointers hoisted (compaction memmoves in place, so they stay valid); any other channel count takes the generic instantiation. - Same append/emit order, same tap::dsp dot kernels: outputs are bit-exact. 57/57 host (gcc + clang -Werror), M55 and M33 bare-metal suites green (and faster: 3.5 -> 2.4 s, 38.5 -> 27.3 s wall), Hexagon 55/55, bluetooth_bridge output identical, clang-tidy clean. Measured (2 s stereo virtual audio per workload): - M55: -31% to -60%. up_q15 118.8M -> 58.5M; Q15 now beats float on M55 (58.5M vs 62.2M), resolving half of M7a's named anomaly — Helium just needed a register-resident loop around the dots. - M33: fixed point -12% to -25% (up_q15 245M -> 216M, q31 -24%); float only -2.8% (soft-double MACs dominate by design — the golden model's double accumulation). - Hexagon: flat to -3.8% — hexagon-clang already generated tight code around the old loop, confirming the overhead was an Arm codegen story. pull() deliberately keeps the generic per-frame loop: its granularity is the pop callback, a different lever. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
…ngths Lever 2 of the optimization campaign (PLAN section 7), the codegen half of "baked tables": the two canonical profiles become constexpr, and process() dispatches once per call on the four pinned taps-per-phase counts (44/78 economy, 96/184 transparent), handing the superblock walk a compile-time trip count that the inlined tap::dsp dot kernels unroll and vectorize against. Custom-taps profiles take the runtime-length instantiation, pinned by a new impulse test of its own (Converter.ImpulseReproducesTableCustomTaps). Measured (baselines re-recorded, all three targets): - Hexagon: fixed point -7% to -12% on all four scenarios (up_q15 49.8M -> 44.0M, up_q31 -11%) — hexagon-clang pipelines the exact-count scalar loops. Float ~-1%. - M55: up_q15 -15% (58.5M -> 49.6M; the 44-tap dot fully unrolls under Helium). The 78/96/184-tap scenarios stay loop-shaped: flat within tolerance. - M33: Q15 -2.6% / -3.4%; float flat (soft-double MAC bound). Coefficient BAKING (committed rodata tables) stays un-pulled: construction is <0.3% of every workload, so its value is MCU boot time and RAM, not instruction counts — deferred until a consumer needs it. Verification: 58/58 host (gcc + clang -Werror), M55 and M33 bare-metal suites green, Hexagon 56/56 under emulation, bluetooth_bridge output identical, clang-tidy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Lever 3 of the optimization campaign (PLAN section 7), the RatioTap half
of the two-PR lever (DspTap grew dot_row_reversed first; this bumps the
pin to c2bfbdb and consumes it).
The linear-phase prototype gives b_p[t] = b_{L-1-p}[T-1-t]: branch p is
branch L-1-p tap-reversed. The table now stores only ceil(L/2) rows (74
of 147 down, 80 of 160 up) and a mirrored phase dots its partner's
stored row backward via tap::dsp::dot_row_reversed — same products, same
accumulation order, bit-identical outputs. Quantization is canonical
over the stored half, so the mirror is exact by construction and the
fixed-point exact-unity row sums carry over (reversal preserves the
multiset). Storage, pinned by test: economy f32 44.8 -> 22.5 KiB (down),
27.5 -> 13.8 KiB (up); transparent 105.7 -> 53.2 / 60.0 -> 30.0 KiB;
Q15 halves those again.
Compute cost: ZERO baseline changes on all three targets — with one
codegen subtlety the ratchet caught and the diagnostics isolated:
inlining both dot arms in the walk broke Arm's unrolled forward codegen
(M55 up_q15 +3.3%, reversed kernel itself measured free), while
outlining the mirrored arm broke Hexagon's (down_q31 +3.3% called, +2.7%
inlined). The mirrored-arm out-lining is therefore gated per target
(TAP_RATIO_MIRRORED_DOT_ATTR), the same measured-per-target pattern as
the tap::dsp kernel gates. Worst residual rides inside the two-sided
gate (Hexagon down_q31 +2.7%); Arm came out slightly ahead (M33 Q31
-2.5%, M55 up_q15 -1.1%).
Verification: 58/58 host (gcc + clang -Werror); M55 and M33 bare-metal
suites green — the M33 leg is the on-target bit-exactness proof of the
SMLALDX swapped-lane pairing (impulse-reproduces-table sweeps every
mirrored phase with the intrinsic active); Hexagon 56/56; every-phase
table battery extended with the exhaustive mirror-identity sweep;
bluetooth_bridge identical; clang-tidy clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
The DspTap merge rewrote the branch commit (c2bfbdb -> dfbe18d, identical tree); the pin must be reachable from dsptap main or recursive submodule checkouts break once the branch ref is pruned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Declares the output-preserving phase of the optimization campaign complete and bumps the version (CMake project + TAP_RATIO_VERSION_*). Four levers landed, each measured by the instruction-count ratchet, outputs bit-identical throughout: the M7a harness itself, the M7b superblock walk, the M7c committed trip counts, and the M7d symmetry-halved tables. Cumulative vs the M7a baselines: M55 Q15 -59%/-60% and float -35%/-37%; M33 Q31 -26%/-27%, Q15 -15%/-16%; Hexagon Q15 -13%/-10%, Q31 -15%/-7%; table storage halved (economy Q15 up: 6.9 KiB). README and PLAN now state the deferral rationale for the remaining levers (multistage, minimum-phase, IIR pre-filter, FFT offline): each changes the output contract or serves a currently-unpressured need, so they wait for a consumer to pull them — a latency need pulls minimum-phase, a storage need pulls multistage, an MCU float consumer pulls the accumulation-contract discussion in DspTap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ldeq57sBySx2nFTQsGcQB6
Synced from the canonical TapHouse copy: the family's shared review prompts -- what changed and why, a Verification section asking what was actually built and run versus what CI will gate (plus the measured-not-remembered rule for performance claims), and a delete-what-does-not-apply list of the recurring cross-repo concerns: contract changes, submodule pin flow, notebook re-execution, package docs/help and universal binaries, and documented per-repo exceptions. Created only-if-missing and deliberately NOT drift-guarded, following the .claude/settings.json precedent: it is prose for humans rather than a machine-enforced config, so this repo is free to tailor it -- trim the bullets that do not apply here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JhhQ93r2E1QTnCx46YfX8j
Refreshes submodules/dsptap from dfbe18d to 28a34a1, the current DspTap main. This repo was the least stale consumer in the family, one commit behind, and that commit is documentation only -- the shared pull-request template synced from taphouse. No substrate change reaches this repo. Pinned to a commit on DspTap's main rather than to any in-flight branch, per the release flow: consumer pins must reference a tree that stays reachable after branch cleanup. Verified against the substrate discipline this repo depends on -- shared code lands in DspTap first and here second, so a pin bump is the only correct way to pick it up: configure and build clean with zero errors, ctest 58/58 passing in Release, which includes the exhaustive phase sweeps and the pinned-eps cross-validation against SampleRateTap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JhhQ93r2E1QTnCx46YfX8j
This repo ships a C ABI under tools/capi that the executed notebook drives via ctypes, but TAP_RATIO_BUILD_CAPI defaults to OFF and CI never turned it on -- so nothing in the pipeline compiled it. A change to a kernel signature could break the ABI and the notebook with it, and CI would stay green. AmbiTap and DspTap already build theirs in CI; this brings RatioTap in line. Enabled on the Linux and macOS legs only. tools/capi carries no __declspec(dllexport) (unlike DspTap's), so an MSVC build would link a DLL that exports nothing -- it would pass without gating anything. That is recorded in the matrix comment and in taphouse's known-divergences list, to be flipped on in the same change that gives the C ABI an export decoration. Verified locally: configure and build succeed with TAP_RATIO_WERROR=ON, the shared library links, it exports its 11 ratio_* entry points, and the full suite passes 58/58. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JhhQ93r2E1QTnCx46YfX8j
…I nits - ratio_create no longer leaks the wrapper when the converter constructor throws (unique_ptr owns it until success), and the C ABI header states the NULL contract explicitly (ratio_destroy(NULL) is the free() no-op). - pull() now enforces its stated noexcept-PopFn requirement with a static_assert instead of terminating at runtime on a throwing callback. - ratio.h's status comment catches up from M3 to the shipped v0.2 state. - test_schedule.cpp includes <vector> for itself instead of leaning on gtest's transitive includes. - The Hexagon toolchain download+verify moves into a shared script so both writers of the digest-keyed cache run identical checks, and the style workflow's clang-tidy loop reads its file list line-wise. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
The economy18 spec-relaxation experiment graduates: the default economy profile moves to an 18 kHz passband at 58/38 taps per phase (minimal even counts meeting the 70 dB stopband with >= 1 dB margin on a 12.5 Hz sweep grid; up-direction 40 taps fails while 38 passes — sidelobe peaking is non-monotonic near threshold, recorded in PLAN section 4). That is 26%/14% fewer MACs per output and -25%/-14% table storage against the previous default, whose design continues bit-for-bit unchanged as balanced(). The 18-19 kHz shelf moves into the transition band (-1.4 dB at 19 kHz going down) — the same species of inaudible speed-first trade that set economy at 19 kHz rather than transparent's 20. Re-measured through every leg: scipy reference vectors regenerated for the six direction x profile cases; cross-validation floors re-pinned at 1.2e-5 down / 3.1e-5 up (-98/-90 dB), still equal at async L=512 and L=1024 (the floor remains the deliberate per-branch DC-normalization difference, which scales with the shorter designs' branch-sum spread); Q15 flagship at 76.5 dB (was 76.1) — the format floor, not the filter; 997 Hz float imaging floor 91.3 dB (was 89.2); bluetooth_bridge locks and recovers at 72.8 dB SNR, total latency 1.93 ms. The converter dispatches on three committed trip counts now (economy/balanced/transparent); the C ABI gains profile tag 2 = balanced; both notebooks re-executed against the shipping engine. Host suite 68/68 green. Contract change: the default output changes for every consumer that does not name a profile; balanced() is the bit-exact escape hatch. Embedded icount baselines for the six economy workloads are re-recorded in the follow-up commit harvested from this push's CI measurements. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
Fourth rung of the v0.3 ladder, from the same fine-grid criterion as the economy re-pin: 40/28 taps per phase (minimal even counts at 70 dB with >= 1 dB margin; 38 and 26 both fail) — half of balanced's MACs, 11.6/8.8 KiB stored f32 tables, 0.42/0.32 ms latency. Unlike economy's inaudible trade, this tier audibly shelves the top octave (-1.4 dB at 18 kHz, -5.9 dB at 19 kHz going down), so it is never a default: opt-in by name, positioned for voice/comms/Bluetooth-class links where the top octave is already gone. In Q15 it is quieter than transparent (77.7 dB measured at 997 Hz) at a fifth of the taps — the format floors near 76 dB regardless of tier, so cheap tiers are the right 16-bit pairing. Wired through every layer the other tiers get: constexpr profile with a committed trip count in the process() dispatch, scipy reference vectors (both directions), the design/phase-table/converter batteries (78 -> 82 host tests, all green), C ABI profile tag 3, ctypes bridge name, two Q15 icount scenarios (up_q15_se / down_q15_se) pinning the tier in its deployment format, and the design-spike notebook re-executed over all eight designs. Baselines for the new scenarios land with the harvest commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
notebooks/profile_ladder.ipynb puts the four v0.3 tiers side by side, all measured against the shipping C++ through the C ABI (family convention): overlaid prototype responses both directions, the top-octave zoom that is the tier-choice decision plot, engine-measured passband sweeps (each tier asserted flat within +/-0.02 dB inside its own passband), the four-way program-material spectra with the alias/image species annotated, grouped MACs/storage/latency cost bars, and a measured quality-vs-compute scatter (997 Hz float SNR: 85.3 / 91.3 / 89.2 / 144.9 dB for super_economy/economy/balanced/transparent). Tier colors are fixed across every figure and CVD-validated (adjacent-pair delta-E >= 8.6); every series is direct-labeled so identity never rides on color alone. Committed executed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
Harvested from the ratchet job's measurements on 924b737 (QEMU counts are deterministic; the job measures all targets even after a gate failure, the designed serial-harvest path). The six economy workloads land at the new 58/38-tap design: M33 -23.5/-25.5% (down q15/float) and -11.6/-13.6% (up), M55 -21.4/-22.6% and -8.7/-10.4%, Hexagon -21.8/-23.3% and -3.7/-10.6% — the tap-count cut, delivered on silicon paths. The two new super_economy Q15 scenarios record their first baselines (e.g. M55 down 47.7M vs economy's 58.3M). The transparent legs moved only inside the gate (+/-1.3% worst) and are re-recorded at measured per --update semantics. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
A branch push and its open PR fired the full matrix twice for the same commit (both run sets are visible on any PR of this branch). Concurrency groups keyed on the head SHA collapse the pair to one surviving run — the later-starting PR run cancels the push run — while a branch with no PR keeps its push-triggered CI, which the pre-PR baseline-harvest flow (PLAN section 7) depends on. Keyed per workflow so CI and the style gate never cancel each other. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
The SHA-keyed cancel-in-progress scheme left the cancelled push run's checks attached to the PR head commit, so the merge box read "10 cancelled, 10 successful — some checks haven't completed yet" on a fully green head. Push runs are now filtered to main: a PR branch gets exactly one run set (pull_request events) and a clean merge box; main keeps push coverage; a branch with no PR yet runs via workflow_dispatch (the pre-PR baseline-harvest path) or by opening the PR first. Concurrency now only cancels superseded in-flight runs per PR/ref — those cancellations attach to the old commit, never the current head. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C1fs1FmoRxYgVLwATD3HB9
DspTap #38 shares the Kaiser window's Bessel series across the prototype design (bit-identical coefficients). RatioTap designs at construction for every profile, so every workload's construction got cheaper: per profile and direction, identical across float/Q15/Q31 -- M33 -37..-299 M, Hexagon -4.7..-35 M, M55 -0.6..-4.6 M. Ten M33 and eight Hexagon scenarios left the two-sided gate (M33 down_q15_eco -28.9%); baselines re-recorded on all three targets. The test-only sampleratetap pin moves in step (no header changes in 5315689..2b4dff1) so both repos keep the identical dsptap tree the dev-only include path relies on. PLAN.md: M7e ledger entry, and a correction. M7c deferred coefficient baking because "construction is <0.3% of every workload"; a construct-only measurement of every scenario shows 3-5% on M55, 5-37% on Hexagon and up to 63% of the M33 Q15 workloads (74% before this change), which dilutes the M33 gate for hot-path regressions ~2-3x. Host: 78/78 under GCC and clang with TAP_RATIO_WERROR; scripts/tidy.sh clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
- icount.py (ported in step with SampleRateTap's): --exact, --json-out and --compare-json for same-job A/B measurement, a failure when a recorded workload has no binary, a fatal error on a missing baselines file, and the workload checksum printed and recorded. Hexagon workloads run from one fixed path with a fixed argv[0] and an empty environment, because qemu-hexagon copies those onto the guest stack and static musl's startup walks them; qemu-hexagon is resolved to an absolute path first. Hexagon baselines are re-recorded separately. - Tests carry a "ratio." ctest prefix and a ratio label (the bare-metal entry too), so names stay unique once they share a tree with SampleRateTap's. - New OutputHash suite: FNV-1a hashes of the converter's output for both directions, every format and every profile, printed for same-job comparison and never pinned. About 8 s on Cortex-M33. - The cross-validation lines now print their tolerance, so a loosened limit is visible in the output and not only in the source. - CI: contents: read permissions; cancellation spares main and the migration PR; every action SHA-pinned (checkout v6, cache v5, as in SampleRateTap); ubuntu-24.04 pinned for QEMU and ratchet jobs, with the image OS in the plugin-qemu cache key and image/toolchain versions logged; --no-tests=error on every ctest; QEMU legs keep and upload full test logs. Step P.2 of the monorepo migration plan (SampleRateTap docs/MONOREPO_PLAN.md). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The previous commit runs Hexagon workloads from a fixed path with a fixed argv[0] and an empty environment. The first isolated CI run (PR #18) measured every scenario 3,724 to 3,809 instructions lower (-0.0004% to -0.0136%): the path and environment strings static musl's startup used to walk. Inside the gate, but re-recorded so exact comparisons start from the new harness; PLAN.md section 7 ledger entry. M33/M55 are unaffected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
- requirements.in / requirements.lock at the repository root: the notebook environment pinned with hashes (numpy 2.4.6, scipy 1.17.1, matplotlib 3.11.2, jupyter/nbconvert; samplerate and soxr for SampleRateTap's comparison notebook, since the file is identical in both repositories). It replaces notebooks/requirements.txt, which named packages without versions. - ratiotap_py rebuilds build_capi/ incrementally on every import instead of only when the library is missing, so a library left over from an older checkout can no longer be measured silently, and the build is quiet (its log used to land in ratio_demo's committed output; it now prints only on failure). - notebooks/figure_digest.py wraps plt.show() to print a digest of each figure's plotted data, quantized to 9 significant digits of each array's peak, so a changed curve shows up as changed text; stable across re-executions. - All three notebooks re-executed in the pinned environment: every committed number reproduces exactly; the only output changes are the added digest lines and the removed build log. Step P.3 of the monorepo migration plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Moves submodules/sampleratetap 2b4dff1 -> 5e2057f, SampleRateTap's main after its step-P PR (#48). That range changes no header under include/ and keeps the same DspTap pin (0eb09fa), so the tree the cross-validation compiles against is unchanged: 82/82 tests pass and the four cross-validation lines are byte-identical to before. The monorepo migration's step 0 snapshots against this pin, so the cross-validation lines it records come from the same async tree the import compiles (step P.4 of the plan). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Proposes merging RatioTap into SampleRateTap as one repository with per-engine directories (async/, ratio/), the tap::sr::<engine> namespace, and a history-preserving import. Records the decisions taken, a measured inventory of both repositories, a gated five-step migration, risks, open questions and an adversarial audit checklist. Nothing is executed yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Folds in the five-reviewer audit (78 findings, 6 blockers). Main changes: - Correct the step-1 recipe: filter-repo rewrites .gitmodules throughout history; a pure-move commit before the merge keeps blame; build glue (root enable_testing, single dsptap, forced test options) and ratio's CI port happen in step 1, not later. - Define gates once (G1-G13): test multisets, exact icount with workload set equality, disassembly and compile-flag identity, full-output hashes, all as same-job A/B against the step-0 SHA. - Add a pre-work step fixing existing breakage (arm64 TSan filter, Pico 2 builds, blame-ignore file) and hardening the harness (Hexagon argv/env, icount --exact) before the snapshot. - Require a merge commit for the final PR; document first-parent bisect. - Replace the chain invariant with a passband rule; restate ratio's charter at 44.1*2^k <-> 48*2^k; keep decimate.h and the async datapath out of the move; forbid rate routing through chain<>. - Correct the inventory (136 commits, functional infra diffs, CI table, exact C ABI counts, 25 DspTap references) and add a file disposition table and an appendix mapping every finding to its resolution. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Adds D13 (one family version, 0.4.0, bit-packed tap_sr_version), D14 (one copyright holder line; ratio/LICENSE retired in the commit that adds it to the root) and D15 (async_sample_rate_converter -> converter), accepts the two-level tap::sr namespace, places icount tables in each engine's README, and records the confirmed outside-repository facts (no external consumers, taphouse syncs RatioTap, merge commits allowed). Every open question is now resolved. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Folds in the round-2 audit (78 findings from an end-to-end dry run of steps 1a-1c, gate prototypes, a CI design review and a coherence check): - Correct the step-1 recipe: resolve the .gitmodules conflict with git add, spell out the duplicate removal, filter only main, and guard engine CMake with a ../submodules/dsptap path so standalone builds and notebook bridges keep working. Root options become defaults, never forced; GTEST_HAS_* stays in each engine's test tree. - Split gates into same-job A/B (icount, codegen, output hashes, flags, notebooks) run by a dedicated migration-gates workflow, and snapshot gates. Scope codegen identity to icount and C ABI binaries, never pin output hashes as constants, and add G14, a rename-only residual diff that catches edits no other gate sees. - Give tests an engine prefix and label so duplicate names no longer alias, select engines at ctest time, and push one gated commit at a time with cancellation disabled on the migration PR. - Rewrite the coverage rule in terms of stopband edge, attenuation and summed ripple; restate the ratio 2^k follow-up as an API change. - Specify D13 version mechanics, correct D14's rationale (the author holds SampleRateTap's copyright) and add banners to every source, and give D15 its full mapping. - Commit the audit's draft workflows and root CMake under docs/migration/drafts/ for review. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
RatioTap's engine becomes `bridge` (tap::sr::bridge, bridge/ from the step-1b import on) and the future 2^a*3^b engine becomes `rational`, replacing `ratio` and `integer`. The Python ctypes modules are now called bindings so "bridge" names only the engine. D16 records that step P landed RatioTap's prefix as `ratio.`, renamed at step 3.4 through the G1/G2 name map, and the CI draft carries a `label` key until then. Step 0 also gains the note that the G2 collector strips ctest's test-number prefix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
S0 = SampleRateTap 5e2057f (140 commits) and R0 = RatioTap 8f19e8b (33), from fresh full clones (docs/migration/tips.txt). docs/migration/ now holds what the gates compare against: - rename.py: the mechanical rename map (paths, namespaces, macros, CMake, C ABI, icount prefixes, D14 banners) and gate G14. Applied through 3.8 to S0 and R0 it builds with -Werror, passes 77/77 and 82/82, reproduces both tips' output hashes and cross-validation lines, and reports a loosened tolerance as one residual hunk. residual/1c.txt seeds 1c's allowlist from the plan. - collect.py: collectors for G1, G2, G6, G9, G10 and G12. - snapshot/: G1 test lists per CI job and labels, G2 on-target [ RUN ] lists from the QEMU artifacts, G6 cross-validation lines, G9 rules, G10 C ABI symbols, G12 history; provenance in runs.md. The plan's step 0 records the result and three map details it left implicit; step 1c's HISTORY.md count follows R0 (33, not 29). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
A pure move: every row 1a of MONOREPO_PLAN.md 4.1 with git mv, nothing edited, so git show -M reports no insertions or deletions. The tree does not build on its own; 1c adds the root build glue. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Joins RatioTap main at R0 = 8f19e8b (33 commits) with an unrelated-histories merge. The history was rewritten by git filter-repo --refs main into bridge/, with .gitmodules rewritten throughout so every imported commit still checks out its submodules; the rewritten tip is 654659e, identical on two fresh clones. RatioTap's reformat commit c0894cf is 89c7eba here (for .git-blame-ignore-revs after step 4). At the merge: root .gitmodules kept (the one add/add conflict), the bridge/submodules gitlinks removed, and the files byte-identical to the root copies deleted (.clang-format, .clang-tidy, STYLE.md, .pre-commit-config.yaml, .claude/, scripts/tidy.sh, the PR template). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The root CMakeLists composes async/ and bridge/ over one tap::dsp; each engine keeps its project() and options until step 3 and still configures on its own (cmake -S async|bridge, bridge/tools/capi) through a guarded dsptap add. bridge's dev-only view of the async headers now points at ../async/include. The M33/M55 toolchain files set both engines' BARE_METAL variables. CI configures the root once per job. Host legs keep each engine's warning policy (MSVC: async /W4, bridge /WX); the QEMU legs run per (target, engine) by ctest label with each engine's own exclusions and -j; the ratchet is one matrix job per engine, each with its own plugin marker and icount.py until step 2, and each engine README carries its icount table. style.yml is RatioTap's body over both engines; the book job also builds the API reference. migration-gates.yml runs docs/migration/gates.py against S0 and R0 rebuilt in the same job. Paths follow the 1a/1b moves: the 52 book includes (the rendered book is byte-identical to S0's), Doxyfile, book-pages filters, compare.yml, the scripts, the notebook bindings, README links and bridge build commands. LICENSE carries D14's holder line as bridge/LICENSE goes; bridge's duplicate lockfile, .gitignore and workflows go; bridge/docs/HISTORY.md maps RatioTap's SHAs to the imported ones. Gates measured locally against S0 and R0 built in the same session: G1, G3+G5 (M33, M55: counts and checksums exact), G4 (C ABI and 17 icount binaries), G5 host hashes, G6, G7, G10, G11 (7 notebooks), G12 and G14 all pass; Hexagon A/B and the macOS/Windows legs run first in CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The drift check guards scripts/tidy.sh as well as the style configs, and 1c had edited it (compile-database flags and the submodule filter). The CI gate in style.yml carries the monorepo-specific configuration; the local mirror stays the shared file and sweeps the default configure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
G12 fails in CI for every file while passing locally on the same history (git 2.55 there, 2.43 here); print what the gated log lacks so the difference is visible in the job log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
git 2.55 (the runner) prints a UTC author date in %aI as 'Z'; git 2.43, which recorded the snapshot, prints '+00:00'. Every entry therefore differed in CI while the history itself was intact (the diagnostic from the previous commit shows the same commits in the same order). Compare '<epoch> <subject>' instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
This PR carries out
docs/MONOREPO_PLAN.md. SampleRateTap and RatioTap become one repository for thetap::srfamily:async/holds the current SampleRateTap engine.bridge/holds RatioTap's engine, imported with its history.The work proceeds one gated commit at a time through steps 1–4. This PR must be merged with "Create a merge commit" (D3). A squash or rebase would erase the imported RatioTap history.
The branch so far:
f2b7d04): the frozen starting points, the snapshot,rename.py(the mechanical rename map and residual check G14), and the gate collectors.5297394):async/as a pure move.mainunderbridge/withgit filter-repoand an unrelated-histories merge. The rewritten tip is654659e, the same on two fresh clones.migration-gates.yml.tidy.sh, and making G12 compare dates as epochs.Why
The engines are built to be composed and cross-checked against each other. Every engine added as a separate repository multiplies submodule pins, harness copies and cross-repository test dependencies. D1 has the full reasoning.
Verification
Step 1 is gated green at
5297394:What the gates showed, with SampleRateTap and RatioTap at their step-0 commits (
S0,R0) rebuilt in the same job:--followreaching every step-0 commit, with no file owned wholesale by a migration commit.docs/migration/residual/1c.txt.[ RUN ]lists: identical to the snapshot in all 30 comparisons, covering the six host jobs and the six QEMU legs. Linux Clang now runs the bridge suite, which is new coverage that matches RatioTap's Linux list.Notes for the reviewer
bridge, and the future 2^a·3^b engine isrational. Until step 3.4, the bridge tests keep RatioTap'sratio.prefix andratiolabel.SrtHandlebecomestap_sr_async_converter, to matchtap_sr_bridge_converter.🤖 Generated with Claude Code
https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA