Skip to content

Correct the M33 budget (stereo Q15 fits one core); refresh stale figures; fix clang sign-conversion - #46

Merged
tap merged 3 commits into
mainfrom
claude/sample-rate-solutions-comparison-mqc190
Sep 26, 2026
Merged

tap merged 3 commits into
mainfrom
claude/sample-rate-solutions-comparison-mqc190

Conversation

@tap

@tap tap commented Sep 26, 2026

Copy link
Copy Markdown
Owner

What this changes

Follow-ups to #45:

  • M33 budget correction. The README, book, HARDWARE_TESTING.md and both Pico examples now state per-frame M33 costs as steady state (servo and FIFO included): 1,138 insn/frame for stereo Q15 and 3,326 for 12 channels. They used to say 5,043 and 10,027. That changes a conclusion: 48 kHz Q15 stereo fits one 150 MHz core; the old text called it infeasible.
  • Stale figures refreshed. The book's −132.1 dB (now −133.9) and its June host-throughput numbers now match docs/COMPARISON.md.
  • Construction cost qualified. Where the text prices construction in seconds, it now notes that the RP2350's DCP coprocessor makes a Pico 2 cheaper than the soft-double count suggests. This is already recorded in docs/PERFORMANCE.md.
  • SRT_SC_SECONDS (default 2) lets the ratchet workload be built at 4 s, so steady state can be separated from construction.
  • clang -Werror fix. Two benchmark SetItemsProcessed calls had a sign conversion that failed clang -Werror.

Why

The old figures divided a 2 s pipeline baseline by its 96,000 frames. That quotient also carries the converter's one-time construction, which is soft-double filter design on the M33, so it overstated the per-frame cost about 4×. I built the June commit c609a0f the same way: its per-frame costs are the same 1,138 / 3,326. The old numbers were inflated from the start. Only construction has grown since (375 M → 1.31 G with the compensated design).

Verification

  • Steady-state counts. M33 counts come from QEMU with the counting plugin (same plugin header SHA as CI).
    • Today: pipeline_q15 1,417,933,654 at 2 s / 1,527,220,570 at 4 s; pipeline12_q15 1,896,402,851 / 2,215,675,798.
    • June c609a0f: 484,257,006 / 593,544,599 and 962,728,307 / 1,281,999,472.
    • My 2 s counts sit +0.36 % above the committed baselines, a constant ~5.05 M on both workloads. It looks like a startup-code difference between toolchains, and it cancels in the difference.
  • Ratchet binaries unchanged. The gated binary's .text is byte-identical to main, both at the default and at an explicit SRT_SC_SECONDS=2, so no baseline moves.
  • Book build. It builds with CI's pinned mdBook v0.4.40 (checksum verified) with 0 warnings/errors, so no stale anchors.
  • clang. A clang -Werror build of every optional target (benchmarks, compare bench, C ABI, shim, icount compare) now has 0 errors.
  • Not re-run. The test suite, since no library code changed. The Pico firmwares are unbuilt (no Pico SDK here); only their comments changed, and clang-format is clean.
  • Not verified on silicon. The budget statements are still instruction-space, pending the pico2_cyccnt measurement.

Notes for the reviewer

  • Contract change (documentation): the README's M33 deployment guidance flips from "stereo wants fast() or the second core" to "stereo fits one core". The 12-channel conclusion (16 kHz in the dual-core example) stands, on a much smaller margin: 3,326 against 3,125 at one instruction per cycle.
  • Pico calibration. pico2_cyccnt's calibration now divides cycles/frame by the steady-state counts. The old denominators included setup the firmware never times.
  • Related DspTap change: tap/DspTap has a bit-identical change that cuts the M33 filter design by 32 % (~400 M instructions per construction). Bumping this repo's pin to it will lower every construction-inclusive icount baseline beyond ±3 %, so that bump needs re-recorded baselines from CI logs.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK


Generated by Claude Code

state.iterations() is already int64 (IterationCount); cast the size_t
block constant instead of the product, so clang -Wconversion (which
implies -Wsign-conversion) accepts both benches under SRT_WERROR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
The book, README and Pico docs derived per-frame M33 costs by dividing a
2 s pipeline baseline by its 96,000 frames. That quotient also carries the
converter's one-time construction (soft-double filter design, input
synthesis), so it overstated the per-frame cost ~4x and led to the
conclusion that 48 kHz Q15 stereo is infeasible on one 150 MHz core.

Measured by building the pipeline workloads at 2 s and 4 s and taking the
difference (new SRT_SC_SECONDS, default 2; the gated binaries' .text is
byte-identical to main):

  pipeline_q15    1,138 insn/frame steady, ~1.31 G construction
  pipeline12_q15  3,326 insn/frame steady, ~1.58 G construction

The June commit c609a0f measures the same per-frame costs (1,138 / 3,326),
so the old 5,043 / 10,027 were construction-inflated from the start; only
construction has grown since (375 M -> 1.31 G with the compensated design).

Against 3,125 cycles/frame at 150 MHz / 48 kHz: stereo fits with room to
spare (~36% at one instruction per cycle); 12 channels still does not fit
at 48 kHz, so the dual-core example's 16 kHz phase stands, on a far
smaller margin than "more than 3x over". Construction is now stated as
its own cost (seconds at boot). All still pending real-silicon validation.

Also refreshes the book's -132.1 dB (now -133.9) and June host-throughput
figures from docs/COMPARISON.md, and describes the 2 s / 4 s method in the
icount chapter and the Pico calibration READMEs (whose denominators were
the inflated quotients).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
The ~1.3 G figure is QEMU's soft-double count. docs/PERFORMANCE.md already
records that the RP2350 routes double arithmetic through its DCP
coprocessor, so a real Pico 2 should pay less; say so wherever the new
text priced construction in seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
@tap
tap merged commit 7658782 into main Sep 26, 2026
32 checks passed
@tap
tap deleted the claude/sample-rate-solutions-comparison-mqc190 branch September 26, 2026 03:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants