Correct the M33 budget (stereo Q15 fits one core); refresh stale figures; fix clang sign-conversion - #46
Merged
Conversation
state.iterations() is already int64 (IterationCount); cast the size_t block constant instead of the product, so clang -Wconversion (which implies -Wsign-conversion) accepts both benches under SRT_WERROR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
The book, README and Pico docs derived per-frame M33 costs by dividing a 2 s pipeline baseline by its 96,000 frames. That quotient also carries the converter's one-time construction (soft-double filter design, input synthesis), so it overstated the per-frame cost ~4x and led to the conclusion that 48 kHz Q15 stereo is infeasible on one 150 MHz core. Measured by building the pipeline workloads at 2 s and 4 s and taking the difference (new SRT_SC_SECONDS, default 2; the gated binaries' .text is byte-identical to main): pipeline_q15 1,138 insn/frame steady, ~1.31 G construction pipeline12_q15 3,326 insn/frame steady, ~1.58 G construction The June commit c609a0f measures the same per-frame costs (1,138 / 3,326), so the old 5,043 / 10,027 were construction-inflated from the start; only construction has grown since (375 M -> 1.31 G with the compensated design). Against 3,125 cycles/frame at 150 MHz / 48 kHz: stereo fits with room to spare (~36% at one instruction per cycle); 12 channels still does not fit at 48 kHz, so the dual-core example's 16 kHz phase stands, on a far smaller margin than "more than 3x over". Construction is now stated as its own cost (seconds at boot). All still pending real-silicon validation. Also refreshes the book's -132.1 dB (now -133.9) and June host-throughput figures from docs/COMPARISON.md, and describes the 2 s / 4 s method in the icount chapter and the Pico calibration READMEs (whose denominators were the inflated quotients). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
The ~1.3 G figure is QEMU's soft-double count. docs/PERFORMANCE.md already records that the RP2350 routes double arithmetic through its DCP coprocessor, so a real Pico 2 should pay less; say so wherever the new text priced construction in seconds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Follow-ups to #45:
HARDWARE_TESTING.mdand both Pico examples now state per-frame M33 costs as steady state (servo and FIFO included): 1,138 insn/frame for stereo Q15 and 3,326 for 12 channels. They used to say 5,043 and 10,027. That changes a conclusion: 48 kHz Q15 stereo fits one 150 MHz core; the old text called it infeasible.docs/COMPARISON.md.docs/PERFORMANCE.md.SRT_SC_SECONDS(default 2) lets the ratchet workload be built at 4 s, so steady state can be separated from construction.-Werrorfix. Two benchmarkSetItemsProcessedcalls had a sign conversion that failed clang-Werror.Why
The old figures divided a 2 s pipeline baseline by its 96,000 frames. That quotient also carries the converter's one-time construction, which is soft-double filter design on the M33, so it overstated the per-frame cost about 4×. I built the June commit
c609a0fthe same way: its per-frame costs are the same 1,138 / 3,326. The old numbers were inflated from the start. Only construction has grown since (375 M → 1.31 G with the compensated design).Verification
pipeline_q151,417,933,654 at 2 s / 1,527,220,570 at 4 s;pipeline12_q151,896,402,851 / 2,215,675,798.c609a0f: 484,257,006 / 593,544,599 and 962,728,307 / 1,281,999,472..textis byte-identical tomain, both at the default and at an explicitSRT_SC_SECONDS=2, so no baseline moves.-Werrorbuild of every optional target (benchmarks, compare bench, C ABI, shim, icount compare) now has 0 errors.pico2_cyccntmeasurement.Notes for the reviewer
fast()or the second core" to "stereo fits one core". The 12-channel conclusion (16 kHz in the dual-core example) stands, on a much smaller margin: 3,326 against 3,125 at one instruction per cycle.pico2_cyccnt's calibration now divides cycles/frame by the steady-state counts. The old denominators included setup the firmware never times.🤖 Generated with Claude Code
https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
Generated by Claude Code