diff --git a/README.md b/README.md index c26973e..d213f8d 100644 --- a/README.md +++ b/README.md @@ -273,7 +273,7 @@ CI builds and tests every push on: channels, against the 3,125 cycles per frame a 150 MHz core has at 48 kHz: 48 kHz Q15 stereo fits one core with room to spare, 12 channels does not (the dual-core example runs it at 16 kHz). Construction is - separate and heavy: ~1.3 G QEMU instructions of soft-double filter + separate and heavy: ~0.9 G QEMU instructions of soft-double filter design, seconds at boot on a generic FP64-less core (less on the RP2350, whose DCP coprocessor handles doubles). Instruction counts are not cycle counts, so treat these as budgets pending real-silicon @@ -327,13 +327,13 @@ Executed instructions per fixed workload (`bench/icount/`), measured under QEMU | Workload | Cortex-M33 | Cortex-M55 | Hexagon | |---|---:|---:|---:| -| `kernel_float` | 2,825,389,539 | 115,540,990 | 470,171,975 | -| `kernel_q15` | 1,515,719,465 | 198,066,650 | 234,098,419 | -| `kernel_q31` | 1,562,793,261 | 226,862,081 | 241,733,702 | -| `pipeline12_q15` | 1,891,351,546 | 403,861,915 | 510,167,719 | -| `pipeline_float` | 2,789,479,856 | 108,833,486 | 467,197,861 | -| `pipeline_q15` | 1,412,881,372 | 143,427,077 | 251,145,660 | -| `pipeline_q31` | 1,495,389,428 | 178,785,673 | 251,993,512 | +| `kernel_float` | 2,427,595,993 | 109,298,076 | 422,164,193 | +| `kernel_q15` | 1,123,119,218 | 192,431,957 | 187,432,836 | +| `kernel_q31` | 1,169,957,251 | 221,164,795 | 194,967,771 | +| `pipeline12_q15` | 1,498,750,507 | 398,227,192 | 463,502,418 | +| `pipeline_float` | 2,391,686,215 | 102,590,477 | 419,190,309 | +| `pipeline_q15` | 1,020,281,310 | 137,792,386 | 204,480,261 | +| `pipeline_q31` | 1,102,554,828 | 173,088,119 | 205,227,716 | diff --git a/bench/baselines.json b/bench/baselines.json index 48fed95..e7b294d 100644 --- a/bench/baselines.json +++ b/bench/baselines.json @@ -1,29 +1,29 @@ { "hexagon": { - "kernel_float": 470171975, - "kernel_q15": 234098419, - "kernel_q31": 241733702, - "pipeline12_q15": 510167719, - "pipeline_float": 467197861, - "pipeline_q15": 251145660, - "pipeline_q31": 251993512 + "kernel_float": 422164193, + "kernel_q15": 187432836, + "kernel_q31": 194967771, + "pipeline12_q15": 463502418, + "pipeline_float": 419190309, + "pipeline_q15": 204480261, + "pipeline_q31": 205227716 }, "m33": { - "kernel_float": 2825389539, - "kernel_q15": 1515719465, - "kernel_q31": 1562793261, - "pipeline12_q15": 1891351546, - "pipeline_float": 2789479856, - "pipeline_q15": 1412881372, - "pipeline_q31": 1495389428 + "kernel_float": 2427595993, + "kernel_q15": 1123119218, + "kernel_q31": 1169957251, + "pipeline12_q15": 1498750507, + "pipeline_float": 2391686215, + "pipeline_q15": 1020281310, + "pipeline_q31": 1102554828 }, "m55": { - "kernel_float": 115540990, - "kernel_q15": 198066650, - "kernel_q31": 226862081, - "pipeline12_q15": 403861915, - "pipeline_float": 108833486, - "pipeline_q15": 143427077, - "pipeline_q31": 178785673 + "kernel_float": 109298076, + "kernel_q15": 192431957, + "kernel_q31": 221164795, + "pipeline12_q15": 398227192, + "pipeline_float": 102590477, + "pipeline_q15": 137792386, + "pipeline_q31": 173088119 } } diff --git a/book/src/part0/budgets.md b/book/src/part0/budgets.md index c496aa4..79c937a 100644 --- a/book/src/part0/budgets.md +++ b/book/src/part0/budgets.md @@ -274,7 +274,9 @@ instructions on this core) over the audio. The fix is to measure the workload at two lengths: the difference is the per-frame cost, the remainder is construction. The per-frame cost was 1,138 instructions both then and now; what changed since is construction, which grew to ~1.3 G -instructions with the compensated design. That is still a real budget +instructions with the compensated design and then fell to ~0.9 G once the +design shared its Kaiser window's Bessel series between kernel builds +(bit-identical coefficients, DspTap #38). That is still a real budget item: seconds of start-up on a generic 150 MHz part with no FP64 help, paid once (the RP2350's DCP coprocessor for double arithmetic should make a Pico 2 cheaper than the emulated count). Knowing both diff --git a/book/src/part2/icount.md b/book/src/part2/icount.md index 9f277ca..d8e604e 100644 --- a/book/src/part2/icount.md +++ b/book/src/part2/icount.md @@ -115,7 +115,7 @@ accumulates a checksum, and ends with: A total is the whole binary's cost, construction included, so a pipeline baseline divided by its 96 000 frames is *not* the per-frame cost: on the -M33 the converter's soft-double filter design alone is over a billion +M33 the converter's soft-double filter design alone is close to a billion instructions. `SRT_SC_SECONDS` (default 2, the only length ever baselined) exists for that question. Build the pipeline scenarios again at `-DSRT_SC_SECONDS=4` and difference the counts: what doubles is the diff --git a/book/src/part4/cortex-m.md b/book/src/part4/cortex-m.md index 0d97d71..0cbf2a3 100644 --- a/book/src/part4/cortex-m.md +++ b/book/src/part4/cortex-m.md @@ -327,11 +327,11 @@ accelerates. **Budgets, stated as instructions, pending cycles.** A baseline divided by its 96,000 frames is *not* the per-frame cost: each workload also constructs its converter, and on the M33 the soft-double filter design -alone is over a billion instructions. Building the pipeline workloads a +alone is close to a billion instructions. Building the pipeline workloads a second time at 4 s of audio (`-DSRT_SC_SECONDS=4`) and taking the difference isolates the steady state: **≈ 1,138 instructions per stereo frame** for `pipeline_q15` and ≈ 3,326 for the 12-channel shape, with -~1.31 G and ~1.58 G of one-time construction. (Earlier revisions quoted +~0.91 G and ~1.18 G of one-time construction. (Earlier revisions quoted 5,043 and 10,027, the construction-inclusive quotients; the per-frame costs have not changed since June.) A 150 MHz core at 48 kHz has 3,125 *cycles* per frame. The README draws the conclusion in diff --git a/docs/COMPARISON.md b/docs/COMPARISON.md index 8b3cab6..9ff711f 100644 --- a/docs/COMPARISON.md +++ b/docs/COMPARISON.md @@ -171,8 +171,8 @@ One-time construction, millions of instructions: | Engine | Cortex-M55 | Cortex-M33 | Hexagon | |---|---:|---:|---:| -| **SampleRateTap** balanced, float | 23.6 | 1,268 | 181 | -| **SampleRateTap** balanced, Q15 | 24.6 | 1,281 | 184 | +| **SampleRateTap** balanced, float | 17.3 | 870 | 133 | +| **SampleRateTap** balanced, Q15 | 18.4 | 884 | 136 | | r8brain 120 dB, default 2 % band | 1.6 | 38.1 | 11.3 | | r8brain 120 dB, 8 % band | 1.8 | 45.5 | 12.6 | | libsamplerate `MEDIUM` | 1.4 | 20.9 | 7.4 | @@ -193,8 +193,11 @@ the Cortex-M builds force-include `bench/icount/r8b_single_thread_mutex.h` per-sample path). Hexagon's musl build uses the real mutex. **Construction is the one column SampleRateTap loses.** Its filter design -(the compensated prototype, run in double at construction) costs ~1.3 G -instructions on the M33 as QEMU emulates it (every double operation a +(the compensated prototype, run in double at construction) costs ~0.9 G +instructions on the M33 as QEMU emulates it (~1.3 G until DspTap shared the +Kaiser window's Bessel series across the design's kernel builds; the +construction rows above are re-measured at that pin, the steady-state rows +did not move by a single instruction) (every double operation a software libcall), against tens of millions for r8brain and libsamplerate. On a generic FP64-less 150 MHz core that is seconds of start-up; the RP2350 routes double arithmetic through its DCP coprocessor, diff --git a/docs/PERFORMANCE.md b/docs/PERFORMANCE.md index 6348602..1d94c66 100644 --- a/docs/PERFORMANCE.md +++ b/docs/PERFORMANCE.md @@ -130,13 +130,29 @@ table is already enforced by test thresholds. algorithm-change rule; the two-sided ratchet flagged both directions during the process (the M55 leg failed as 'IMPROVED beyond tolerance' when the one-pass reduction landed on top of the two-pass baselines). +- [x] **Shared Kaiser-window Bessel series (DspTap #38, pin bump + 28a34a1 → 0eb09fa)** — both prototype designs now evaluate the window's + `bessel_i0` series once, for half the taps, and mirror it; the + compensated design used to evaluate it per tap in each of its two + kernel builds. Bit-identical coefficients (FNV-1a-64 over the whole + prototype, on M33 and on x86 under GCC and clang), so the audio path is + untouched: the comparison workloads' steady state (4 s minus 2 s, + docs/COMPARISON.md) is identical to the instruction on all three + targets. Construction only, constant across all seven scenarios per + target to within the scenarios' own spread: M33 −393 to −398M, Hexagon + −47M, M55 −6M instructions (M33 `pipeline_q15` −27.8%, the largest + relative move). Baselines re-recorded with this entry as the + justification; the ratchet flagged every scenario 'IMPROVED beyond + tolerance' before the update. The pin also carries DspTap's FFT + stages, which SampleRateTap does not use. ## Known debt - **Constructor cost on soft-FP64 targets, and ratchet sensitivity.** The compensated design costs ~7M double flops at construction — pennies on - hosts and the M55 (hardware FP64), but ~1G instructions on the QEMU - M33 (soft-double libcalls): seconds of boot on an M33-class part, and + hosts and the M55 (hardware FP64), but ~0.9G instructions on the QEMU + M33 (soft-double libcalls; ~1.3G before the shared-window entry + above): seconds of boot on an M33-class part, and a large share of each M33 icount scenario's total, which dilutes the +/-3% gate's sensitivity to hot-path regressions on that target. Mitigations queued: split construction into its own ratcheted scenario diff --git a/examples/pico2_cyccnt/README.md b/examples/pico2_cyccnt/README.md index e7dcec9..fffc8de 100644 --- a/examples/pico2_cyccnt/README.md +++ b/examples/pico2_cyccnt/README.md @@ -17,7 +17,7 @@ steady-state instruction counts (`pipeline_q15` 1,138/frame, cycles" ratio, turning every current and future M33 baseline into a real cycle budget. The steady state is the difference of the workload built at 2 s and at 4 s (`-DSRT_SC_SECONDS=4`); a committed baseline divided by its -96,000 frames also carries the converter's one-time construction (~1.3 G +96,000 frames also carries the converter's one-time construction (~0.9 G instructions of soft-double filter design), which this firmware does not time. It also tests the README's claim directly: Q15 stereo fits a 150 MHz core with room to spare, 12 channels at 48 kHz does not. diff --git a/examples/pico2_dualcore/README.md b/examples/pico2_dualcore/README.md index d4391ee..e2f22e1 100644 --- a/examples/pico2_dualcore/README.md +++ b/examples/pico2_dualcore/README.md @@ -111,7 +111,7 @@ SRT_PICO2_DUALCORE_DONE workload built at 2 s and at 4 s, `-DSRT_SC_SECONDS=4`, and differenced). A committed baseline divided by its 96,000 frames is larger — it also carries one-time setup (the soft-double filter design, input synthesis, - ~1.3–1.6 G instructions) — so the difference is the right counterpart of + ~0.9–1.2 G instructions) — so the difference is the right counterpart of the steady-state loop this firmware times; `cyc_frame ÷ insns/frame` from the sibling `examples/pico2_cyccnt` run gives the cycles-per-instruction diff --git a/submodules/dsptap b/submodules/dsptap index 28a34a1..0eb09fa 160000 --- a/submodules/dsptap +++ b/submodules/dsptap @@ -1 +1 @@ -Subproject commit 28a34a18c40bda74b5d8cea0c54e09f31fbdc94b +Subproject commit 0eb09fa1edf5ce4f83043a05ce04ed8f3a653dbc