Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 8 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -273,7 +273,7 @@ CI builds and tests every push on:
channels, against the 3,125 cycles per frame a 150 MHz core has at
48 kHz: 48 kHz Q15 stereo fits one core with room to spare, 12 channels
does not (the dual-core example runs it at 16 kHz). Construction is
separate and heavy: ~1.3 G QEMU instructions of soft-double filter
separate and heavy: ~0.9 G QEMU instructions of soft-double filter
design, seconds at boot on a generic FP64-less core (less on the RP2350,
whose DCP coprocessor handles doubles). Instruction counts are not
cycle counts, so treat these as budgets pending real-silicon
Expand Down Expand Up @@ -327,13 +327,13 @@ Executed instructions per fixed workload (`bench/icount/`), measured under QEMU

| Workload | Cortex-M33 | Cortex-M55 | Hexagon |
|---|---:|---:|---:|
| `kernel_float` | 2,825,389,539 | 115,540,990 | 470,171,975 |
| `kernel_q15` | 1,515,719,465 | 198,066,650 | 234,098,419 |
| `kernel_q31` | 1,562,793,261 | 226,862,081 | 241,733,702 |
| `pipeline12_q15` | 1,891,351,546 | 403,861,915 | 510,167,719 |
| `pipeline_float` | 2,789,479,856 | 108,833,486 | 467,197,861 |
| `pipeline_q15` | 1,412,881,372 | 143,427,077 | 251,145,660 |
| `pipeline_q31` | 1,495,389,428 | 178,785,673 | 251,993,512 |
| `kernel_float` | 2,427,595,993 | 109,298,076 | 422,164,193 |
| `kernel_q15` | 1,123,119,218 | 192,431,957 | 187,432,836 |
| `kernel_q31` | 1,169,957,251 | 221,164,795 | 194,967,771 |
| `pipeline12_q15` | 1,498,750,507 | 398,227,192 | 463,502,418 |
| `pipeline_float` | 2,391,686,215 | 102,590,477 | 419,190,309 |
| `pipeline_q15` | 1,020,281,310 | 137,792,386 | 204,480,261 |
| `pipeline_q31` | 1,102,554,828 | 173,088,119 | 205,227,716 |
<!-- ICOUNT:END -->

<!-- PERF:BEGIN -->
Expand Down
42 changes: 21 additions & 21 deletions bench/baselines.json
Original file line number Diff line number Diff line change
@@ -1,29 +1,29 @@
{
"hexagon": {
"kernel_float": 470171975,
"kernel_q15": 234098419,
"kernel_q31": 241733702,
"pipeline12_q15": 510167719,
"pipeline_float": 467197861,
"pipeline_q15": 251145660,
"pipeline_q31": 251993512
"kernel_float": 422164193,
"kernel_q15": 187432836,
"kernel_q31": 194967771,
"pipeline12_q15": 463502418,
"pipeline_float": 419190309,
"pipeline_q15": 204480261,
"pipeline_q31": 205227716
},
"m33": {
"kernel_float": 2825389539,
"kernel_q15": 1515719465,
"kernel_q31": 1562793261,
"pipeline12_q15": 1891351546,
"pipeline_float": 2789479856,
"pipeline_q15": 1412881372,
"pipeline_q31": 1495389428
"kernel_float": 2427595993,
"kernel_q15": 1123119218,
"kernel_q31": 1169957251,
"pipeline12_q15": 1498750507,
"pipeline_float": 2391686215,
"pipeline_q15": 1020281310,
"pipeline_q31": 1102554828
},
"m55": {
"kernel_float": 115540990,
"kernel_q15": 198066650,
"kernel_q31": 226862081,
"pipeline12_q15": 403861915,
"pipeline_float": 108833486,
"pipeline_q15": 143427077,
"pipeline_q31": 178785673
"kernel_float": 109298076,
"kernel_q15": 192431957,
"kernel_q31": 221164795,
"pipeline12_q15": 398227192,
"pipeline_float": 102590477,
"pipeline_q15": 137792386,
"pipeline_q31": 173088119
}
}
4 changes: 3 additions & 1 deletion book/src/part0/budgets.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,7 +274,9 @@ instructions on this core) over the audio. The fix is to measure the
workload at two lengths: the difference is the per-frame cost, the
remainder is construction. The per-frame cost was 1,138 instructions both
then and now; what changed since is construction, which grew to ~1.3 G
instructions with the compensated design. That is still a real budget
instructions with the compensated design and then fell to ~0.9 G once the
design shared its Kaiser window's Bessel series between kernel builds
(bit-identical coefficients, DspTap #38). That is still a real budget
item: seconds of start-up on a generic 150 MHz part with no FP64 help,
paid once (the RP2350's DCP coprocessor for double arithmetic should make
a Pico 2 cheaper than the emulated count). Knowing both
Expand Down
2 changes: 1 addition & 1 deletion book/src/part2/icount.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ accumulates a checksum, and ends with:

A total is the whole binary's cost, construction included, so a pipeline
baseline divided by its 96 000 frames is *not* the per-frame cost: on the
M33 the converter's soft-double filter design alone is over a billion
M33 the converter's soft-double filter design alone is close to a billion
instructions. `SRT_SC_SECONDS` (default 2, the only length ever baselined)
exists for that question. Build the pipeline scenarios again at
`-DSRT_SC_SECONDS=4` and difference the counts: what doubles is the
Expand Down
4 changes: 2 additions & 2 deletions book/src/part4/cortex-m.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,11 +327,11 @@ accelerates.
**Budgets, stated as instructions, pending cycles.** A baseline divided
by its 96,000 frames is *not* the per-frame cost: each workload also
constructs its converter, and on the M33 the soft-double filter design
alone is over a billion instructions. Building the pipeline workloads a
alone is close to a billion instructions. Building the pipeline workloads a
second time at 4 s of audio (`-DSRT_SC_SECONDS=4`) and taking the
difference isolates the steady state: **≈ 1,138 instructions per stereo
frame** for `pipeline_q15` and ≈ 3,326 for the 12-channel shape, with
~1.31 G and ~1.58 G of one-time construction. (Earlier revisions quoted
~0.91 G and ~1.18 G of one-time construction. (Earlier revisions quoted
5,043 and 10,027, the construction-inclusive quotients; the per-frame
costs have not changed since June.) A 150 MHz core at 48 kHz has 3,125
*cycles* per frame. The README draws the conclusion in
Expand Down
11 changes: 7 additions & 4 deletions docs/COMPARISON.md
Original file line number Diff line number Diff line change
Expand Up @@ -171,8 +171,8 @@ One-time construction, millions of instructions:

| Engine | Cortex-M55 | Cortex-M33 | Hexagon |
|---|---:|---:|---:|
| **SampleRateTap** balanced, float | 23.6 | 1,268 | 181 |
| **SampleRateTap** balanced, Q15 | 24.6 | 1,281 | 184 |
| **SampleRateTap** balanced, float | 17.3 | 870 | 133 |
| **SampleRateTap** balanced, Q15 | 18.4 | 884 | 136 |
| r8brain 120 dB, default 2 % band | 1.6 | 38.1 | 11.3 |
| r8brain 120 dB, 8 % band | 1.8 | 45.5 | 12.6 |
| libsamplerate `MEDIUM` | 1.4 | 20.9 | 7.4 |
Expand All @@ -193,8 +193,11 @@ the Cortex-M builds force-include `bench/icount/r8b_single_thread_mutex.h`
per-sample path). Hexagon's musl build uses the real mutex.

**Construction is the one column SampleRateTap loses.** Its filter design
(the compensated prototype, run in double at construction) costs ~1.3 G
instructions on the M33 as QEMU emulates it (every double operation a
(the compensated prototype, run in double at construction) costs ~0.9 G
instructions on the M33 as QEMU emulates it (~1.3 G until DspTap shared the
Kaiser window's Bessel series across the design's kernel builds; the
construction rows above are re-measured at that pin, the steady-state rows
did not move by a single instruction) (every double operation a
software libcall), against tens of millions for r8brain and
libsamplerate. On a generic FP64-less 150 MHz core that is seconds of
start-up; the RP2350 routes double arithmetic through its DCP coprocessor,
Expand Down
20 changes: 18 additions & 2 deletions docs/PERFORMANCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,13 +130,29 @@ table is already enforced by test thresholds.
algorithm-change rule; the two-sided ratchet flagged both directions
during the process (the M55 leg failed as 'IMPROVED beyond tolerance'
when the one-pass reduction landed on top of the two-pass baselines).
- [x] **Shared Kaiser-window Bessel series (DspTap #38, pin bump
28a34a1 → 0eb09fa)** — both prototype designs now evaluate the window's
`bessel_i0` series once, for half the taps, and mirror it; the
compensated design used to evaluate it per tap in each of its two
kernel builds. Bit-identical coefficients (FNV-1a-64 over the whole
prototype, on M33 and on x86 under GCC and clang), so the audio path is
untouched: the comparison workloads' steady state (4 s minus 2 s,
docs/COMPARISON.md) is identical to the instruction on all three
targets. Construction only, constant across all seven scenarios per
target to within the scenarios' own spread: M33 −393 to −398M, Hexagon
−47M, M55 −6M instructions (M33 `pipeline_q15` −27.8%, the largest
relative move). Baselines re-recorded with this entry as the
justification; the ratchet flagged every scenario 'IMPROVED beyond
tolerance' before the update. The pin also carries DspTap's FFT
stages, which SampleRateTap does not use.

## Known debt

- **Constructor cost on soft-FP64 targets, and ratchet sensitivity.** The
compensated design costs ~7M double flops at construction — pennies on
hosts and the M55 (hardware FP64), but ~1G instructions on the QEMU
M33 (soft-double libcalls): seconds of boot on an M33-class part, and
hosts and the M55 (hardware FP64), but ~0.9G instructions on the QEMU
M33 (soft-double libcalls; ~1.3G before the shared-window entry
above): seconds of boot on an M33-class part, and
a large share of each M33 icount scenario's total, which dilutes the
+/-3% gate's sensitivity to hot-path regressions on that target.
Mitigations queued: split construction into its own ratcheted scenario
Expand Down
2 changes: 1 addition & 1 deletion examples/pico2_cyccnt/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ steady-state instruction counts (`pipeline_q15` 1,138/frame,
cycles" ratio, turning every current and future M33 baseline into a real
cycle budget. The steady state is the difference of the workload built at
2 s and at 4 s (`-DSRT_SC_SECONDS=4`); a committed baseline divided by its
96,000 frames also carries the converter's one-time construction (~1.3 G
96,000 frames also carries the converter's one-time construction (~0.9 G
instructions of soft-double filter design), which this firmware does not
time. It also tests the README's claim directly: Q15 stereo fits a 150 MHz
core with room to spare, 12 channels at 48 kHz does not.
Expand Down
2 changes: 1 addition & 1 deletion examples/pico2_dualcore/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ SRT_PICO2_DUALCORE_DONE
workload built at 2 s and at 4 s, `-DSRT_SC_SECONDS=4`, and differenced).
A committed baseline divided by its 96,000 frames is larger — it also
carries one-time setup (the soft-double filter design, input synthesis,
~1.3–1.6 G instructions) — so the difference is the right counterpart of
~0.9–1.2 G instructions) — so the difference is the right counterpart of
the steady-state loop this firmware times; `cyc_frame ÷ insns/frame`
from the
sibling `examples/pico2_cyccnt` run gives the cycles-per-instruction
Expand Down
2 changes: 1 addition & 1 deletion submodules/dsptap
Submodule dsptap updated 100 files
Loading