diff --git a/README.md b/README.md index 648d102..c26973e 100644 --- a/README.md +++ b/README.md @@ -268,11 +268,17 @@ CI builds and tests every push on: M33 has no FP64 and no Helium, and the instruction baselines make the consequences concrete: the float datapath costs ~19× the M55's instructions (soft-double accumulation) — on Pico-class parts use - Q15/Q31. The instruction baselines suggest 48 kHz Q15 mono fits a - 150 MHz core and stereo wants the `fast()` preset or the RP2350's - second core — instruction counts are not cycle counts, so treat these - as budgets pending real-silicon validation: `examples/pico2_cyccnt/` - is a flashable DWT.CYCCNT harness built to measure exactly this, and + Q15/Q31. In steady state the full Q15 converter (servo and FIFO + included) costs ~1,140 instructions per stereo frame and ~3,330 at 12 + channels, against the 3,125 cycles per frame a 150 MHz core has at + 48 kHz: 48 kHz Q15 stereo fits one core with room to spare, 12 channels + does not (the dual-core example runs it at 16 kHz). Construction is + separate and heavy: ~1.3 G QEMU instructions of soft-double filter + design, seconds at boot on a generic FP64-less core (less on the RP2350, + whose DCP coprocessor handles doubles). Instruction counts are not + cycle counts, so treat these as budgets pending real-silicon + validation: `examples/pico2_cyccnt/` is a flashable DWT.CYCCNT harness + built to measure exactly this, and `examples/pico2_dualcore/` validates the one-clock-domain-per-core deployment shape. - **Arm Cortex-M55**, bare metal (newlib + semihosting, no OS/threads), diff --git a/bench/bench_asrc.cpp b/bench/bench_asrc.cpp index 6c7d13e..8f18617 100644 --- a/bench/bench_asrc.cpp +++ b/bench/bench_asrc.cpp @@ -80,7 +80,7 @@ namespace { step(); benchmark::DoNotOptimize(out.data()); } - state.SetItemsProcessed(static_cast(state.iterations()) * kBlock); + state.SetItemsProcessed(state.iterations() * static_cast(kBlock)); if (asrc.status().underruns != 0) state.SkipWithError("underrun during steady-state benchmark"); } diff --git a/bench/compare/bench_compare.cpp b/bench/compare/bench_compare.cpp index 86a4ac9..f4da5cc 100644 --- a/bench/compare/bench_compare.cpp +++ b/bench/compare/bench_compare.cpp @@ -120,7 +120,7 @@ namespace { if (got != kBlock) state.SkipWithError("source ran dry"); } - state.SetItemsProcessed(static_cast(state.iterations()) * kBlock); + state.SetItemsProcessed(state.iterations() * static_cast(kBlock)); } void lsrBench(benchmark::State& state, int converter, std::size_t channels) { diff --git a/bench/icount/icount_main.cpp b/bench/icount/icount_main.cpp index b92c810..ff1f456 100644 --- a/bench/icount/icount_main.cpp +++ b/bench/icount/icount_main.cpp @@ -55,6 +55,14 @@ namespace { #ifndef SRT_SC_CH #define SRT_SC_CH 2 +#endif + +// Pipeline length in seconds of virtual audio. The ratchet always builds the +// default; a second build at -DSRT_SC_SECONDS=4 separates steady-state cost +// from one-time construction (the difference of the two counts), the method +// docs/COMPARISON.md uses. Only the gated default is ever baselined. +#ifndef SRT_SC_SECONDS +#define SRT_SC_SECONDS 2 #endif template @@ -70,7 +78,7 @@ namespace { double sink = 0.0; std::size_t off = 0; - const std::size_t blocks = 2 * 48000 / kBlock; // 2 s of virtual audio + const std::size_t blocks = SRT_SC_SECONDS * 48000 / kBlock; // virtual audio for (std::size_t b = 0; b < blocks; ++b) { asrc.push(input.data() + off, kBlock); asrc.pull(out.data(), kBlock); diff --git a/book/src/appendix/bibliography.md b/book/src/appendix/bibliography.md index e18c14b..6d2d631 100644 --- a/book/src/appendix/bibliography.md +++ b/book/src/appendix/bibliography.md @@ -58,7 +58,7 @@ integrate the residual across the audio band for THD+N, measure dynamic range at −60 dBFS with A-weighting. The comparison notebook implements an AES17-style procedure (exact fit plus ±20 Hz notch, 20 Hz–20 kHz integration) and calibrates it against synthetic signals before use — the -standard is what makes the −132 dB figure commensurable with silicon +standard is what makes the −134 dB figure commensurable with silicon datasheets rather than a house metric. ## The measured competitors diff --git a/book/src/appendix/glossary.md b/book/src/appendix/glossary.md index a67f4f4..a215b08 100644 --- a/book/src/appendix/glossary.md +++ b/book/src/appendix/glossary.md @@ -229,7 +229,7 @@ yielding the deterministic per-workload counts the ratchet gates. **THD+N (total harmonic distortion plus noise)** — everything that is not the test signal — harmonics, spurs, noise — integrated over the audio band and expressed relative to the signal. The AES17 measurement -the comparison document reports (−132 dB at the 24-bit interface). +the comparison document reports (−134 dB at the 24-bit interface). **ThreadSanitizer (TSan)** — a compiler-instrumented data-race detector that observes the ordering annotations actually used. It certifies only diff --git a/book/src/part0/budgets.md b/book/src/part0/budgets.md index a6e18b1..c496aa4 100644 --- a/book/src/part0/budgets.md +++ b/book/src/part0/budgets.md @@ -256,16 +256,29 @@ What does an instruction budget *mean* on a 150 MHz M33? Divide. A 150 MHz core executing (optimistically) one instruction per cycle retires 150 million instructions per second, and a 48 kHz stream demands a frame every 20.8 µs — about 3,100 instructions of total budget per frame, forever, -before the rest of the firmware has run at all. Against that, the measured -comparison workloads put the full Q15 converter — servo and FIFO included -— at roughly 5,043 instructions per stereo frame on the M33: about 242 -million instructions per second for stereo, over the core's ceiling even -at ideal IPC. Mono, at roughly half that, fits. This is exactly the -README's guidance, now visible as arithmetic rather than advice: 48 kHz -Q15 mono fits a 150 MHz M33; stereo wants the `fast()` preset or the -RP2350's second core. On a Xeon the same library is a rounding error; on -the M33 the default preset is *infeasible in stereo*, and knowing that -before flashing hardware is the entire point of keeping the budget in a +before the rest of the firmware has run at all. Against that, the +steady-state cost of the full Q15 converter — servo and FIFO included — +is about 1,140 instructions per stereo frame on the M33: roughly 55 +million instructions per second, about a third of the core's ceiling at +ideal IPC. Twelve channels cost about 3,330 per frame, 160 million per +second — just over the ceiling. This is the README's guidance visible as +arithmetic rather than advice: 48 kHz Q15 stereo fits a 150 MHz M33, and +the 12-channel shape does not fit at 48 kHz on one converter instance. + +The arithmetic needs the right denominator, and an earlier edition of this +chapter used the wrong one. It quoted 5,043 instructions per stereo frame +and concluded stereo was *infeasible*. That figure divided a whole 2 s +workload by its frame count, which spreads the converter's one-time +construction (its soft-double filter design, hundreds of millions of +instructions on this core) over the audio. The fix is to measure the +workload at two lengths: the difference is the per-frame cost, the +remainder is construction. The per-frame cost was 1,138 instructions both +then and now; what changed since is construction, which grew to ~1.3 G +instructions with the compensated design. That is still a real budget +item: seconds of start-up on a generic 150 MHz part with no FP64 help, +paid once (the RP2350's DCP coprocessor for double arithmetic should make +a Pico 2 cheaper than the emulated count). Knowing both +numbers before flashing hardware is the point of keeping the budget in a table. The honesty clause matters as much as the numbers, and `docs/PERFORMANCE.md` diff --git a/book/src/part0/two-crystals.md b/book/src/part0/two-crystals.md index cb96738..8be1d17 100644 --- a/book/src/part0/two-crystals.md +++ b/book/src/part0/two-crystals.md @@ -69,7 +69,7 @@ the residual integrated across the 20 Hz–20 kHz band. The naive FIFO measures **−34.7 dB THD+N** and 94.7 dB of A-weighted dynamic range. The converter this book describes, on the same signal and -the same clocks, measures −132.1 dB. +the same clocks, measures −133.9 dB. What does −34.7 dB sound like? The number means that the error left after subtracting the test tone sits only 34.7 dB below the tone itself — a @@ -90,7 +90,7 @@ floor lies. That row of the table is the cost of doing nothing, and it calibrates everything else in this book. Every design decision in the chapters ahead -is ultimately justified by the distance between −34.7 dB and −132.1 dB. +is ultimately justified by the distance between −34.7 dB and −133.9 dB. ## The two industry answers @@ -195,10 +195,10 @@ the constructor. The computational tables in `docs/COMPARISON.md` measure what that is worth. Against libsamplerate — the closest architectural analog, a streaming time-domain polyphase resampler — at the matched ~120 dB quality -tier, SampleRateTap converts 2.9–3.6× more frames per second (mono/stereo; -2.1× at 8 channels, where both engines amortize), while carrying half the +tier, SampleRateTap converts 3.1–3.9× more frames per second (stereo/mono; +1.5× at 8 channels, where both engines amortize), while carrying half the algorithmic latency: 24 frames (0.50 ms) of filter group delay against 46 -frames (0.96 ms). At the ~140 dB tier the gap widens to 6.2× in throughput +frames (0.96 ms). At the ~140 dB tier the gap widens to 6.1× in throughput and to 40 frames against 143 in latency. That is the near-unity dividend, and the comparison document names its mechanism exactly: a 48-tap window with a creeping phase, instead of general-ratio machinery. On targets @@ -211,12 +211,15 @@ converter's one-time construction into a 2 s workload; the comparison document now reports steady state and construction separately.) The soxr rows teach a different lesson, and reading them honestly is a -preview of the next chapter. At the ~120 dB tier soxr converts 32.4 -million stereo frames per second on the same host to SampleRateTap's 10.5 +preview of the next chapter. At the ~120 dB tier soxr converts 52.9 +million stereo frames per second on the same host to SampleRateTap's 14.6 million — soxr wins raw throughput, decisively, by processing in large -SIMD-friendly internal batches. The latency column is the price: 556 to -607 frames of algorithmic delay, 11.6 to 12.6 ms, rising to 777 frames -(16.2 ms) at its highest quality tier. Those are fine numbers for batch +SIMD-friendly internal batches. The latency column is the price: 424 to +788 frames of algorithmic delay, 8.8 to 16.4 ms, and 433 frames (9.0 ms) +at its highest quality tier. r8brain-free-src tells the same story in a +different key: it out-runs SampleRateTap on the desktop by 1.1–1.7×, and +its lowest delay still flat to 20 kHz is 200 frames (4.2 ms) against +`balanced`'s 24. Those are fine numbers for batch conversion and impossible ones inside a 1–2 ms live-monitoring budget, and — as `docs/COMPARISON.md` puts it — there is no setting that buys soxr's throughput at SampleRateTap's latency. Throughput, latency, and quality @@ -226,13 +229,13 @@ allocated, and different tools have allocated it for different lives. One more number from the measured table completes the picture, because this book does not deal in free lunches. Fed by its own servo rather than an oracle, running causally at 1.5 ms of total design latency, -SampleRateTap measures −132.1 dB THD+N against the oracle-fed libraries' -−143.5 dB. The ~11 dB gap is the measured price of solving the *whole* +SampleRateTap measures −133.9 dB THD+N against the oracle-fed libraries' +−143.5 dB. The ~10 dB gap is the measured price of solving the *whole* problem — discovering the ratio from buffer occupancy in real time instead of being told it — and the comparison document presents it as exactly -that. Eleven decibels, spent 132 dB below the signal, purchasing the half +that. Ten decibels, spent 134 dB below the signal, purchasing the half of the problem that was actually hard. The rest of this book is an account -of how both numbers — the 132 and the 11 — were achieved, measured, and +of how both numbers — the 134 and the 10 — were achieved, measured, and defended. ## Watching the invisible diff --git a/book/src/part2/icount.md b/book/src/part2/icount.md index 4bc07b7..9f277ca 100644 --- a/book/src/part2/icount.md +++ b/book/src/part2/icount.md @@ -113,6 +113,14 @@ accumulates a checksum, and ends with: std::printf("SRT_ICOUNT_DONE ok=%d checksum=%.17g\n", ok ? 1 : 0, checksum); ``` +A total is the whole binary's cost, construction included, so a pipeline +baseline divided by its 96 000 frames is *not* the per-frame cost: on the +M33 the converter's soft-double filter design alone is over a billion +instructions. `SRT_SC_SECONDS` (default 2, the only length ever baselined) +exists for that question. Build the pipeline scenarios again at +`-DSRT_SC_SECONDS=4` and difference the counts: what doubles is the +per-frame steady state, and what stays is construction. + The three gated targets each run under the QEMU mode that matches their deployment reality. Hexagon binaries are Linux user-space processes, so `qemu-hexagon` (user-mode emulation) runs them directly. The two Cortex-M diff --git a/book/src/part2/notebooks.md b/book/src/part2/notebooks.md index a9441b6..1c6afb6 100644 --- a/book/src/part2/notebooks.md +++ b/book/src/part2/notebooks.md @@ -251,8 +251,8 @@ Every subject's output is measured both ways, and `docs/COMPARISON.md` leads with the 24-bit columns as the chip-comparable condition. The result reads differently than bravado would: at that interface the oracle-fed libraries measure at the 24-bit format ceiling itself (~−143.5 dB THD+N), -all three real converters share the identical 149.1 dB A-weighted -dynamic-range ceiling, and SampleRateTap's −132.1 dB sits ~11 dB behind the +every real converter shares the identical 149.1 dB A-weighted +dynamic-range ceiling, and SampleRateTap's −133.9 dB sits ~10 dB behind the oracles — a gap the document does not explain away but *prices*: it is the measured cost of solving the clock-recovery half of the problem, which the libraries do not attempt. Even so, the caveats refuse the flattering frame diff --git a/book/src/part4/cortex-m.md b/book/src/part4/cortex-m.md index 9701be3..0d97d71 100644 --- a/book/src/part4/cortex-m.md +++ b/book/src/part4/cortex-m.md @@ -324,13 +324,20 @@ bounded: the M33's Q15 frame cost is dominated by the coefficient blend's 64-bit products and transport, not by the dot product the intrinsic accelerates. -**Budgets, stated as instructions, pending cycles.** Dividing the -baselines out: `pipeline_q15` is 484,146,844 instructions per 96,000 -frames ≈ **5,043 instructions per stereo frame**; the 12-channel shape is -≈ 10,027. A 150 MHz core at 48 kHz has 3,125 *cycles* per frame. The -README draws the honest conclusion in instruction-space — Q15 mono fits -a 150 MHz core, stereo wants the `fast()` preset or the RP2350's second -core — and then refuses to pretend the units match: instructions are not +**Budgets, stated as instructions, pending cycles.** A baseline divided +by its 96,000 frames is *not* the per-frame cost: each workload also +constructs its converter, and on the M33 the soft-double filter design +alone is over a billion instructions. Building the pipeline workloads a +second time at 4 s of audio (`-DSRT_SC_SECONDS=4`) and taking the +difference isolates the steady state: **≈ 1,138 instructions per stereo +frame** for `pipeline_q15` and ≈ 3,326 for the 12-channel shape, with +~1.31 G and ~1.58 G of one-time construction. (Earlier revisions quoted +5,043 and 10,027, the construction-inclusive quotients; the per-frame +costs have not changed since June.) A 150 MHz core at 48 kHz has 3,125 +*cycles* per frame. The README draws the conclusion in +instruction-space — Q15 stereo fits a 150 MHz core, the 12-channel shape +does not at 48 kHz — and then refuses to pretend the units match: +instructions are not cycles, the ratio between them is an empirical property of real silicon, and the guidance is explicitly a budget *pending real-silicon validation*. @@ -340,8 +347,8 @@ the bridge from this chapter's emulated world to Part V's hardware: - **`examples/pico2_cyccnt`** runs the same fixed pipeline workloads on a real Pico 2 and times each 32-frame block with the M33's DWT.CYCCNT - hardware cycle counter. Its output divided by the committed baselines - (5,043 and 10,027 instructions per frame) yields the + hardware cycle counter. Its output divided by the steady-state + instruction counts (1,138 and 3,326 per frame) yields the cycles-per-QEMU-instruction calibration constant that turns *every* M33 baseline, current and future, into a real cycle budget. - **`examples/pico2_dualcore`** is the "second core" clause made @@ -358,9 +365,10 @@ the bridge from this chapter's emulated world to Part V's hardware: M33, 64-bit `std::atomic` is not lock-free, the same fact the startup file's PRIMASK helpers exist to paper over on *one* core and which no single-core trick can fix across two. Even the firmware's 12-channel - phase runs at 16 kHz *by arithmetic, not caution*: 10,027 + phase runs at 16 kHz *by arithmetic, not caution*: 3,326 instructions per frame against a 3,125-cycle budget cannot fit at - 48 kHz on one core, and `pull()` of one converter instance is one + 48 kHz on one core even at one instruction per cycle, and `pull()` of + one converter instance is one consumer by contract — a second core buys one clock domain per core, not more datapath than one core has. diff --git a/book/src/part5/hardware.md b/book/src/part5/hardware.md index 8852306..4357d68 100644 --- a/book/src/part5/hardware.md +++ b/book/src/part5/hardware.md @@ -232,9 +232,12 @@ is the wrong datapath on an FP64-less core" — the QEMU baselines already price float at roughly 3.8× the Q15 instruction count, and a cycle figure makes the guidance concrete rather than rhetorical. -The deeper purpose is calibration. The committed M33 baselines divide out -to 5,043 instructions per frame for the stereo Q15 pipeline and 10,027 for -the 12-channel one. Divide the firmware's measured cycles-per-frame by +The deeper purpose is calibration. In steady state the M33 pipelines cost +1,138 instructions per frame for the stereo Q15 pipeline and 3,326 for the +12-channel one (the difference of 2 s and 4 s workloads; a baseline +divided by its frame count also carries the converter's construction, and +an earlier edition used those larger quotients here). Divide the +firmware's measured cycles-per-frame by those figures and you get the constant the whole ratchet has been waiting for: *one QEMU instruction ≈ N RP2350 cycles*. That single ratio converts every current and future M33 instruction baseline into a real cycle @@ -329,11 +332,12 @@ Phase B is the 12-channel reference-microphone/AVB shape... at **16 kHz**, not 48. Its README records why, and the passage is a model of how to scope a demo honestly: -> Phase B is 16 kHz **by arithmetic, not caution**: the M33 QEMU baseline -> puts `pipeline12_q15` at 10,027 insns/frame against a 150 MHz / 48 kHz -> budget of 3,125 cycles/frame — more than 3× over, and `pull()` of a -> single instance is one consumer by contract, so no core assignment can -> split it across cores. Dual-core buys one clock domain per core, not +> Phase B is 16 kHz **by arithmetic, not caution**: in steady state the +> M33 QEMU count puts `pipeline12_q15` at 3,326 insns/frame against a +> 150 MHz / 48 kHz budget of 3,125 cycles/frame — over budget even at one +> instruction per cycle, and `pull()` of a single instance is one consumer +> by contract, so no core assignment can split it across cores. +> Dual-core buys one clock domain per core, not > more datapath than one core has. That last sentence is the chapter's most important deployment fact. The diff --git a/docs/COMPARISON.md b/docs/COMPARISON.md index 603b1ea..8b3cab6 100644 --- a/docs/COMPARISON.md +++ b/docs/COMPARISON.md @@ -194,10 +194,13 @@ per-sample path). Hexagon's musl build uses the real mutex. **Construction is the one column SampleRateTap loses.** Its filter design (the compensated prototype, run in double at construction) costs ~1.3 G -instructions on the M33, where double is emulated — seconds of start-up on -a 150 MHz part (instructions are not cycles), against tens of millions for -r8brain and libsamplerate. It is paid once per converter, never on the -audio path, but it is a real cost for devices that construct at boot. +instructions on the M33 as QEMU emulates it (every double operation a +software libcall), against tens of millions for r8brain and +libsamplerate. On a generic FP64-less 150 MHz core that is seconds of +start-up; the RP2350 routes double arithmetic through its DCP coprocessor, +so a Pico 2 should pay less than the count suggests (`pico2_cyccnt` can +measure it). It is paid once per converter, never on the audio path, but +it is a real cost for devices that construct at boot. ## The landscape diff --git a/docs/HARDWARE_TESTING.md b/docs/HARDWARE_TESTING.md index 1bab42e..92b3ae1 100644 --- a/docs/HARDWARE_TESTING.md +++ b/docs/HARDWARE_TESTING.md @@ -69,9 +69,10 @@ Two things this proves that emulation cannot: gives deterministic *instruction* counts, not cycles, and real cycles need hardware counters. The RP2350 has DWT.CYCCNT: wrapping `pull()` in CYCCNT reads gives real cycles-per-block at 150 MHz — directly - testing the README's claim that Q15 mono fits comfortably and stereo - is tight on one core. Correlating CYCCNT against the QEMU instruction - baselines also calibrates the ratchet ("1 QEMU instruction ≈ N RP2350 + testing the README's claim that Q15 stereo fits comfortably on one + core (~1,140 steady-state instructions per frame against 3,125 cycles). + Correlating CYCCNT against the QEMU steady-state instruction counts + also calibrates the ratchet ("1 QEMU instruction ≈ N RP2350 cycles") for all future M33 numbers. - **Dual-core deployment** (harness shipped: [`examples/pico2_dualcore/`](../examples/pico2_dualcore/), self-validating diff --git a/examples/pico2_cyccnt/README.md b/examples/pico2_cyccnt/README.md index 5dc54c5..e7dcec9 100644 --- a/examples/pico2_cyccnt/README.md +++ b/examples/pico2_cyccnt/README.md @@ -11,13 +11,16 @@ hardware cycle counter. [docs/PERFORMANCE.md](../../docs/PERFORMANCE.md) gates regressions on QEMU *instruction* counts because they are deterministic and noise-free — but silicon budgets are spent in *cycles*, and QEMU cannot provide those. This -firmware closes that loop: dividing its measured cycles/frame by the -committed M33 instruction baselines (`pipeline_q15` 484,146,844 insns per -96,000 frames = 5,043/frame; `pipeline12_q15` 962,613,655 = 10,027/frame) -calibrates the "1 QEMU instruction ≈ N RP2350 cycles" ratio, turning every -current and future M33 baseline into a real cycle budget. It also tests the -README's claim directly: Q15 mono fits a 150 MHz core with room to spare, -stereo is tighter. +firmware closes that loop: dividing its measured cycles/frame by the M33 +steady-state instruction counts (`pipeline_q15` 1,138/frame, +`pipeline12_q15` 3,326/frame) calibrates the "1 QEMU instruction ≈ N RP2350 +cycles" ratio, turning every current and future M33 baseline into a real +cycle budget. The steady state is the difference of the workload built at +2 s and at 4 s (`-DSRT_SC_SECONDS=4`); a committed baseline divided by its +96,000 frames also carries the converter's one-time construction (~1.3 G +instructions of soft-double filter design), which this firmware does not +time. It also tests the README's claim directly: Q15 stereo fits a 150 MHz +core with room to spare, 12 channels at 48 kHz does not. ## Build @@ -75,7 +78,7 @@ SRT_PICO2_DONE ## Reading the numbers -- **cyc/frame ÷ 5,043** (Q15 balanced 2ch) and **÷ 10,027** (12ch) give the +- **cyc/frame ÷ 1,138** (Q15 balanced 2ch) and **÷ 3,326** (12ch) give the silicon cycles-per-QEMU-instruction ratio — the calibration constant for all M33 instruction baselines in the README table. - **%core@48k** is the headline budget figure. The float rows exist to put a diff --git a/examples/pico2_cyccnt/main.cpp b/examples/pico2_cyccnt/main.cpp index 843df66..7b18855 100644 --- a/examples/pico2_cyccnt/main.cpp +++ b/examples/pico2_cyccnt/main.cpp @@ -6,11 +6,12 @@ // Calibration purpose: docs/PERFORMANCE.md gates regressions on QEMU // *instruction* counts because they are deterministic; real cost is in // *cycles*, which only hardware counters give. Dividing the mean cycles/frame -// printed here by the committed M33 QEMU baselines (bench/baselines.json, -// 2 s of 48 kHz audio = 96,000 frames per workload): +// printed here by the M33 QEMU steady-state instruction counts (the workload +// built at 2 s and at 4 s, -DSRT_SC_SECONDS=4, differenced so the one-time +// construction a baseline also carries drops out; measured 2026-09-26): // -// pipeline_q15 (2ch, balanced) 484,146,844 insns = 5,043 insns/frame -// pipeline12_q15 (12ch, balanced) 962,613,655 insns = 10,027 insns/frame +// pipeline_q15 (2ch, balanced) 1,138 insns/frame +// pipeline12_q15 (12ch, balanced) 3,326 insns/frame // // yields the "1 QEMU instruction ~= N RP2350 cycles" ratio that converts // every M33 instruction baseline into a real cycle budget. diff --git a/examples/pico2_dualcore/README.md b/examples/pico2_dualcore/README.md index 49257b9..d4391ee 100644 --- a/examples/pico2_dualcore/README.md +++ b/examples/pico2_dualcore/README.md @@ -36,13 +36,14 @@ Two phases, ~30 s each: | Phase | config | Rates | Why | |---|---|---|---| -| A | Q15 stereo `balanced()` | 48 kHz out, +200 ppm in | the config the README calls tight on one core | +| A | Q15 stereo `balanced()` | 48 kHz out, +200 ppm in | the README's stereo config (~1,140 insns/frame steady state, about a third of one core) | | B | Q15 12-channel, `balanced()` band edges and servo scaled ×16/48 | 16 kHz out, +200 ppm in | the reference-microphone/AVB 12-channel shape at its deployment rate | -Phase B is 16 kHz **by arithmetic, not caution**: the M33 QEMU baseline puts -`pipeline12_q15` at 10,027 insns/frame against a 150 MHz / 48 kHz budget of -3,125 cycles/frame — more than 3× over, and `pull()` of a single instance is -one consumer by contract, so no core assignment can split it across cores. +Phase B is 16 kHz **by arithmetic, not caution**: in steady state the M33 +QEMU count puts `pipeline12_q15` at 3,326 insns/frame against a 150 MHz / +48 kHz budget of 3,125 cycles/frame — over budget even at one instruction per +cycle, and `pull()` of a single instance is one consumer by contract, so no +core assignment can split it across cores. Dual-core buys one clock domain per core, not more datapath than one core has. At 16 kHz the budget is 9,375 cycles/frame. The measured cycles/block is rate-independent, so phase B still produces the real-silicon counterpart @@ -105,17 +106,20 @@ SRT_PICO2_DUALCORE_DONE reported sys clock, `cyc_frame × rate / 150 MHz`. It prices `pull()` only — by design, since `push()` is a ring write and the producer core's real budget goes to whatever feeds it (here: telemetry). -- **Relation to the QEMU baselines** (`bench/baselines.json`, 2 s = 96,000 - frames per workload): `pipeline_q15` 484,146,844 insns = **5,043 - insns/frame**, `pipeline12_q15` 962,613,655 = **10,027 insns/frame**. - Those figures amortize one-time setup (soft-double Kaiser design, input - synthesis) over the workload, so they are upper bounds for the - steady-state loop this firmware times; `cyc_frame ÷ insns/frame` from the +- **Relation to the QEMU counts**: in steady state `pipeline_q15` costs + **1,138 insns/frame** and `pipeline12_q15` **3,326 insns/frame** (each + workload built at 2 s and at 4 s, `-DSRT_SC_SECONDS=4`, and differenced). + A committed baseline divided by its 96,000 frames is larger — it also + carries one-time setup (the soft-double filter design, input synthesis, + ~1.3–1.6 G instructions) — so the difference is the right counterpart of + the steady-state loop this firmware times; `cyc_frame ÷ insns/frame` + from the sibling `examples/pico2_cyccnt` run gives the cycles-per-instruction calibration that converts every M33 baseline into a cycle budget. - Against the budgets: 3,125 cycles/frame buys one 48 kHz frame at 150 MHz, 9,375 one 16 kHz frame. Phase A's `core_pct` is the measured - version of "stereo balanced() is tight on one core": whatever it reads, + version of "stereo balanced() fits one core with room to spare": + whatever it reads, that is the share of core1 a deployment must reserve — and on a *single*-core deployment the same cycles would contend with the producer side and the rest of the application, which is exactly why the input diff --git a/examples/pico2_dualcore/main.cpp b/examples/pico2_dualcore/main.cpp index f3c4dd3..0a098b0 100644 --- a/examples/pico2_dualcore/main.cpp +++ b/examples/pico2_dualcore/main.cpp @@ -509,9 +509,10 @@ int main() { // A: 32/48000 s = 2000/3 us; B: 32/16000 s = 2000/1 us // // Phase B pins the 12-channel shape at 16 kHz — the README's - // reference-microphone/AVB deployment rate — not 48 kHz: the M33 QEMU - // baseline puts pipeline12_q15 at 10,027 insns/frame against a - // 150 MHz / 48 kHz budget of 3,125 cycles/frame, more than 3x over, and + // reference-microphone/AVB deployment rate — not 48 kHz: in steady state + // the M33 QEMU count puts pipeline12_q15 at 3,326 insns/frame against a + // 150 MHz / 48 kHz budget of 3,125 cycles/frame, over even at one + // instruction per cycle, and // pull() of one instance is a single consumer by contract — no core // assignment can split it. At 16 kHz the budget is 9,375 cycles/frame. // The measured cycles/block is rate-independent either way, so phase B