Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 11 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -268,11 +268,17 @@ CI builds and tests every push on:
M33 has no FP64 and no Helium, and the instruction baselines make the
consequences concrete: the float datapath costs ~19× the M55's
instructions (soft-double accumulation) — on Pico-class parts use
Q15/Q31. The instruction baselines suggest 48 kHz Q15 mono fits a
150 MHz core and stereo wants the `fast()` preset or the RP2350's
second core — instruction counts are not cycle counts, so treat these
as budgets pending real-silicon validation: `examples/pico2_cyccnt/`
is a flashable DWT.CYCCNT harness built to measure exactly this, and
Q15/Q31. In steady state the full Q15 converter (servo and FIFO
included) costs ~1,140 instructions per stereo frame and ~3,330 at 12
channels, against the 3,125 cycles per frame a 150 MHz core has at
48 kHz: 48 kHz Q15 stereo fits one core with room to spare, 12 channels
does not (the dual-core example runs it at 16 kHz). Construction is
separate and heavy: ~1.3 G QEMU instructions of soft-double filter
design, seconds at boot on a generic FP64-less core (less on the RP2350,
whose DCP coprocessor handles doubles). Instruction counts are not
cycle counts, so treat these as budgets pending real-silicon
validation: `examples/pico2_cyccnt/` is a flashable DWT.CYCCNT harness
built to measure exactly this, and
`examples/pico2_dualcore/` validates the one-clock-domain-per-core
deployment shape.
- **Arm Cortex-M55**, bare metal (newlib + semihosting, no OS/threads),
Expand Down
2 changes: 1 addition & 1 deletion bench/bench_asrc.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@ namespace {
step();
benchmark::DoNotOptimize(out.data());
}
state.SetItemsProcessed(static_cast<std::int64_t>(state.iterations()) * kBlock);
state.SetItemsProcessed(state.iterations() * static_cast<std::int64_t>(kBlock));
if (asrc.status().underruns != 0)
state.SkipWithError("underrun during steady-state benchmark");
}
Expand Down
2 changes: 1 addition & 1 deletion bench/compare/bench_compare.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -120,7 +120,7 @@ namespace {
if (got != kBlock)
state.SkipWithError("source ran dry");
}
state.SetItemsProcessed(static_cast<std::int64_t>(state.iterations()) * kBlock);
state.SetItemsProcessed(state.iterations() * static_cast<std::int64_t>(kBlock));
}

void lsrBench(benchmark::State& state, int converter, std::size_t channels) {
Expand Down
10 changes: 9 additions & 1 deletion bench/icount/icount_main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,14 @@ namespace {

#ifndef SRT_SC_CH
#define SRT_SC_CH 2
#endif

// Pipeline length in seconds of virtual audio. The ratchet always builds the
// default; a second build at -DSRT_SC_SECONDS=4 separates steady-state cost
// from one-time construction (the difference of the two counts), the method
// docs/COMPARISON.md uses. Only the gated default is ever baselined.
#ifndef SRT_SC_SECONDS
#define SRT_SC_SECONDS 2
#endif

template <typename S>
Expand All @@ -70,7 +78,7 @@ namespace {

double sink = 0.0;
std::size_t off = 0;
const std::size_t blocks = 2 * 48000 / kBlock; // 2 s of virtual audio
const std::size_t blocks = SRT_SC_SECONDS * 48000 / kBlock; // virtual audio
for (std::size_t b = 0; b < blocks; ++b) {
asrc.push(input.data() + off, kBlock);
asrc.pull(out.data(), kBlock);
Expand Down
2 changes: 1 addition & 1 deletion book/src/appendix/bibliography.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ integrate the residual across the audio band for THD+N, measure dynamic
range at −60 dBFS with A-weighting. The comparison notebook implements an
AES17-style procedure (exact fit plus ±20 Hz notch, 20 Hz–20 kHz
integration) and calibrates it against synthetic signals before use — the
standard is what makes the −132 dB figure commensurable with silicon
standard is what makes the −134 dB figure commensurable with silicon
datasheets rather than a house metric.

## The measured competitors
Expand Down
2 changes: 1 addition & 1 deletion book/src/appendix/glossary.md
Original file line number Diff line number Diff line change
Expand Up @@ -229,7 +229,7 @@ yielding the deterministic per-workload counts the ratchet gates.
**THD+N (total harmonic distortion plus noise)** — everything that is
not the test signal — harmonics, spurs, noise — integrated over the
audio band and expressed relative to the signal. The AES17 measurement
the comparison document reports (−132 dB at the 24-bit interface).
the comparison document reports (−134 dB at the 24-bit interface).

**ThreadSanitizer (TSan)** — a compiler-instrumented data-race detector
that observes the ordering annotations actually used. It certifies only
Expand Down
33 changes: 23 additions & 10 deletions book/src/part0/budgets.md
Original file line number Diff line number Diff line change
Expand Up @@ -256,16 +256,29 @@ What does an instruction budget *mean* on a 150 MHz M33? Divide. A 150 MHz
core executing (optimistically) one instruction per cycle retires 150
million instructions per second, and a 48 kHz stream demands a frame every
20.8 µs — about 3,100 instructions of total budget per frame, forever,
before the rest of the firmware has run at all. Against that, the measured
comparison workloads put the full Q15 converter — servo and FIFO included
— at roughly 5,043 instructions per stereo frame on the M33: about 242
million instructions per second for stereo, over the core's ceiling even
at ideal IPC. Mono, at roughly half that, fits. This is exactly the
README's guidance, now visible as arithmetic rather than advice: 48 kHz
Q15 mono fits a 150 MHz M33; stereo wants the `fast()` preset or the
RP2350's second core. On a Xeon the same library is a rounding error; on
the M33 the default preset is *infeasible in stereo*, and knowing that
before flashing hardware is the entire point of keeping the budget in a
before the rest of the firmware has run at all. Against that, the
steady-state cost of the full Q15 converter — servo and FIFO included —
is about 1,140 instructions per stereo frame on the M33: roughly 55
million instructions per second, about a third of the core's ceiling at
ideal IPC. Twelve channels cost about 3,330 per frame, 160 million per
second — just over the ceiling. This is the README's guidance visible as
arithmetic rather than advice: 48 kHz Q15 stereo fits a 150 MHz M33, and
the 12-channel shape does not fit at 48 kHz on one converter instance.

The arithmetic needs the right denominator, and an earlier edition of this
chapter used the wrong one. It quoted 5,043 instructions per stereo frame
and concluded stereo was *infeasible*. That figure divided a whole 2 s
workload by its frame count, which spreads the converter's one-time
construction (its soft-double filter design, hundreds of millions of
instructions on this core) over the audio. The fix is to measure the
workload at two lengths: the difference is the per-frame cost, the
remainder is construction. The per-frame cost was 1,138 instructions both
then and now; what changed since is construction, which grew to ~1.3 G
instructions with the compensated design. That is still a real budget
item: seconds of start-up on a generic 150 MHz part with no FP64 help,
paid once (the RP2350's DCP coprocessor for double arithmetic should make
a Pico 2 cheaper than the emulated count). Knowing both
numbers before flashing hardware is the point of keeping the budget in a
table.

The honesty clause matters as much as the numbers, and `docs/PERFORMANCE.md`
Expand Down
31 changes: 17 additions & 14 deletions book/src/part0/two-crystals.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,7 @@ the residual integrated across the 20 Hz–20 kHz band.

The naive FIFO measures **−34.7 dB THD+N** and 94.7 dB of A-weighted
dynamic range. The converter this book describes, on the same signal and
the same clocks, measures −132.1 dB.
the same clocks, measures −133.9 dB.

What does −34.7 dB sound like? The number means that the error left after
subtracting the test tone sits only 34.7 dB below the tone itself — a
Expand All @@ -90,7 +90,7 @@ floor lies.

That row of the table is the cost of doing nothing, and it calibrates
everything else in this book. Every design decision in the chapters ahead
is ultimately justified by the distance between −34.7 dB and −132.1 dB.
is ultimately justified by the distance between −34.7 dB and −133.9 dB.

## The two industry answers

Expand Down Expand Up @@ -195,10 +195,10 @@ the constructor.
The computational tables in `docs/COMPARISON.md` measure what that is
worth. Against libsamplerate — the closest architectural analog, a
streaming time-domain polyphase resampler — at the matched ~120 dB quality
tier, SampleRateTap converts 2.9–3.6× more frames per second (mono/stereo;
2.1× at 8 channels, where both engines amortize), while carrying half the
tier, SampleRateTap converts 3.1–3.9× more frames per second (stereo/mono;
1.5× at 8 channels, where both engines amortize), while carrying half the
algorithmic latency: 24 frames (0.50 ms) of filter group delay against 46
frames (0.96 ms). At the ~140 dB tier the gap widens to 6.2× in throughput
frames (0.96 ms). At the ~140 dB tier the gap widens to 6.1× in throughput
and to 40 frames against 143 in latency. That is the near-unity dividend,
and the comparison document names its mechanism exactly: a 48-tap window
with a creeping phase, instead of general-ratio machinery. On targets
Expand All @@ -211,12 +211,15 @@ converter's one-time construction into a 2 s workload; the comparison
document now reports steady state and construction separately.)

The soxr rows teach a different lesson, and reading them honestly is a
preview of the next chapter. At the ~120 dB tier soxr converts 32.4
million stereo frames per second on the same host to SampleRateTap's 10.5
preview of the next chapter. At the ~120 dB tier soxr converts 52.9
million stereo frames per second on the same host to SampleRateTap's 14.6
million — soxr wins raw throughput, decisively, by processing in large
SIMD-friendly internal batches. The latency column is the price: 556 to
607 frames of algorithmic delay, 11.6 to 12.6 ms, rising to 777 frames
(16.2 ms) at its highest quality tier. Those are fine numbers for batch
SIMD-friendly internal batches. The latency column is the price: 424 to
788 frames of algorithmic delay, 8.8 to 16.4 ms, and 433 frames (9.0 ms)
at its highest quality tier. r8brain-free-src tells the same story in a
different key: it out-runs SampleRateTap on the desktop by 1.1–1.7×, and
its lowest delay still flat to 20 kHz is 200 frames (4.2 ms) against
`balanced`'s 24. Those are fine numbers for batch
conversion and impossible ones inside a 1–2 ms live-monitoring budget, and
— as `docs/COMPARISON.md` puts it — there is no setting that buys soxr's
throughput at SampleRateTap's latency. Throughput, latency, and quality
Expand All @@ -226,13 +229,13 @@ allocated, and different tools have allocated it for different lives.
One more number from the measured table completes the picture, because
this book does not deal in free lunches. Fed by its own servo rather than
an oracle, running causally at 1.5 ms of total design latency,
SampleRateTap measures −132.1 dB THD+N against the oracle-fed libraries'
−143.5 dB. The ~11 dB gap is the measured price of solving the *whole*
SampleRateTap measures −133.9 dB THD+N against the oracle-fed libraries'
−143.5 dB. The ~10 dB gap is the measured price of solving the *whole*
problem — discovering the ratio from buffer occupancy in real time instead
of being told it — and the comparison document presents it as exactly
that. Eleven decibels, spent 132 dB below the signal, purchasing the half
that. Ten decibels, spent 134 dB below the signal, purchasing the half
of the problem that was actually hard. The rest of this book is an account
of how both numbers — the 132 and the 11 — were achieved, measured, and
of how both numbers — the 134 and the 10 — were achieved, measured, and
defended.

## Watching the invisible
Expand Down
8 changes: 8 additions & 0 deletions book/src/part2/icount.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,14 @@ accumulates a checksum, and ends with:
std::printf("SRT_ICOUNT_DONE ok=%d checksum=%.17g\n", ok ? 1 : 0, checksum);
```

A total is the whole binary's cost, construction included, so a pipeline
baseline divided by its 96 000 frames is *not* the per-frame cost: on the
M33 the converter's soft-double filter design alone is over a billion
instructions. `SRT_SC_SECONDS` (default 2, the only length ever baselined)
exists for that question. Build the pipeline scenarios again at
`-DSRT_SC_SECONDS=4` and difference the counts: what doubles is the
per-frame steady state, and what stays is construction.

The three gated targets each run under the QEMU mode that matches their
deployment reality. Hexagon binaries are Linux user-space processes, so
`qemu-hexagon` (user-mode emulation) runs them directly. The two Cortex-M
Expand Down
4 changes: 2 additions & 2 deletions book/src/part2/notebooks.md
Original file line number Diff line number Diff line change
Expand Up @@ -251,8 +251,8 @@ Every subject's output is measured both ways, and `docs/COMPARISON.md`
leads with the 24-bit columns as the chip-comparable condition. The result
reads differently than bravado would: at that interface the oracle-fed
libraries measure at the 24-bit format ceiling itself (~−143.5 dB THD+N),
all three real converters share the identical 149.1 dB A-weighted
dynamic-range ceiling, and SampleRateTap's −132.1 dB sits ~11 dB behind the
every real converter shares the identical 149.1 dB A-weighted
dynamic-range ceiling, and SampleRateTap's −133.9 dB sits ~10 dB behind the
oracles — a gap the document does not explain away but *prices*: it is the
measured cost of solving the clock-recovery half of the problem, which the
libraries do not attempt. Even so, the caveats refuse the flattering frame
Expand Down
30 changes: 19 additions & 11 deletions book/src/part4/cortex-m.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,13 +324,20 @@ bounded: the M33's Q15 frame cost is dominated by the coefficient blend's
64-bit products and transport, not by the dot product the intrinsic
accelerates.

**Budgets, stated as instructions, pending cycles.** Dividing the
baselines out: `pipeline_q15` is 484,146,844 instructions per 96,000
frames ≈ **5,043 instructions per stereo frame**; the 12-channel shape is
≈ 10,027. A 150 MHz core at 48 kHz has 3,125 *cycles* per frame. The
README draws the honest conclusion in instruction-space — Q15 mono fits
a 150 MHz core, stereo wants the `fast()` preset or the RP2350's second
core — and then refuses to pretend the units match: instructions are not
**Budgets, stated as instructions, pending cycles.** A baseline divided
by its 96,000 frames is *not* the per-frame cost: each workload also
constructs its converter, and on the M33 the soft-double filter design
alone is over a billion instructions. Building the pipeline workloads a
second time at 4 s of audio (`-DSRT_SC_SECONDS=4`) and taking the
difference isolates the steady state: **≈ 1,138 instructions per stereo
frame** for `pipeline_q15` and ≈ 3,326 for the 12-channel shape, with
~1.31 G and ~1.58 G of one-time construction. (Earlier revisions quoted
5,043 and 10,027, the construction-inclusive quotients; the per-frame
costs have not changed since June.) A 150 MHz core at 48 kHz has 3,125
*cycles* per frame. The README draws the conclusion in
instruction-space — Q15 stereo fits a 150 MHz core, the 12-channel shape
does not at 48 kHz — and then refuses to pretend the units match:
instructions are not
cycles, the ratio between them is an empirical property of real silicon,
and the guidance is explicitly a budget *pending real-silicon
validation*.
Expand All @@ -340,8 +347,8 @@ the bridge from this chapter's emulated world to Part V's hardware:

- **`examples/pico2_cyccnt`** runs the same fixed pipeline workloads on a
real Pico 2 and times each 32-frame block with the M33's DWT.CYCCNT
hardware cycle counter. Its output divided by the committed baselines
(5,043 and 10,027 instructions per frame) yields the
hardware cycle counter. Its output divided by the steady-state
instruction counts (1,138 and 3,326 per frame) yields the
cycles-per-QEMU-instruction calibration constant that turns *every*
M33 baseline, current and future, into a real cycle budget.
- **`examples/pico2_dualcore`** is the "second core" clause made
Expand All @@ -358,9 +365,10 @@ the bridge from this chapter's emulated world to Part V's hardware:
M33, 64-bit `std::atomic` is not lock-free, the same fact the startup
file's PRIMASK helpers exist to paper over on *one* core and which no
single-core trick can fix across two. Even the firmware's 12-channel
phase runs at 16 kHz *by arithmetic, not caution*: 10,027
phase runs at 16 kHz *by arithmetic, not caution*: 3,326
instructions per frame against a 3,125-cycle budget cannot fit at
48 kHz on one core, and `pull()` of one converter instance is one
48 kHz on one core even at one instruction per cycle, and `pull()` of
one converter instance is one
consumer by contract — a second core buys one clock domain per core,
not more datapath than one core has.

Expand Down
20 changes: 12 additions & 8 deletions book/src/part5/hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,9 +232,12 @@ is the wrong datapath on an FP64-less core" — the QEMU baselines already
price float at roughly 3.8× the Q15 instruction count, and a cycle figure
makes the guidance concrete rather than rhetorical.

The deeper purpose is calibration. The committed M33 baselines divide out
to 5,043 instructions per frame for the stereo Q15 pipeline and 10,027 for
the 12-channel one. Divide the firmware's measured cycles-per-frame by
The deeper purpose is calibration. In steady state the M33 pipelines cost
1,138 instructions per frame for the stereo Q15 pipeline and 3,326 for the
12-channel one (the difference of 2 s and 4 s workloads; a baseline
divided by its frame count also carries the converter's construction, and
an earlier edition used those larger quotients here). Divide the
firmware's measured cycles-per-frame by
those figures and you get the constant the whole ratchet has been waiting
for: *one QEMU instruction ≈ N RP2350 cycles*. That single ratio converts
every current and future M33 instruction baseline into a real cycle
Expand Down Expand Up @@ -329,11 +332,12 @@ Phase B is the 12-channel reference-microphone/AVB shape... at **16 kHz**,
not 48. Its README records why, and the passage is a model of how to scope
a demo honestly:

> Phase B is 16 kHz **by arithmetic, not caution**: the M33 QEMU baseline
> puts `pipeline12_q15` at 10,027 insns/frame against a 150 MHz / 48 kHz
> budget of 3,125 cycles/frame — more than 3× over, and `pull()` of a
> single instance is one consumer by contract, so no core assignment can
> split it across cores. Dual-core buys one clock domain per core, not
> Phase B is 16 kHz **by arithmetic, not caution**: in steady state the
> M33 QEMU count puts `pipeline12_q15` at 3,326 insns/frame against a
> 150 MHz / 48 kHz budget of 3,125 cycles/frame — over budget even at one
> instruction per cycle, and `pull()` of a single instance is one consumer
> by contract, so no core assignment can split it across cores.
> Dual-core buys one clock domain per core, not
> more datapath than one core has.

That last sentence is the chapter's most important deployment fact. The
Expand Down
11 changes: 7 additions & 4 deletions docs/COMPARISON.md
Original file line number Diff line number Diff line change
Expand Up @@ -194,10 +194,13 @@ per-sample path). Hexagon's musl build uses the real mutex.

**Construction is the one column SampleRateTap loses.** Its filter design
(the compensated prototype, run in double at construction) costs ~1.3 G
instructions on the M33, where double is emulated — seconds of start-up on
a 150 MHz part (instructions are not cycles), against tens of millions for
r8brain and libsamplerate. It is paid once per converter, never on the
audio path, but it is a real cost for devices that construct at boot.
instructions on the M33 as QEMU emulates it (every double operation a
software libcall), against tens of millions for r8brain and
libsamplerate. On a generic FP64-less 150 MHz core that is seconds of
start-up; the RP2350 routes double arithmetic through its DCP coprocessor,
so a Pico 2 should pay less than the count suggests (`pico2_cyccnt` can
measure it). It is paid once per converter, never on the audio path, but
it is a real cost for devices that construct at boot.

## The landscape

Expand Down
Loading
Loading