Skip to content

Add r8brain-free-src to the resampler comparison - #45

Merged
tap merged 2 commits into
mainfrom
claude/sample-rate-solutions-comparison-mqc190
Sep 26, 2026
Merged

tap merged 2 commits into
mainfrom
claude/sample-rate-solutions-comparison-mqc190

Conversation

@tap

@tap tap commented Sep 26, 2026

Copy link
Copy Markdown
Owner

What this changes

Adds r8brain-free-src (7.5, pinned at commit 9e73d2dd in cmake/r8brain.cmake) to every part of docs/COMPARISON.md: the quality notebook, the host throughput bench and the embedded instruction counts. The embedded counts now report steady-state cost separately from one-time construction.

Why

We had no measured comparison against r8brain, a widely used MIT resampler. Re-measuring also turned up stale numbers in the comparison. The old embedded per-frame figures included one-time construction, so the headline "~9.8× cheaper than libsamplerate on M33" no longer matched the library.

Verification

All of this was built and run in one container on 2026-09-25/26:

  • Quality: notebooks/asrc_comparison.ipynb is re-executed and committed executed, and its assertions pass.
    • r8brain CDSPResampler24 measures −143.9 dB THD+N at 24-bit I/O, −150.8 dB at float I/O, and 149.1 dB DR.
    • The new "Latency vs. passband" cell sweeps r8brain's transition band. At the default it holds back 789 input frames. The lowest-latency setting still flat to 20 kHz is 8 % at 200 frames, and the ~1 ms settings are −35 dB at 20 kHz.
  • Host throughput: srt_bench_compare, median of 5, all engines in one session. The numbers are in COMPARISON.md.
  • Embedded counts: QEMU plus the counting plugin, all 36 runs with ok=1, on M55 (mps3-an547), M33 (mps2-an505) and Hexagon (qemu-hexagon built with plugins from the checksum-pinned 8.2.2 source, toolchain 19.1.5). The toolchain and plugin-header pins match compare.yml.
    • The libsamplerate totals under the old metric reproduce the previous table exactly: 2,218 / 6,400 on M55, 49,424 / 149,426 on M33, 9,102 / 26,959 on Hexagon.
    • I have not dispatched compare.yml on GitHub. The recorded counts come from this local reproduction, and one manual run would confirm them on CI.
  • Default build: configured with -DSRT_WERROR=ON; all 73 tests pass.
  • compare-smoke steps: run locally in fresh directories, both pass.
  • Formatting and tidy: clang-format 18.1.3 is clean on all changed C++. clang-tidy is clean on the new shim.
  • clang -Werror: my new code builds clean. The build still fails on existing lines in bench/bench_asrc.cpp:83 and in the old srtBench in bench_compare.cpp, both int64_t/size_t signedness. I didn't touch these, and CI doesn't build them with clang.

Notes for the reviewer

  • Notebooks re-executed. SampleRateTap's own row moved from −132.1 dB to −133.9 dB. The notebook hadn't been re-run since the compensated prototype landed (a45043a, 2026-07-04), so this is a re-measurement, not a change this PR makes.
  • Metric change in the embedded table. Each engine now builds at 2 s and 4 s. The difference is the steady-state cost per frame and the remainder is construction.
    • Steady state, balanced float beats r8brain 1.1–2.2× on every embedded target.
    • Q15 on M33 is 879 insn/frame, against ~26.6k for r8brain and ~49.2k for libsamplerate MEDIUM.
    • Construction is where we lose: about 1.3 G instructions on M33 for the soft-double filter design, against tens of millions for the competitors. The doc says so plainly, and it may deserve its own issue.
  • r8brain on x86. r8brain is faster than balanced on x86 (1.1–1.7×), at 8–33× the filter delay. The doc states this too.
  • No-op mutex for Cortex-M. The Cortex-M comparison builds force-include bench/icount/r8b_single_thread_mutex.h. r8brain's filter cache uses std::mutex with no hook to replace it, and thread-less newlib doesn't declare one. The stand-in is exact for the single-threaded workload and off the per-sample path.
  • Fetch method. r8brain is fetched with GIT_REPOSITORY plus a full commit SHA rather than a tarball URL_HASH. Its newest tag (6.5) predates 7.x, and the archive download was blocked in my environment. It's comparison-only and never linked into the library; the README provenance section now says so.
  • Stale text left for a follow-up. Only the "9.8×" sentence in book/src/part0/two-crystals.md is corrected here. Other stale figures (−132.1 dB, 5,043 insn/frame, June host throughput) remain in the book (two-crystals.md, budgets.md, cortex-m.md, notebooks.md, hardware.md, glossary) and in the Pico example READMEs.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK


Generated by Claude Code

r8brain-free-src (7.5, pinned by commit in cmake/r8brain.cmake) joins
libsamplerate and soxr in every leg of docs/COMPARISON.md:

- Quality (notebooks/asrc_comparison.ipynb): CDSPResampler24 through a
  two-function ctypes shim (tools/compare_shim/, SRT_BUILD_COMPARE_SHIM)
  measures at the format ceilings like the other oracle-fed libraries
  (-143.9 dB THD+N at 24-bit IO, 149.1 dB DR). A new section sweeps its
  transition band and measures latency against the 20 kHz passband: 789
  input frames at the default, 200 at the lowest-latency setting still flat
  to 20 kHz (8%), ~1 ms only at a 45% band that is -35 dB at 20 kHz.
- Host throughput (bench/compare): 120 dB default and passband-matched
  rows plus the 16- and 24-bit presets. r8brain out-runs balanced on x86
  (1.1-1.7x) at 8-33x the filter delay.
- Embedded instruction counts (bench/icount, compare.yml): r8brain on
  M55/M33/Hexagon. It builds for Cortex-M only with a no-op std::mutex
  (forced include; its filter cache has no lock hook and thread-less
  newlib has none). Each engine now also builds at 4 s so steady state
  separates from one-time construction, and SampleRateTap's Q15 datapath
  joins as an engine. Steady state, balanced float beats r8brain 1.1-2.2x;
  Q15 on M33 is 879 insn/frame vs r8brain's ~26.6k.

Re-measurement also corrects stale numbers: balanced now measures -133.9
dB (was -132.1; the compensated design landed after the last run), and
the old per-frame embedded figures folded construction in. Construction
is now reported on its own: ~1.3 G instructions on M33, the one column
where SampleRateTap loses. The libsamplerate totals reproduce the old
table exactly, confirming the harness.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
- compare-smoke now also builds the r8brain ctypes shim on the host and
  the Q15 and r8brain comparison workloads for M55, so the bare-metal
  no-op-mutex path and the new engines can't silently rot between manual
  compare.yml runs.
- README provenance: r8brain-free-src (MIT) is fetched at a commit pin by
  the opt-in comparison builds only, never linked into the library.
- Book (two-crystals): replace the stale "about 9.8x" M33 claim with the
  steady-state figure (~56x the Q15 datapath's 879 insn/frame) and say why
  the number changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G3HxzEGiZp7jStoMuYNuBK
@tap
tap merged commit 838f99b into main Sep 26, 2026
32 checks passed
@tap
tap deleted the claude/sample-rate-solutions-comparison-mqc190 branch September 26, 2026 02:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants