Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion .git-blame-ignore-revs
Original file line number Diff line number Diff line change
@@ -1,2 +1,4 @@
# Bulk clang-format reformat under the shared Tap house style.
34bb89eb806a91667fcd895b56dea8c0e701739a
b84020e738f771c7689ffc1e592f8448b9ce063f
# Reflow of clang-format alignment after the namespace rename.
e2f5a484712699592396c1df705c47c9ddab30a3
7 changes: 5 additions & 2 deletions .github/workflows/ci-arm64.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ jobs:
run: cmake --build build -j 4

- name: Full test suite
run: ctest --test-dir build --output-on-failure
run: ctest --test-dir build --output-on-failure --no-tests=error

# The ring stress under TSan on weakly-ordered silicon is the point
# of this workflow: run it several times for schedule diversity.
Expand All @@ -57,8 +57,11 @@ jobs:
env:
TSAN_OPTIONS: halt_on_error=1
run: |
# The suite is spsc_ring (it was SpscRing before the snake_case
# rename, when this filter silently started matching nothing);
# --no-tests=error makes a stale filter fail instead of passing.
for i in 1 2 3 4 5; do
ctest --test-dir build-tsan -R 'SpscRing' --output-on-failure
ctest --test-dir build-tsan -R 'spsc_ring' --output-on-failure --no-tests=error
done

# Scheduled runs have no PR or push audience; without this a weekly
Expand Down
115 changes: 99 additions & 16 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -1,15 +1,25 @@
name: CI

# Push runs only for main; branches get CI through their pull request (and
# on demand). Before, a push to a PR branch ran everything twice: once for
# the push and once for the pull_request event, in different groups.
on:
push:
branches: [main]
pull_request:
workflow_dispatch:

permissions:
contents: read

concurrency:
# Superseded pushes cancel their own in-flight runs: the ratchet stack
# (three QEMU targets, a toolchain download, minutes of runner time) was
# piling up once per push during baseline harvests.
group: ci-${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
# One group per PR (or per ref for main/dispatch). Superseded PR runs are
# cancelled: the ratchet stack (three QEMU targets, a toolchain download,
# minutes of runner time) piled up once per push during baseline
# harvests. Never on main, and never on the monorepo migration PR, whose
# plan requires a complete run for every pushed commit.
group: ci-${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' && github.head_ref != 'claude/sample-rate-expansion-strategies-ezqzu6' }}

jobs:
build-and-test:
Expand Down Expand Up @@ -64,7 +74,7 @@ jobs:
run: cmake --build build --config Release -j 4

- name: Test
run: ctest --test-dir build -C Release --output-on-failure
run: ctest --test-dir build -C Release --output-on-failure --no-tests=error

sanitizers:
name: ${{ matrix.name }}
Expand Down Expand Up @@ -99,7 +109,7 @@ jobs:
env:
TSAN_OPTIONS: halt_on_error=1
UBSAN_OPTIONS: print_stacktrace=1
run: ctest --test-dir build --output-on-failure
run: ctest --test-dir build --output-on-failure --no-tests=error

# Cross-compile for Qualcomm Hexagon (Linux/musl) with the open-source
# toolchain and run a subset of the suite under qemu-hexagon user-mode
Expand All @@ -110,7 +120,7 @@ jobs:
# neither speeds up nor measures meaningfully.
hexagon-qemu:
name: Hexagon cross (QEMU)
runs-on: ubuntu-latest
runs-on: ubuntu-24.04
timeout-minutes: 45
env:
# Prebuilt open-source toolchain (BSD-3) published by Qualcomm/Quicinc;
Expand Down Expand Up @@ -187,24 +197,35 @@ jobs:
- name: Build
run: cmake --build build -j 4

# -V --output-log keeps every test's own output ([ RUN ] lines,
# [ measured ] numbers), which --output-on-failure prints only for
# failures; the log is uploaded below as evidence.
- name: Test under emulation
run: >
ctest --test-dir build --output-on-failure
ctest --test-dir build --output-on-failure --no-tests=error
-V --output-log ctest-hexagon.log
-E 'AsrcQuality|AsrcLock|TwoThreadStress|TransparentPrototypeMeetsSpec|MultiChannel\.|Feasibility|Reset\.|ConfigValidation'
# ConfigValidation: this static-musl toolchain cannot unwind across
# frames — the constructor throws correctly but EXPECT_THROW never
# catches and libc++abi terminates. Validation is target-independent
# and covered on every other leg; limitation tracked in
# docs/PERFORMANCE.md "Known debt".

- name: Upload test log
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ctest-hexagon
path: ctest-hexagon.log

# Cross-compile for Arm Cortex-M55 (bare metal, newlib + semihosting) and
# run the emulation-sized test subset on QEMU's MPS3 AN547 board model.
# Validates the library on a 32-bit MCU-class target with no OS, no
# threads and no double-precision FPU; the fixed-point datapaths are the
# performance-appropriate formats here.
cortex-m55-qemu:
name: Cortex-M55 cross (QEMU)
runs-on: ubuntu-latest
runs-on: ubuntu-24.04
timeout-minutes: 30
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
Expand All @@ -227,16 +248,27 @@ jobs:
run: cmake --build build -j 4

- name: Test under emulation
run: ctest --test-dir build --output-on-failure
run: >
ctest --test-dir build --output-on-failure --no-tests=error
-V --output-log ctest-m55.log

- name: Upload test log
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ctest-m55
path: ctest-m55.log

# Cortex-M33 (Raspberry Pi Pico 2 / RP2350 class: single-precision FPU,
# no FP64, no MVE) on QEMU's MPS2+ AN505 model. Shares the Armv8-M
# startup with the M55 target; quantifies the soft-double float path and
# anchors the Q15/Q31 budgets for Pico-class parts.
cortex-m33-qemu:
name: Cortex-M33 cross (QEMU)
runs-on: ubuntu-latest
timeout-minutes: 30
runs-on: ubuntu-24.04
# ~23 min of emulated suite (plus ~1 min of output hashes) against the
# old 30: headroom for a slower runner.
timeout-minutes: 40
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
Expand All @@ -258,7 +290,16 @@ jobs:
run: cmake --build build -j 4

- name: Test under emulation
run: ctest --test-dir build --output-on-failure
run: >
ctest --test-dir build --output-on-failure --no-tests=error
-V --output-log ctest-m33.log

- name: Upload test log
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ctest-m33
path: ctest-m33.log

# ------------------------------------------------------------------------
# Template: genuine Tensilica HiFi4/HiFi5 coverage. The HiFi audio ISA,
Expand Down Expand Up @@ -297,7 +338,9 @@ jobs:
# so a hard >3% gate is safe on shared runners.
icount-ratchet:
name: Instruction-count ratchet
runs-on: ubuntu-latest
# Pinned image: the counts are a function of the apt toolchain, the
# plugin build and QEMU, so the job must not straddle an image rollout.
runs-on: ubuntu-24.04
timeout-minutes: 45
env:
# Commit the v8.2.2 tag pointed at when pinned (tags are movable;
Expand Down Expand Up @@ -332,6 +375,15 @@ jobs:
gcc -shared -fPIC $(pkg-config --cflags glib-2.0) -I/tmp \
-o /tmp/libinsncount.so tools/qemu_insn_plugin/insn_count.c

# What produced the counts: the runner image and the toolchain packages
# (the counts move when either does; see docs/PERFORMANCE.md).
- name: Record image and toolchain versions
id: image
run: |
echo "image: ${ImageOS:-unknown} ${ImageVersion:-unknown}"
dpkg-query -W gcc-arm-none-eabi qemu-system-arm
echo "os=${ImageOS:-unknown}" >> "$GITHUB_OUTPUT"

# Release (-O2), matching how the baselines were recorded.
- name: Build M55 workloads
run: >
Expand Down Expand Up @@ -376,7 +428,10 @@ jobs:
uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5
with:
path: ~/qemu-hexagon-plugins
key: qemu-hexagon-plugins-${{ env.QEMU_SRC_URL }}-1
# The image OS is part of the key: the cached binary links the
# image's glib, so a new Ubuntu release must not restore an old
# build. (Weekly image updates keep the same glib ABI.)
key: qemu-hexagon-plugins-${{ steps.image.outputs.os }}-${{ env.QEMU_SRC_URL }}-1

- name: Build plugin-enabled qemu-hexagon
if: ${{ !cancelled() && steps.qemu-hex.outputs.cache-hit != 'true' }}
Expand Down Expand Up @@ -515,6 +570,34 @@ jobs:
&& cmake --build build-m55 -j 4 --target cmp_icount_lsr_medium
cmp_icount_srt_q15 cmp_icount_r8b_120

# The Pico 2 firmware examples are standalone projects (Pico SDK fetched at
# configure time), deliberately outside the root build. Building them here
# keeps them from silently rotting again: they had stopped compiling after
# the snake_case rename and the DspTap substrate move, with nothing
# noticing. Build-only; running them needs the hardware
# (docs/HARDWARE_TESTING.md).
pico2-build:
name: Pico 2 firmware build
runs-on: ubuntu-24.04
timeout-minutes: 20
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
submodules: recursive

- name: Install toolchain
run: sudo apt-get update -q && sudo apt-get install -y -q gcc-arm-none-eabi

- name: Build pico2_cyccnt
run: >
cmake -S examples/pico2_cyccnt -B build-cyccnt -DPICO_BOARD=pico2
&& cmake --build build-cyccnt -j 4

- name: Build pico2_dualcore
run: >
cmake -S examples/pico2_dualcore -B build-dualcore -DPICO_BOARD=pico2
&& cmake --build build-dualcore -j 4

clang-format:
name: clang-format
runs-on: ubuntu-latest
Expand Down
11 changes: 9 additions & 2 deletions .github/workflows/style.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,14 @@ name: Tap House Style
# Enforces the shared Tap House Rules. clang-format is already checked in
# ci.yml; this adds (1) a drift check against the canonical TapHouse configs
# and (2) clang-tidy naming + mandatory-braces enforcement.
on: [push, pull_request]
on:
push:
branches: [main]
pull_request:
workflow_dispatch:

permissions:
contents: read

jobs:
drift:
Expand All @@ -15,7 +22,7 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
submodules: recursive
- name: Install tools
Expand Down
26 changes: 13 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,9 +56,9 @@ transparency vs. a naive FIFO, spectrograms, latency, drift tracking,
dropout recovery — see
[notebooks/asrc_demo.ipynb](notebooks/asrc_demo.ipynb), which drives the
library through its C ABI (`-DSRT_BUILD_CAPI=ON`, `tools/capi/`) via ctypes
(Python needs `numpy` and `matplotlib`; the comparison notebook below
additionally needs the `samplerate` and `soxr` packages; the first cell
builds the shared library if missing). A second notebook,
(the notebook environment is pinned, with hashes, in `requirements.lock`:
`pip install --require-hashes -r requirements.lock`; the first cell
rebuilds the shared library incrementally on every run). A second notebook,
[notebooks/asrc_block_size_study.ipynb](notebooks/asrc_block_size_study.ipynb),
measures how processing block size (32 / 64 / 240 frames) trades latency
against servo observability — including per-impulse latency-breathing
Expand Down Expand Up @@ -159,7 +159,7 @@ theoretically unavailable from counts alone, so the servo deliberately stays
in Track, where the block beat is phase-tracked mostly as benign latency
breathing, the remainder as cent-scale low-rate FM (measured in
[notebooks/asrc_block_size_study.ipynb](notebooks/asrc_block_size_study.ipynb):
~0.9 cents rms / 61 dB wideband at 32-frame blocks, ~1.3 cents rms / 53 dB
~0.9 cents rms / 62 dB wideband at 32-frame blocks, ~1.3 cents rms / 54 dB
at 5 ms blocks). Promotion to Quiet is gated on the
cascade-smoothed error, which is exactly the discriminator between the two
regimes.
Expand Down Expand Up @@ -327,13 +327,13 @@ Executed instructions per fixed workload (`bench/icount/`), measured under QEMU

| Workload | Cortex-M33 | Cortex-M55 | Hexagon |
|---|---:|---:|---:|
| `kernel_float` | 2,427,595,993 | 109,298,076 | 422,164,193 |
| `kernel_q15` | 1,123,119,218 | 192,431,957 | 187,432,836 |
| `kernel_q31` | 1,169,957,251 | 221,164,795 | 194,967,771 |
| `pipeline12_q15` | 1,498,750,507 | 398,227,192 | 463,502,418 |
| `pipeline_float` | 2,391,686,215 | 102,590,477 | 419,190,309 |
| `pipeline_q15` | 1,020,281,310 | 137,792,386 | 204,480,261 |
| `pipeline_q31` | 1,102,554,828 | 173,088,119 | 205,227,716 |
| `kernel_float` | 2,427,595,993 | 109,298,076 | 422,160,503 |
| `kernel_q15` | 1,123,119,218 | 192,431,957 | 187,429,180 |
| `kernel_q31` | 1,169,957,251 | 221,164,795 | 194,964,115 |
| `pipeline12_q15` | 1,498,750,507 | 398,227,192 | 463,498,694 |
| `pipeline_float` | 2,391,686,215 | 102,590,477 | 419,186,585 |
| `pipeline_q15` | 1,020,281,310 | 137,792,386 | 204,476,571 |
| `pipeline_q31` | 1,102,554,828 | 173,088,119 | 205,224,026 |
<!-- ICOUNT:END -->

<!-- PERF:BEGIN -->
Expand Down Expand Up @@ -422,8 +422,8 @@ inferred from a float ratio:
exactly, 2.0 ms total latency).

The two converters check each other: RatioTap's suite cross-validates its
output against this library's async engine at −109 dB (down) / −99 dB (up)
over every polyphase phase.
output against this library's async engine at −98 dB (down) / −90 dB (up)
on its default `economy` profile, over every polyphase phase.

## Limitations

Expand Down
14 changes: 7 additions & 7 deletions bench/baselines.json
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
{
"hexagon": {
"kernel_float": 422164193,
"kernel_q15": 187432836,
"kernel_q31": 194967771,
"pipeline12_q15": 463502418,
"pipeline_float": 419190309,
"pipeline_q15": 204480261,
"pipeline_q31": 205227716
"kernel_float": 422160503,
"kernel_q15": 187429180,
"kernel_q31": 194964115,
"pipeline12_q15": 463498694,
"pipeline_float": 419186585,
"pipeline_q15": 204476571,
"pipeline_q31": 205224026
},
"m33": {
"kernel_float": 2427595993,
Expand Down
4 changes: 2 additions & 2 deletions book/src/part1/pi-servo.md
Original file line number Diff line number Diff line change
Expand Up @@ -446,8 +446,8 @@ beat: most of the sawtooth is absorbed as **latency breathing** — the
buffer level, and hence the delay, swaying by a fraction of the block at
the beat rate, inaudible by construction. The remainder leaks into ε̂ as
low-rate FM, and the study put calibrated numbers on it: **~0.9 cents rms
of frequency wobble (61 dB wideband quality) at 32-frame blocks, ~1.3
cents / 53 dB at 5 ms blocks**, as the README reports. Cent-scale wobble
of frequency wobble (62 dB wideband quality) at 32-frame blocks, ~1.3
cents / 54 dB at 5 ms blocks**, as the README reports. Cent-scale wobble
at sub-hertz rates is at the edge of perception for sustained pure tones
and irrelevant for program material — but it is a real ceiling, and it is
a *sensor* ceiling, not a servo defect. The README's limitations section
Expand Down
9 changes: 5 additions & 4 deletions book/src/part2/notebooks.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,12 +29,12 @@ the naive-FIFO disaster, then walks lock acquisition, transparency,
spectrograms, latency, drift tracking, and dropout recovery. Its committed
outputs are where the README's "what does it sound like" numbers come from:
clicks roughly ten times per second at 29 dB SNR for the naive path,
126.4 dB for the converter under the notebook's instrument.
125.9 dB for the converter under the notebook's instrument.

**`asrc_block_size_study.ipynb`** answers a deployment question: what
happens at block sizes 32, 64, and 240 frames? Its committed conclusion —
Track-stage operation turns block quantization into cent-scale, low-rate FM
over a 53–61 dB wideband floor, while designed latency scales as roughly
over a 54–62 dB wideband floor, while designed latency scales as roughly
`2·B/fs + 0.5 ms` — is quoted by `docs/COMPARISON.md` whenever coarse-block
operation comes up.

Expand Down Expand Up @@ -267,7 +267,7 @@ The last trap is the quietest, and this project walked into it. The demo
notebook's measurement cell printed, in its committed output:

```text
ASRC SNR: 126.4 dB | naive: 29.4 dB | improvement: 97 dB
ASRC SNR: 125.9 dB | naive: 29.4 dB | improvement: 97 dB
```

with `assert snr_asrc > 125.0` enforcing it. The *summary table* at the
Expand All @@ -280,7 +280,8 @@ results. (The measured 135 dB figure from the test suite is real, but it is
a *different instrument* — a tracked global fit over a different window —
and a summary must quote its own cell, not the best number available
elsewhere in the repo.) The fix was the boring, correct one: the summary
now states 126.4 dB and points at the assertion.
now states the measured figure (125.9 dB since the compensated
prototype design; re-executed 2026-09) and points at the assertion.

The lesson generalizes beyond notebooks: **summaries drift from cells the
same way READMEs drift from benchmarks and comments drift from code.**
Expand Down
Loading
Loading