Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ the FFT came from). Seven primitives today: the real FFT (`fft.h`), the YIN pitc
the log-mel/PCEN feature extractor (`log_mel.h`) and the fixed-ratio decimators to
16 kHz (`decimate.h`) — and the dense/GRU inference kernels (`nn.h`) — plus the FIR substrate carried from
SampleRateTap for the two rate converters (SampleRateTap, RatioTap): Kaiser prototype design
(`kaiser.h`), the sample-format traits (`sample_traits.h`: float/Q15/Q31), the FIR dot kernels
(`kaiser.h`), the sample-format traits (`sample_traits.h`: double/float/Q15/Q31), the FIR dot kernels
(`fir_kernels.h`), row-sum-preserving quantization (`quantize.h`), and the measurement
instruments (`analysis/`). See `README.md` for each asset's contract summary.

Expand Down
52 changes: 45 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,8 +170,8 @@ arithmetic on the float path: the RP2350's Cortex-M33 has no FP64.
front end: `basic_decimator<Sample, M>` for M = 2, 3, 6 (32 / 48 / 96 kHz in),
in RatioTap's pattern — ratio as a type, Kaiser-windowed sinc from
`kaiser.h` with the cutoff at the output Nyquist and DC gain exactly 1,
`fir_kernels.h`'s `dot_row` over the `sample_traits.h` formats (float golden,
Q15 / Q31 with row-sum-preserving quantization). It is deliberately not
`fir_kernels.h`'s `dot_row` over the `sample_traits.h` formats (double golden,
float embedded, Q15 / Q31 with row-sum-preserving quantization). It is deliberately not
RatioTap, whose charter is 44.1 ↔ 48 only; a 44.1 kHz host composes RatioTap's
44.1 → 48 in front of the by-3 stage. Odd tap counts, integer group delay
`(taps - 1) / 2`, one output as input `k*M` arrives.
Expand Down Expand Up @@ -228,7 +228,10 @@ suppressor's cross-precision pin must be unchanged by the promotion.

Five headers carried from **SampleRateTap** (where they design and run the
ASRC's polyphase datapath) and promoted here so **RatioTap**'s fixed-ratio
44.1↔48 converter — and any future FIR consumer — shares one implementation.
44.1↔48 converter — and any future FIR consumer — shares one implementation,
plus the FFT's butterfly arithmetic trait (`fft/fft_arith.h`), which is built
over the same sample formats and documented here until the Stage 3b README
rewrite moves it to the FFT section's profiles table.
The performance-sensitive pieces are regression-gated in SampleRateTap's
instruction-count CI (Cortex-M33/M55, Hexagon, ±3%); treat measured claims in
the header comments as contracts.
Expand All @@ -245,16 +248,30 @@ constexpr (the header's design note does the arithmetic); run it in a
constructor, off the audio path. Also exports `solve_dense`, the small dense
solver the compensated design and the analysis instruments share.

### `tap/dsp/sample_traits.h` — sample formats: float, Q15, Q31
### `tap/dsp/sample_traits.h` — sample formats: double, float, Q15, Q31

The family's sample-format substrate: how each sample type stores
coefficients, accumulates dot products, and rounds/saturates back to samples.
`double` is a sample format because it is the golden model of every
primitive: a traits-based primitive instantiates its reference profile
through the same substrate as its embedded profiles, and the cross-precision
pins measure float/Q15/Q31 against it.

| Type | Coefficients | Accumulation | Output |
|---|---|---|---|
| `double` | double | double | identity (the golden model) |
| `float` | float | double | plain cast |
| `std::int16_t` | Q1.14 | int64, exact | single Q29→Q15 round-half-up, saturating |
| `std::int32_t` | Q1.30 | int64, products pre-shifted to Q45 | Q45→Q31, saturating |
| `std::int32_t` | Q1.30 | int64, products pre-shifted to Q45 | single Q45→Q31 round-half-up, saturating |

The Q ladder is spelled as named constants on each fixed-point
specialization (`k_sample_frac_bits`, `k_coeff_frac_bits`,
`k_accum_pre_shift`) with the accumulator format and the single 14-bit
finalize shift derived from them and `static_assert`ed; `k_coeff_scale` is
`2^k_coeff_frac_bits`, and `k_is_fixed_point` is what selects the
fixed-point algorithm in `quantize.h`. Every member is `constexpr`, and the
`sample_type` concept requires a value-initialized accumulator to be the
additive identity.

**Fixed point is a first-class embedded direction, not a legacy path.** The
Q15/Q31 profiles exist for targets where double (sometimes any float) is
Expand All @@ -273,6 +290,24 @@ This header is the format core only. Engine-specific extensions (e.g.
SampleRateTap's inter-phase coefficient blending) derive from these
specializations and refine the `tap::dsp::sample_type` concept.

### `tap/dsp/fft/fft_arith.h` — butterfly arithmetic for the FFT profiles

The sibling trait the fixed-point real FFT is written against (its docstrings
are that kernel's specification — shift-before-butterfly, the magnitude bound,
the rounding count per complex product, the BFP headroom rule, the Q31 input
pre-shift, the twiddle generator; `tests/test_fft_arith.cpp` pins every
number). Documented here until the Stage 3b README rewrite; it moves to the
FFT section's profiles table then. One int32 kernel serves both fixed profiles: `fft_arith<std::int32_t>`
carries `mul_coeff` (int32 × Q1.30 → int64, `>> 30` with one round-half-up,
saturating), saturating `add` / `sub`, `shr_round` (round-half-up), and
`headroom_bits` (the block's shared redundant sign bits, via
`std::countl_zero`); `fft_arith<std::int16_t>` is the Q15 I/O width —
`widen` (`<< 14`, two guard bits) and `narrow` (round-half-up, saturating) —
and names the int32 trait as its `work`. Twiddles are
`sample_traits<std::int32_t>::coeff` (Q1.30) for both fixed profiles, and 1.0
is representable. The `float` / `double` specializations are the same names
over plain arithmetic. Everything is `constexpr` and `noexcept`.

### `tap/dsp/fir_kernels.h` — dot-product kernels

The FIR hot loops, target-gated the way SampleRateTap's optimization campaign
Expand All @@ -290,8 +325,11 @@ targets prefer which layout.
`quantize_row_preserving_sum`: quantizes one polyphase branch to a fixed-point
coefficient format while preserving the row's DC sum *exactly*
(largest-remainder distribution of the rounding residual — "the coefficients
of every phase must add to one", R. Bristow-Johnson, music-dsp). Plain
conversion for float. Design-time code.
of every phase must add to one", R. Bristow-Johnson, music-dsp). Selected by
the trait's `k_is_fixed_point`: plain conversion for double and float. A tap
at the format's rail is never wrapped: each step goes to the largest-remainder
tap that can still move in that direction, so the sum is preserved whenever
one exists. Design-time code.

### `tap/dsp/analysis/` — measurement instruments

Expand Down
18 changes: 12 additions & 6 deletions include/tap/dsp/decimate.h
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,9 @@
// 96 kHz in, 16 kHz out. Built in RatioTap's pattern — the ratio is a
// compile-time type, the prototype is a Kaiser-windowed sinc from kaiser.h,
// the hot loop is fir_kernels.h's dot_row over the sample_traits.h formats
// (float golden, Q15/Q31 fixed point) — but deliberately NOT RatioTap: that
// library's charter is 44.1 <-> 48 only. A 44.1 kHz host composes RatioTap's
// 44.1 -> 48 in front of the by-3 stage here.
// (double golden, float embedded, Q15/Q31 fixed point) — but deliberately
// NOT RatioTap: that library's charter is 44.1 <-> 48 only. A 44.1 kHz host
// composes RatioTap's 44.1 -> 48 in front of the by-3 stage here.
//
// Contract, as numbers:
// - Ratios: 2, 3 and 6 (input 32 / 48 / 96 kHz for a 16 kHz output). The
Expand Down Expand Up @@ -37,10 +37,16 @@
// bands stop at 7.6 kHz and its features carry little above 7 kHz, and
// economy costs 121 MACs per 16 kHz output (by 3) — under 20k MACs per
// 10 ms hop. transparent exists for offline use.
// - Sample formats: float (double accumulation, the golden model, pinned
// - Sample formats: double (the golden model, as for every primitive),
// float (double accumulation, the embedded profile, pinned
// sample-for-sample against a committed numpy reference), Q15 and Q31
// through sample_traits.h with row-sum-preserving quantization so DC
// gain stays exactly 1. Mono: the consumer is single-channel by charter.
// gain stays exactly 1; the Q15 and Q31 profiles are pinned against
// double as measured numbers, and their coefficient tables are bit-pinned
// (row sum and FNV-1a-64 per ratio and profile). The decimate_by_*
// aliases stay float: they name the
// 16 kHz front end's deployed profile. Mono: the consumer is
// single-channel by charter.
//
// Construction designs the filter (runtime double, off the audio path) and
// allocates; process() and reset() are noexcept and allocation-free.
Expand Down Expand Up @@ -185,7 +191,7 @@ namespace tap::dsp {
std::size_t m_phase = 0; ///< inputs since the last emitted output, in [0, M)
};

using decimate_by_2 = basic_decimator<float, 2>; ///< 32 kHz -> 16 kHz, float golden
using decimate_by_2 = basic_decimator<float, 2>; ///< 32 kHz -> 16 kHz, float embedded profile
using decimate_by_3 = basic_decimator<float, 3>; ///< 48 kHz -> 16 kHz
using decimate_by_6 = basic_decimator<float, 6>; ///< 96 kHz -> 16 kHz

Expand Down
Loading
Loading