Skip to content

Split-band variants: run the 16 kHz models full-band at 48 kHz - #2

Open
bigbruno wants to merge 3 commits into
mainfrom
feat/split-band-16k
Open

bigbruno wants to merge 3 commits into
mainfrom
feat/split-band-16k

Conversation

@bigbruno

@bigbruno bigbruno commented Sep 9, 2026

Copy link
Copy Markdown
Member

The 16 kHz DPDFNet networks denoise well but are offline-only: inside a 48 kHz graph they need a resampler and their output is capped at 8 kHz (telephone band). This adds <model>_sb variants that run the network at 48 kHz with no resampler and reconstruct the band above 8 kHz, the same way the GTCRN wrapper does.

  • The build registry now carries a model rate and a separate STFT/host rate. A split-band entry pairs a 48 kHz host rate with the 16 kHz model rate and reuses the base model's IR unchanged; build.rs emits the STFT geometry (nfft 960) alongside the model geometry (161 bins) and a spectrum scale.
  • process_frame runs the STFT at the host geometry, feeds the network the low 161 bins scaled to the magnitude it was trained on, scales the enhanced band back up, reconstructs 8-24 kHz with the ported high-band processor (spectral gate / air exciter) and a raised-cosine crossfade, then inverse transforms at 48 kHz. A normal build is the special case where the two geometries are equal and there is no high band — its path and output are unchanged.
  • Four split packages (dpdfnet-ladspa-<model>-sb) reuse the base models' IR, so they need no ONNX of their own and no prepare() step.

Validated: feeding the scaled 48 kHz low bins matches the native 16 kHz model to -39 dB against the ONNX. dpdfnet2_sb measures a constant 70 ms across blocks (declared == delivered), reconstructs the 8-16 kHz band to within a dB of the clean reference, costs 2.48 s of a core per 10 s of audio against the 48 kHz hi-res variant's 3.14 s, and runs in a real 48 kHz graph with ERR 0 and no clipping. Every model, split-band and normal, builds clippy-clean and passes the callback-deadline and degradation tests.

The 16 kHz DPDFNet models denoise well but are offline-only: inside a
48 kHz graph they need a resampler and their output is capped at 8 kHz
(telephone band). This adds `<model>_sb` variants that run the network at
48 kHz with no resampler and reconstruct the band above 8 kHz, the same way
the GTCRN wrapper does.

The registry now carries a model rate and a separate STFT/host rate. A
split-band entry pairs a 48 kHz host rate with the 16 kHz model rate and
reuses the base model's IR unchanged; build.rs emits the STFT geometry
(nfft 960) alongside the model geometry (161 bins) and a spectrum scale.

process_frame runs the STFT at the host geometry, feeds the network the low
161 bins scaled to the magnitude it was trained on, scales the enhanced band
back up, reconstructs 8-24 kHz with the ported high-band processor
(spectral gate / air exciter) and a raised-cosine crossfade, then inverse
transforms at 48 kHz. A normal build is the special case where the two
geometries are equal and there is no high band — its path and output are
unchanged.

Validated against the ONNX: feeding the scaled 48 kHz low bins matches the
native 16 kHz model to -39 dB. dpdfnet2_sb measures a constant 70 ms across
blocks, declared equal to delivered, with the 8-16 kHz band reconstructed to
within a dB of the clean reference. Every model, split-band and normal,
builds clippy-clean and passes the callback-deadline and degradation tests.
Each reuses its base 16 kHz model's IR, so they need no ONNX of their own
and no prepare() step — a `_split_band` list drives their build() after the
base models have populated model/<name>/, and one package_*() each installs
libdpdfnet_<model>_sb_ladspa.so.
The high-band processor had two modes: pass the captured >8 kHz through a
gate, or discard it and synthesize a band from the enhanced low band (mirror
4-8 kHz, square for a second harmonic). The mode was chosen per frame by the
high band's SNR.

Measured on speech, the synthesis fired on 68-89% of frames, clean speech
included: voice carries almost no energy above 8 kHz outside fricatives, so
the SNR is low nearly all the time. The SNR test was detecting "not a
fricative right now", not "the high band is destroyed" — so the full-band
output was mostly fabricated, not the captured voice. For a natural-capture
denoiser that is the wrong behaviour: it changes the voice's character
silently, and if most of the band is synthetic, resampling to 16 kHz would be
honest and cheaper.

Drop the synthesizer. The high band is now always the captured signal, gated
by the speech probability and a per-bin SNR gate: measured, it keeps the real
high band on speech (-2 dB, fricatives preserved) and attenuates it in pauses
(-4 to -5 dB, background noise removed). Latency is unchanged at 70 ms; a
synthetic highs effect, if ever wanted, belongs in its own opt-in stage.

The GTCRN wrapper carries the same air exciter and wants the same change,
tracked separately.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant