Repository navigation
Conversation
The 16 kHz DPDFNet models denoise well but are offline-only: inside a 48 kHz graph they need a resampler and their output is capped at 8 kHz (telephone band). This adds `<model>_sb` variants that run the network at 48 kHz with no resampler and reconstruct the band above 8 kHz, the same way the GTCRN wrapper does. The registry now carries a model rate and a separate STFT/host rate. A split-band entry pairs a 48 kHz host rate with the 16 kHz model rate and reuses the base model's IR unchanged; build.rs emits the STFT geometry (nfft 960) alongside the model geometry (161 bins) and a spectrum scale. process_frame runs the STFT at the host geometry, feeds the network the low 161 bins scaled to the magnitude it was trained on, scales the enhanced band back up, reconstructs 8-24 kHz with the ported high-band processor (spectral gate / air exciter) and a raised-cosine crossfade, then inverse transforms at 48 kHz. A normal build is the special case where the two geometries are equal and there is no high band — its path and output are unchanged. Validated against the ONNX: feeding the scaled 48 kHz low bins matches the native 16 kHz model to -39 dB. dpdfnet2_sb measures a constant 70 ms across blocks, declared equal to delivered, with the 8-16 kHz band reconstructed to within a dB of the clean reference. Every model, split-band and normal, builds clippy-clean and passes the callback-deadline and degradation tests.
Each reuses its base 16 kHz model's IR, so they need no ONNX of their own and no prepare() step — a `_split_band` list drives their build() after the base models have populated model/<name>/, and one package_*() each installs libdpdfnet_<model>_sb_ladspa.so.
The high-band processor had two modes: pass the captured >8 kHz through a gate, or discard it and synthesize a band from the enhanced low band (mirror 4-8 kHz, square for a second harmonic). The mode was chosen per frame by the high band's SNR. Measured on speech, the synthesis fired on 68-89% of frames, clean speech included: voice carries almost no energy above 8 kHz outside fricatives, so the SNR is low nearly all the time. The SNR test was detecting "not a fricative right now", not "the high band is destroyed" — so the full-band output was mostly fabricated, not the captured voice. For a natural-capture denoiser that is the wrong behaviour: it changes the voice's character silently, and if most of the band is synthetic, resampling to 16 kHz would be honest and cheaper. Drop the synthesizer. The high band is now always the captured signal, gated by the speech probability and a per-bin SNR gate: measured, it keeps the real high band on speech (-2 dB, fricatives preserved) and attenuates it in pauses (-4 to -5 dB, background noise removed). Latency is unchanged at 70 ms; a synthetic highs effect, if ever wanted, belongs in its own opt-in stage. The GTCRN wrapper carries the same air exciter and wants the same change, tracked separately.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 16 kHz DPDFNet networks denoise well but are offline-only: inside a 48 kHz graph they need a resampler and their output is capped at 8 kHz (telephone band). This adds
<model>_sbvariants that run the network at 48 kHz with no resampler and reconstruct the band above 8 kHz, the same way the GTCRN wrapper does.dpdfnet-ladspa-<model>-sb) reuse the base models' IR, so they need no ONNX of their own and no prepare() step.Validated: feeding the scaled 48 kHz low bins matches the native 16 kHz model to -39 dB against the ONNX. dpdfnet2_sb measures a constant 70 ms across blocks (declared == delivered), reconstructs the 8-16 kHz band to within a dB of the clean reference, costs 2.48 s of a core per 10 s of audio against the 48 kHz hi-res variant's 3.14 s, and runs in a real 48 kHz graph with ERR 0 and no clipping. Every model, split-band and normal, builds clippy-clean and passes the callback-deadline and degradation tests.