Small local pipeline for:
- downloading raw recordings from ClickUp
- trimming boundary silence and optional spoken opening dates while preserving titles
- denoising and exporting polished MP3s
- uploading the finished MP3s back to ClickUp
- Run setup once:
bash setup.sh- Run the end-to-end pipeline:
python3 run_pipeline.pyThat single command runs:
workflow1_download.pyprepare_raw_audio.py runprocess_audio.sh --skip-prepareworkflow2_upload.py
Useful end-to-end variations:
python3 run_pipeline.py --stop-after prepare
python3 run_pipeline.py --skip-upload
python3 run_pipeline.py --trim-method ffmpeg-vad-asr --trim-debug --trim-force
python3 run_pipeline.py --cleanUse --stop-after prepare for a safe first real-data pass if you want to inspect trim-artifacts/manifest.csv and a few sidecars before denoise/export/upload.
Use --clean to automatically remove audio working files after a successful upload, keeping the workspace ready for the next session. Metadata sidecars are preserved.
The scheduled Process ClickUp Audio workflow is subject to GitHub's public repository inactivity policy: scheduled workflows are automatically disabled after 60 days without repository activity. There is no workflow-level setting to extend that inactivity window. If GitHub disables the schedule, re-enable the workflow in the Actions tab, or make a commit that changes the workflow cron schedule to reactivate it.
If you want to run or debug the stages manually:
python3 workflow1_download.py
python3 prepare_raw_audio.py run --debug --skip-existing
bash process_audio.sh
python3 workflow2_upload.pyprocess_audio.sh auto-runs the raw-audio prepare step, so step 3 is optional if you just want the end-to-end flow.
audio-input/: original downloaded source audio using the task/attachment stemaudio-input/.metadata/: download-time metadata with ClickUptask_idand attachment detailsaudio-prepared/: trimmed WAVs using the same human-readable stemaudio-prepared/.metadata/: trimmed-audio metadata used to recovertask_iddownstreamaudio-output/: final MP3 exports using the same human-readable stemtrim-artifacts/: sidecars, manifests, logs, debug artifactstrim-experiments/: experiment runs comparing trim methods
prepare_raw_audio.py adds a layered preprocessing stage:
- ffmpeg silence baseline
- optional WebRTC VAD boundary detection
- optional Faster-Whisper intro classification for
date | title | content
When ASR is enabled, the opening transcript is classified so that:
dateis removabletitleis detected but preservedcontentis preserved
Operational safeguards:
- malformed
classified_segmentsare rejected before they can affect trim metadata - a single-file prepare failure is written to a failure sidecar and manifest row instead of aborting the whole batch
- if a file fails during prepare, any stale prepared WAV for that file is removed so downstream processing does not reuse it accidentally
Useful commands:
python3 prepare_raw_audio.py run --dry-run --debug
python3 prepare_raw_audio.py experiment --sample-size 12Configuration lives in config/trim_audio.toml.
The design note is in docs/raw-audio-trimming.md.
Download and naming behavior:
- filenames are derived from the audio attachment filename, not the volunteer-entered task name, to prevent date/name errors from propagating through the pipeline
- if the task name and attachment filename don't match, a note is posted on the ClickUp task activity panel so the team can review
- if two tasks share the same attachment filename, the pipeline appends the task ID to disambiguate and posts a note on the affected task
- the same stem is preserved through
audio-prepared/andaudio-output/ - ClickUp
task_idlives in metadata sidecars, so upload matching stays reliable even though the visible files are human-readable - the downloader chooses audio attachments deterministically, prefers non-
copyvariants, and preserves the real extension from the source attachment
process_audio.sh runs after trimming and applies:
- DeepFilterNet neural noise reduction
- Residual leading silence removal (catches anything the trim step missed)
- Loudness normalization to -20 LUFS / -1 dBTP
- Exactly 4 seconds of clean silence prepended for a consistent start
- Export as 192kbps stereo MP3 at 44.1kHz
Every output has the same format and lead-in regardless of how the raw recording was captured.
Each batch is a self-contained session:
python3 run_pipeline.py— downloads, trims, processes, uploadspython3 run_pipeline.py --clean— same, but clears audio folders after upload
The --clean flag removes files from audio-input/, audio-prepared/, audio-output/, and trim-artifacts/ but preserves .metadata/ sidecars. To review output before uploading, use --stop-after process first.
Baseline trimming only needs ffmpeg and ffprobe.
For advanced VAD and ASR trimming:
pip install webrtcvad-wheels faster-whisperAfter installing advanced backends, do one forced re-run so older silence-only sidecars are refreshed:
python3 prepare_raw_audio.py run --method ffmpeg-vad-asr --debug --forcefaster-whisper also needs a Whisper model available locally. On the first run it may download the configured model, or you can point config/trim_audio.toml asr.model_size at a local model path.
When ffmpeg-vad or ffmpeg-vad-asr is requested, the pipeline now fails fast if the required backend is not actually ready. It will not silently fall back to plain ffmpeg trimming.
Raw recordings, prepared audio, exports, and experiment artifacts are intentionally ignored by git in .gitignore.