Summary
When using this pipeline, it's difficult to know upfront:
- What inputs each script requires
- What outputs each script produces
- Dependencies between scripts
- External tool requirements
Currently this information is scattered across shell scripts and Python files, requiring users to read through the code to understand the full picture. This makes it harder to:
- Integrate the pipeline into larger workflows
- Debug issues when outputs are missing
- Know which scripts need to run before others
Proposal: Add pipeline.yaml
A single YAML file at the repo root that declares inputs/outputs for each script. This would serve as:
- Living documentation that's easy to scan
- A contract that can be validated against actual outputs
- A reference for anyone integrating the pipeline into their workflow
Example structure:
# pipeline.yaml
version: "1.0"
description: "Entropy-based cell calling for scATAC-seq"
environment:
required_tools:
- name: ripgrep
command: rg
- name: samtools
command: samtools
- name: bedtools
command: bedtools
- name: Genrich
path_var: Genrich # from config.sh
optional_tools:
- name: HOMER
command: annotatePeaks.pl
note: "Requires genome installation (e.g., hg38)"
scripts:
auto_process.sh:
description: "Calculate entropy, auto-threshold, filter fragments"
inputs:
- flag: -f
name: fragments_file
required: true
format: "*.tsv.gz or *.tsv"
- flag: -o
name: output_dir
required: true
- flag: -g
name: species
required: true
values: [hg38, mm10]
- flag: -c
name: chromosome
default: chr1
- flag: -w
name: window_size
default: 3000
outputs:
- "{output_dir}/fragments.tsv"
- "{output_dir}/{chromosome}_barcode_entropy_df.tsv"
- "{output_dir}/entropy_cutoff.csv"
- "{output_dir}/filtered_fragments.tsv"
- "{output_dir}/Entropy_filtered_bc_CBZ.txt"
- "{output_dir}/figures/{chromosome}_entropy_knee_plot_color_crbc.png"
- "{output_dir}/figures/{chromosome}_entropy_knee_plot_logscale.png"
- "{output_dir}/figures/{chromosome}_Entropy_histogram.png"
- "{output_dir}/figures/entropy_violin_EmptyCRcell.png"
- "{output_dir}/figures/log10_entropy_violin_EmptyvsCell.png"
- "{output_dir}/figures/frag_length_distribution.png"
compare_cellcalling.sh:
description: "Compare entropy vs CellRanger cell calling"
inputs:
- flag: -d
name: output_dir
required: true
- flag: -c
name: chromosome
required: true
- flag: -b
name: cellranger_barcode_file
required: true
outputs:
- "{output_dir}/figures/{chromosome}_entropy_knee_plot_color_crbc.png"
requires:
- auto_process.sh
compare_peakcalling.sh:
description: "Compare peaks before/after entropy filtering"
inputs:
- flag: -d
name: output_dir
required: true
- flag: -g
name: species
required: true
- flag: -s
name: bam_file
required: true
- flag: -c
name: chromosome
required: true
outputs:
- "{output_dir}/peaks/peaks_w_blacklistregion.bed"
- "{output_dir}/peaks_entropy_filtered/peaks_w_blacklistregion.bed"
- "{output_dir}/peaks_comparison/newly_discovered_peaks.bed"
- "{output_dir}/annotation/annStats_peaks.txt"
- "{output_dir}/figures/fragments_overlap_peaks.png"
- "{output_dir}/figures/fragments_overlap_peaks_scatter.png"
- "{output_dir}/figures/fragments_overlap_entropypeaks_colorentropy.png"
requires:
- auto_process.sh
external_tools:
- Genrich
- HOMER (optional, for annotation)
downstream_analysis.sh:
description: "PIC matrix and PEAKVI clustering"
inputs:
- flag: -d
name: output_dir
required: true
- flag: -b
name: barcode_file
required: true
outputs:
- "{output_dir}/DownstreamReanalysis/filtered_fragments_sorted.tsv.gz"
- "{output_dir}/DownstreamReanalysis/entropy_bc_peak_count.h5"
- "{output_dir}/DownstreamReanalysis/downstream_complete.log"
requires:
- auto_process.sh
- compare_peakcalling.sh
Additional Options (author's choice)
These are additional improvements that could complement the manifest:
Option A: Add --dry-run flag
Scripts could support a dry-run mode that prints expected outputs without executing:
bash auto_process.sh --dry-run -f input.tsv.gz -o /output -g hg38
# Prints: Would create: /output/fragments.tsv, /output/chr1_barcode_entropy_df.tsv, ...
Option B: Input validation at script start
Each script could validate required inputs exist before starting long operations:
# At start of compare_peakcalling.sh
required_files=("${output_dir}/filtered_fragments.tsv" "${output_dir}/Entropy_filtered_bc_CBZ.txt")
for f in "${required_files[@]}"; do
[[ -f "$f" ]] || { echo "ERROR: Missing required input: $f" >&2; exit 1; }
done
Option C: Status files after completion
Write a JSON status file after each script completes:
// .auto_process_status.json
{"status": "success", "outputs_created": [...], "timestamp": "..."}
Questions
- Is
pipeline.yaml an acceptable filename, or would you prefer something else?
- Are there outputs I missed in the example above?
- Would any of the additional options (dry-run, validation, status files) be useful?
Summary
When using this pipeline, it's difficult to know upfront:
Currently this information is scattered across shell scripts and Python files, requiring users to read through the code to understand the full picture. This makes it harder to:
Proposal: Add
pipeline.yamlA single YAML file at the repo root that declares inputs/outputs for each script. This would serve as:
Example structure:
Additional Options (author's choice)
These are additional improvements that could complement the manifest:
Option A: Add
--dry-runflagScripts could support a dry-run mode that prints expected outputs without executing:
bash auto_process.sh --dry-run -f input.tsv.gz -o /output -g hg38 # Prints: Would create: /output/fragments.tsv, /output/chr1_barcode_entropy_df.tsv, ...Option B: Input validation at script start
Each script could validate required inputs exist before starting long operations:
Option C: Status files after completion
Write a JSON status file after each script completes:
Questions
pipeline.yamlan acceptable filename, or would you prefer something else?