Skip to content

Add pipeline.yaml for machine-readable I/O specification #30

Description

@jasegehring

Summary

When using this pipeline, it's difficult to know upfront:

  • What inputs each script requires
  • What outputs each script produces
  • Dependencies between scripts
  • External tool requirements

Currently this information is scattered across shell scripts and Python files, requiring users to read through the code to understand the full picture. This makes it harder to:

  • Integrate the pipeline into larger workflows
  • Debug issues when outputs are missing
  • Know which scripts need to run before others

Proposal: Add pipeline.yaml

A single YAML file at the repo root that declares inputs/outputs for each script. This would serve as:

  • Living documentation that's easy to scan
  • A contract that can be validated against actual outputs
  • A reference for anyone integrating the pipeline into their workflow

Example structure:

# pipeline.yaml
version: "1.0"
description: "Entropy-based cell calling for scATAC-seq"

environment:
  required_tools:
    - name: ripgrep
      command: rg
    - name: samtools
      command: samtools
    - name: bedtools
      command: bedtools
    - name: Genrich
      path_var: Genrich  # from config.sh
  optional_tools:
    - name: HOMER
      command: annotatePeaks.pl
      note: "Requires genome installation (e.g., hg38)"

scripts:
  auto_process.sh:
    description: "Calculate entropy, auto-threshold, filter fragments"
    inputs:
      - flag: -f
        name: fragments_file
        required: true
        format: "*.tsv.gz or *.tsv"
      - flag: -o
        name: output_dir
        required: true
      - flag: -g
        name: species
        required: true
        values: [hg38, mm10]
      - flag: -c
        name: chromosome
        default: chr1
      - flag: -w
        name: window_size
        default: 3000
    outputs:
      - "{output_dir}/fragments.tsv"
      - "{output_dir}/{chromosome}_barcode_entropy_df.tsv"
      - "{output_dir}/entropy_cutoff.csv"
      - "{output_dir}/filtered_fragments.tsv"
      - "{output_dir}/Entropy_filtered_bc_CBZ.txt"
      - "{output_dir}/figures/{chromosome}_entropy_knee_plot_color_crbc.png"
      - "{output_dir}/figures/{chromosome}_entropy_knee_plot_logscale.png"
      - "{output_dir}/figures/{chromosome}_Entropy_histogram.png"
      - "{output_dir}/figures/entropy_violin_EmptyCRcell.png"
      - "{output_dir}/figures/log10_entropy_violin_EmptyvsCell.png"
      - "{output_dir}/figures/frag_length_distribution.png"

  compare_cellcalling.sh:
    description: "Compare entropy vs CellRanger cell calling"
    inputs:
      - flag: -d
        name: output_dir
        required: true
      - flag: -c
        name: chromosome
        required: true
      - flag: -b
        name: cellranger_barcode_file
        required: true
    outputs:
      - "{output_dir}/figures/{chromosome}_entropy_knee_plot_color_crbc.png"
    requires:
      - auto_process.sh

  compare_peakcalling.sh:
    description: "Compare peaks before/after entropy filtering"
    inputs:
      - flag: -d
        name: output_dir
        required: true
      - flag: -g
        name: species
        required: true
      - flag: -s
        name: bam_file
        required: true
      - flag: -c
        name: chromosome
        required: true
    outputs:
      - "{output_dir}/peaks/peaks_w_blacklistregion.bed"
      - "{output_dir}/peaks_entropy_filtered/peaks_w_blacklistregion.bed"
      - "{output_dir}/peaks_comparison/newly_discovered_peaks.bed"
      - "{output_dir}/annotation/annStats_peaks.txt"
      - "{output_dir}/figures/fragments_overlap_peaks.png"
      - "{output_dir}/figures/fragments_overlap_peaks_scatter.png"
      - "{output_dir}/figures/fragments_overlap_entropypeaks_colorentropy.png"
    requires:
      - auto_process.sh
    external_tools:
      - Genrich
      - HOMER (optional, for annotation)

  downstream_analysis.sh:
    description: "PIC matrix and PEAKVI clustering"
    inputs:
      - flag: -d
        name: output_dir
        required: true
      - flag: -b
        name: barcode_file
        required: true
    outputs:
      - "{output_dir}/DownstreamReanalysis/filtered_fragments_sorted.tsv.gz"
      - "{output_dir}/DownstreamReanalysis/entropy_bc_peak_count.h5"
      - "{output_dir}/DownstreamReanalysis/downstream_complete.log"
    requires:
      - auto_process.sh
      - compare_peakcalling.sh

Additional Options (author's choice)

These are additional improvements that could complement the manifest:

Option A: Add --dry-run flag

Scripts could support a dry-run mode that prints expected outputs without executing:

bash auto_process.sh --dry-run -f input.tsv.gz -o /output -g hg38
# Prints: Would create: /output/fragments.tsv, /output/chr1_barcode_entropy_df.tsv, ...

Option B: Input validation at script start

Each script could validate required inputs exist before starting long operations:

# At start of compare_peakcalling.sh
required_files=("${output_dir}/filtered_fragments.tsv" "${output_dir}/Entropy_filtered_bc_CBZ.txt")
for f in "${required_files[@]}"; do
    [[ -f "$f" ]] || { echo "ERROR: Missing required input: $f" >&2; exit 1; }
done

Option C: Status files after completion

Write a JSON status file after each script completes:

// .auto_process_status.json
{"status": "success", "outputs_created": [...], "timestamp": "..."}

Questions

  1. Is pipeline.yaml an acceptable filename, or would you prefer something else?
  2. Are there outputs I missed in the example above?
  3. Would any of the additional options (dry-run, validation, status files) be useful?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions