Skip to content

Enhancement: Checkpoint system should detect when input files are updated #21

Description

@jasegehring

Problem

The checkpoint system only checks if output files exist, not whether input files have been modified since the last run. This means if you update input data, the pipeline uses stale cached results without reprocessing.

Location

Multiple checkpoint locations, example: auto_process.sh:84-97

Current Behavior

checkpoint_file=${output_dir}/.finish_format_fragments
if [[ \! -f ${checkpoint_file} ]] ; then
    # Process fragments...
    touch ${checkpoint_file}  # create empty file as checkpoint
fi

Issue

If a user:

  1. Runs pipeline with fragments_v1.tsv.gz → checkpoint created
  2. Gets better data, updates to fragments_v2.tsv.gz
  3. Reruns pipeline → checkpoint exists, so reformatting is skipped
  4. Pipeline uses old processed data with new input file

This is confusing and can lead to using stale results with updated input data.

Impact

  • Research workflows often involve refining input data
  • Users may unknowingly use stale cached results
  • No warning that input is newer than cached output
  • Difficult to debug when results don't match expectations

Expected Behavior

Checkpoint system should detect when input files are newer than output files and reprocess accordingly.

Proposed Fix

Use timestamp comparison with -nt (newer than):

checkpoint_file=${output_dir}/.finish_format_fragments
frag_file=${output_dir}/fragments.tsv

# Reprocess if: checkpoint doesn't exist OR input is newer than checkpoint
if [[ \! -f ${checkpoint_file} ]] || [[ ${fragment_file} -nt ${checkpoint_file} ]] ; then
    echo "Processing fragments file..."
    if [[ ${fragment_file} == *.gz ]]; then
        cat ${fragment_file} | zcat | awk -F '\t' '{if (NF == 5) print $0}' - > ${frag_file}  && \
        touch ${checkpoint_file}
    # ... rest of processing
fi

Locations to Update

Apply this pattern to all checkpoint logic:

  • auto_process.sh:84 - fragment formatting
  • auto_process.sh:112 - entropy calculation
  • auto_process.sh:123 - threshold finding
  • auto_process.sh:139 - fragment filtering
  • calculate_entropy.sh:34 - chromosome splitting
  • calculate_entropy.sh:47 - insertion frequency
  • calculate_entropy.sh:59 - entropy calculation per chromosome

Additional Context

This is standard practice for build systems (Make, etc.) and bioinformatics pipelines. Improves user experience and prevents confusion in iterative research workflows.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions