Problem
The checkpoint system only checks if output files exist, not whether input files have been modified since the last run. This means if you update input data, the pipeline uses stale cached results without reprocessing.
Location
Multiple checkpoint locations, example: auto_process.sh:84-97
Current Behavior
checkpoint_file=${output_dir}/.finish_format_fragments
if [[ \! -f ${checkpoint_file} ]] ; then
# Process fragments...
touch ${checkpoint_file} # create empty file as checkpoint
fi
Issue
If a user:
- Runs pipeline with
fragments_v1.tsv.gz → checkpoint created
- Gets better data, updates to
fragments_v2.tsv.gz
- Reruns pipeline → checkpoint exists, so reformatting is skipped
- Pipeline uses old processed data with new input file
This is confusing and can lead to using stale results with updated input data.
Impact
- Research workflows often involve refining input data
- Users may unknowingly use stale cached results
- No warning that input is newer than cached output
- Difficult to debug when results don't match expectations
Expected Behavior
Checkpoint system should detect when input files are newer than output files and reprocess accordingly.
Proposed Fix
Use timestamp comparison with -nt (newer than):
checkpoint_file=${output_dir}/.finish_format_fragments
frag_file=${output_dir}/fragments.tsv
# Reprocess if: checkpoint doesn't exist OR input is newer than checkpoint
if [[ \! -f ${checkpoint_file} ]] || [[ ${fragment_file} -nt ${checkpoint_file} ]] ; then
echo "Processing fragments file..."
if [[ ${fragment_file} == *.gz ]]; then
cat ${fragment_file} | zcat | awk -F '\t' '{if (NF == 5) print $0}' - > ${frag_file} && \
touch ${checkpoint_file}
# ... rest of processing
fi
Locations to Update
Apply this pattern to all checkpoint logic:
auto_process.sh:84 - fragment formatting
auto_process.sh:112 - entropy calculation
auto_process.sh:123 - threshold finding
auto_process.sh:139 - fragment filtering
calculate_entropy.sh:34 - chromosome splitting
calculate_entropy.sh:47 - insertion frequency
calculate_entropy.sh:59 - entropy calculation per chromosome
Additional Context
This is standard practice for build systems (Make, etc.) and bioinformatics pipelines. Improves user experience and prevents confusion in iterative research workflows.
Problem
The checkpoint system only checks if output files exist, not whether input files have been modified since the last run. This means if you update input data, the pipeline uses stale cached results without reprocessing.
Location
Multiple checkpoint locations, example:
auto_process.sh:84-97Current Behavior
Issue
If a user:
fragments_v1.tsv.gz→ checkpoint createdfragments_v2.tsv.gzThis is confusing and can lead to using stale results with updated input data.
Impact
Expected Behavior
Checkpoint system should detect when input files are newer than output files and reprocess accordingly.
Proposed Fix
Use timestamp comparison with
-nt(newer than):Locations to Update
Apply this pattern to all checkpoint logic:
auto_process.sh:84- fragment formattingauto_process.sh:112- entropy calculationauto_process.sh:123- threshold findingauto_process.sh:139- fragment filteringcalculate_entropy.sh:34- chromosome splittingcalculate_entropy.sh:47- insertion frequencycalculate_entropy.sh:59- entropy calculation per chromosomeAdditional Context
This is standard practice for build systems (Make, etc.) and bioinformatics pipelines. Improves user experience and prevents confusion in iterative research workflows.