Skip to content

Enhancement: Add validation warnings for unusual data patterns #26

Description

@jasegehring

Problem

The pipeline does not warn users when data patterns are unusual or outside expected ranges. Users cannot tell if their results are reasonable or if there are data quality issues that need attention.

User Feedback

From testing notes:

"Warnings for unexpected situations? Like read depth or entropy outside of expected range?"

Current Behavior

Pipeline processes data without validating:

  • Entropy distribution shape
  • Number of cells passing threshold
  • Read depth per cell
  • Fragment length distribution
  • Unusual parameter values

Impact

  • Users cannot tell if results are reasonable
  • Data quality issues go undetected
  • No guidance on whether to trust results
  • Difficult to catch upstream data problems
  • May produce meaningless results without warning

Expected Behavior

Pipeline should provide warnings when:

  1. Data patterns are unusual
  2. Results are outside expected ranges
  3. Parameters may be inappropriate for the dataset

Proposed Warnings

1. Entropy Distribution Warnings

# In calculate_entropy.py after computing all entropies
entropy_values = list(barcode_entropy.values())
median_entropy = np.median(entropy_values)
iqr = np.percentile(entropy_values, 75) - np.percentile(entropy_values, 25)

if median_entropy < 0.5:
    print(f"WARNING: Low median entropy ({median_entropy:.2f}). Data may be low quality or parameters need adjustment.")
    
if iqr < 0.1:
    print(f"WARNING: Very narrow entropy distribution (IQR={iqr:.2f}). All cells may be similar quality.")

2. Cell Filtering Warnings

# In filter_fragments.sh after filtering
cells_passing=$(wc -l < ${filtered_bc_file})
cells_total=$(wc -l < ${entropy_df_file})
percent=$((cells_passing * 100 / cells_total))

if [[ ${percent} -lt 10 ]]; then
    echo "WARNING: Only ${percent}% of cells passed filter. Threshold may be too strict."
elif [[ ${percent} -gt 90 ]]; then
    echo "WARNING: ${percent}% of cells passed filter. Threshold may be too permissive."
fi

if [[ ${cells_passing} -lt 100 ]]; then
    echo "WARNING: Very few cells (${cells_passing}) passed filter. Check input data quality."
fi

3. Read Depth Warnings

# In count_perchrom_tn5_insertions.py
total_insertions = np.sum([record.sum() for record in record.values()])
avg_per_barcode = total_insertions / len(record)

if avg_per_barcode < 100:
    print(f"WARNING: Low average read depth per cell ({avg_per_barcode:.0f} insertions). "
          f"Consider increasing sequencing depth or using larger chromosome.")
    
if avg_per_barcode > 100000:
    print(f"WARNING: Very high average read depth ({avg_per_barcode:.0f} insertions). "
          f"Check for index hopping or aggregation artifacts.")

4. Parameter Warnings

# In auto_process.sh when validating parameters
if [[ ${window_size} -lt 1000 ]] || [[ ${window_size} -gt 10000 ]]; then
    echo "WARNING: Unusual window size (${window_size}). Typical range is 1000-10000 bp."
fi

if [[ ${chromosome} != "chr1" ]] && [[ ${chromosome} != "all" ]]; then
    echo "NOTE: Using non-standard chromosome '${chromosome}'. Chr1 typically provides most stable entropy estimates."
fi

5. Threshold Detection Warnings

# In find_threshold.py
if len(inflection_regions) == 0:
    print("WARNING: No clear inflection point found in entropy distribution. "
          "Manual threshold selection may be needed.")
    
if rank_cutoff < 100:
    print(f"WARNING: Very few cells selected by automatic threshold (rank={rank_cutoff}). "
          f"Consider manual inspection.")

Warning Categories

Data Quality:

  • Low/high read depth
  • Unusual entropy distribution
  • Very few/many cells passing

Parameter Issues:

  • Window size outside typical range
  • Unusual chromosome choice
  • Missing expected files

Method Limitations:

  • No clear threshold detected
  • Insufficient data for reliable estimation
  • Edge cases in statistical model

Implementation Notes

  1. Use consistent format:

    • WARNING: for issues needing attention
    • NOTE: for informational messages
    • Include actionable suggestions
  2. Don't fail pipeline:

    • Warnings inform but don't stop execution
    • Users can choose to ignore or investigate
  3. Add to log file:

    • All warnings should be captured in log
    • Easy to review after pipeline completes
  4. Documentation:

    • Explain what warnings mean
    • Provide guidance on when to worry
    • Link to troubleshooting guide

Expected Ranges (from scATAC-seq best practices)

  • Cells per experiment: 500 - 50,000
  • Reads per cell: 1,000 - 50,000
  • Fragments in peaks: 30-70%
  • Cell filter rate: 20-70% (depends on method)
  • Entropy (log10): Typically 0.5 - 2.0 range

Additional Context

This is especially valuable for:

  • Users new to scATAC-seq
  • Validating unusual experimental protocols
  • Catching upstream data processing errors
  • Publication QC documentation

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions