Problem
The pipeline does not warn users when data patterns are unusual or outside expected ranges. Users cannot tell if their results are reasonable or if there are data quality issues that need attention.
User Feedback
From testing notes:
"Warnings for unexpected situations? Like read depth or entropy outside of expected range?"
Current Behavior
Pipeline processes data without validating:
- Entropy distribution shape
- Number of cells passing threshold
- Read depth per cell
- Fragment length distribution
- Unusual parameter values
Impact
- Users cannot tell if results are reasonable
- Data quality issues go undetected
- No guidance on whether to trust results
- Difficult to catch upstream data problems
- May produce meaningless results without warning
Expected Behavior
Pipeline should provide warnings when:
- Data patterns are unusual
- Results are outside expected ranges
- Parameters may be inappropriate for the dataset
Proposed Warnings
1. Entropy Distribution Warnings
# In calculate_entropy.py after computing all entropies
entropy_values = list(barcode_entropy.values())
median_entropy = np.median(entropy_values)
iqr = np.percentile(entropy_values, 75) - np.percentile(entropy_values, 25)
if median_entropy < 0.5:
print(f"WARNING: Low median entropy ({median_entropy:.2f}). Data may be low quality or parameters need adjustment.")
if iqr < 0.1:
print(f"WARNING: Very narrow entropy distribution (IQR={iqr:.2f}). All cells may be similar quality.")
2. Cell Filtering Warnings
# In filter_fragments.sh after filtering
cells_passing=$(wc -l < ${filtered_bc_file})
cells_total=$(wc -l < ${entropy_df_file})
percent=$((cells_passing * 100 / cells_total))
if [[ ${percent} -lt 10 ]]; then
echo "WARNING: Only ${percent}% of cells passed filter. Threshold may be too strict."
elif [[ ${percent} -gt 90 ]]; then
echo "WARNING: ${percent}% of cells passed filter. Threshold may be too permissive."
fi
if [[ ${cells_passing} -lt 100 ]]; then
echo "WARNING: Very few cells (${cells_passing}) passed filter. Check input data quality."
fi
3. Read Depth Warnings
# In count_perchrom_tn5_insertions.py
total_insertions = np.sum([record.sum() for record in record.values()])
avg_per_barcode = total_insertions / len(record)
if avg_per_barcode < 100:
print(f"WARNING: Low average read depth per cell ({avg_per_barcode:.0f} insertions). "
f"Consider increasing sequencing depth or using larger chromosome.")
if avg_per_barcode > 100000:
print(f"WARNING: Very high average read depth ({avg_per_barcode:.0f} insertions). "
f"Check for index hopping or aggregation artifacts.")
4. Parameter Warnings
# In auto_process.sh when validating parameters
if [[ ${window_size} -lt 1000 ]] || [[ ${window_size} -gt 10000 ]]; then
echo "WARNING: Unusual window size (${window_size}). Typical range is 1000-10000 bp."
fi
if [[ ${chromosome} != "chr1" ]] && [[ ${chromosome} != "all" ]]; then
echo "NOTE: Using non-standard chromosome '${chromosome}'. Chr1 typically provides most stable entropy estimates."
fi
5. Threshold Detection Warnings
# In find_threshold.py
if len(inflection_regions) == 0:
print("WARNING: No clear inflection point found in entropy distribution. "
"Manual threshold selection may be needed.")
if rank_cutoff < 100:
print(f"WARNING: Very few cells selected by automatic threshold (rank={rank_cutoff}). "
f"Consider manual inspection.")
Warning Categories
Data Quality:
- Low/high read depth
- Unusual entropy distribution
- Very few/many cells passing
Parameter Issues:
- Window size outside typical range
- Unusual chromosome choice
- Missing expected files
Method Limitations:
- No clear threshold detected
- Insufficient data for reliable estimation
- Edge cases in statistical model
Implementation Notes
-
Use consistent format:
WARNING: for issues needing attention
NOTE: for informational messages
- Include actionable suggestions
-
Don't fail pipeline:
- Warnings inform but don't stop execution
- Users can choose to ignore or investigate
-
Add to log file:
- All warnings should be captured in log
- Easy to review after pipeline completes
-
Documentation:
- Explain what warnings mean
- Provide guidance on when to worry
- Link to troubleshooting guide
Expected Ranges (from scATAC-seq best practices)
- Cells per experiment: 500 - 50,000
- Reads per cell: 1,000 - 50,000
- Fragments in peaks: 30-70%
- Cell filter rate: 20-70% (depends on method)
- Entropy (log10): Typically 0.5 - 2.0 range
Additional Context
This is especially valuable for:
- Users new to scATAC-seq
- Validating unusual experimental protocols
- Catching upstream data processing errors
- Publication QC documentation
Problem
The pipeline does not warn users when data patterns are unusual or outside expected ranges. Users cannot tell if their results are reasonable or if there are data quality issues that need attention.
User Feedback
From testing notes:
Current Behavior
Pipeline processes data without validating:
Impact
Expected Behavior
Pipeline should provide warnings when:
Proposed Warnings
1. Entropy Distribution Warnings
2. Cell Filtering Warnings
3. Read Depth Warnings
4. Parameter Warnings
5. Threshold Detection Warnings
Warning Categories
Data Quality:
Parameter Issues:
Method Limitations:
Implementation Notes
Use consistent format:
WARNING:for issues needing attentionNOTE:for informational messagesDon't fail pipeline:
Add to log file:
Documentation:
Expected Ranges (from scATAC-seq best practices)
Additional Context
This is especially valuable for: