a Sequence Structured Labeling Framework for Large-Scale Genomic Insertion Detection Leveraging Time-Distributed Dual Transformer
Overall workflow and network architecture of SSL-InsNet for large-scale genomic insertion variant calling.
SSL-InsNet is a novel sequence structured labeling framework engineered for large-scale genomic insertion detection in third-generation long-read sequencing data. The framework reformulates insertion calling as a structured tagging task, integrating a Multi-level Spatial Perception (MSP) module and a Time-distributed Transformer (T-Trans) within a dual-transformer architecture. By incorporating an advanced confidence-gated fusion layer, the network effectively suppresses alignment artifacts to deliver highly precise, chromosome-wide structural variant profiles.
⏳ Sequence Structured Labeling: Reformulates large-scale genomic insertion detection into an optimized sequence structured tagging task across continuous long-read tracks.
🧬 Multi-level Spatial Perception (MSP): Integrates a Spatial Transformer (S-Trans) with deep convolutions via a hierarchical attention mechanism to simultaneously resolve micro-scale breakpoint motifs and regional genomic contexts.
🤝 Confidence-gated Feature Integration (CFI): Implements an elite adaptive Feature Fusion module that leverages learnable descriptors to fuse heterogeneous spatial features, effectively suppressing sequencer background noise and alignment artifacts.
⏳ Time-distributed Transformer (T-Trans): Tracks macro-scale long-distance sequence dependencies along the temporal axis across consecutive segments, ensuring global structural consistency for ultra-long variants under standard GPU footprints.
🛡️ Neighborhood-Weighted Focal Loss (NWFL): Utilizes a custom structural optimization loss function to supervise continuous position tags and drastically refine boundary alignment precision.
- Python 3.9.1
- CUDA 12.4+
- PyTorch 2.5
- 11+ GB VRAM recommended
conda create -n ssl_insnet python=3.9.1 -y conda activate ssl_insnet
pip install torch==2.5.0 torchvision==0.20.0 --index-url https://download.pytorch.org/whl/cu124 pip install pandas==2.2.3 numpy==1.26.4
pip install pysam==0.23.0
| Package | Purpose | Official Repository |
|---|---|---|
| BAM/CRAM file processing and header extraction | GitHub | |
| Window feature metrics management | GitHub | |
| Vectorized genomic array serialization | GitHub |
Extract alignment signatures and base matrices into continuous serialized chunks across target chromosomes.
python SSL_InsNet.py generate_feature bam_file output_path contigs_list(default:[](all chromosomes)) max_worker vcf_file
bam_file: the path of the alignment file about the reference and the long read set;
output_path: a folder which is used to store generated features data;
contigs_list: the list of contig to preform detection. (default: [], all contig are used);
max_worker: the number of threads to use;
vcf_file: the gold standard file for standard data.
eg: python SSL_InsNet.py generate_feature ./HG002_PB_5x_RG_HP10XtrioRTG.bam ./features_dir [12,13] 5 ./HG002_SVs_Tier1_v0.6.vcf.gz
Stream window blocks through the Time-Distributed network to run sequence decoding and merge insertion candidates.
python SSL_InsNet.py call_insertion gpu_name save_length timesteps ins_predict_weight data_path bam_file out_vcf_file contigs support
gpu_name: num of the GPU to use (e.g., '0' or '1,2');
save_length: the feature file spans across nucleotide base sequence lengths;
timesteps: time step of time-distributed network;
ins_predict_weight: path of insert predict weight file (.pth);
data_path: a folder for storing evaluation feature files;
bam_file: path of the alignment file about the reference and the long read set;
out_vcf_file: the path of output vcf file;
contigs: the list of contig to preform detection. (default: [], all contig are used);
support: min support reads.
eg: python SSL_InsNet.py call_insertion '1,2' 10000000 100 ./ins_predict_weight.pth ./features_dir /home/laicx/00.dataset/HG002_PB_10x_RG_HP10XtrioRTG.bam ./out_vcf_file.vcf [12,13] 5
