Arachne is the platform-agnostic successor to the Lariat and EMA aligners for barcoded linked reads. Both were originally written for the 10X Genomics GEMcode platform, which was discontinued in 2019 and Arachne drops support for 10X-style data in favor of a standard linked-read data format. Conversion to standard format from TELLseq, Haplotagging, 10X, and stLFR are provided in Djinn. See the documentation for the full description of Lariat and EMA and the rationale behind Arachne.
- New logo/identity
- Modernize Go idioms
- Replace custom FASTQ reader with
fastx(used by seqkit) - Rewrite internals to match Standard FASTQ format
- Create
preprocesssubcommand - Output SAM to
stdoutinstead of to many files - Create test data
- Get everything to compile and run
- Add build and run tests
- Replace bwa with minibwa (bwa is kept in
archive/) - Expose minibwa index for convenience
- validate output
- user validation
This assumes an environment has already been created with conda create
for
mamba: just swapcondawithmamba
conda install -c bioconda -c conda-forge bioconda::arachneThis assumes a pixi project was already created with pixi init and has the conda-forge and bioconda channels added
pixi add arachneRequires:
- Go installation
- a C compiler and zlib
- a CPU with SSE4.2 (x86-64) or NEON (arm64), which minibwa requires
- optionally OpenMP (
libgomp), used byarachne indexfor multi-threaded index construction. It is detected automatically when building; without it indexing is single-threaded
git clone --recursive https://github.com/pdimens/arachne.git
cd arachne
make # Build arachne
bin/arachne # Show helpAll the dependencies are provided, so it's just a matter of cloning and building.
git clone --recursive https://github.com/pdimens/arachne.git
cd arachne
pixi run buildTL;DR: The only distinction between the 'standard' linked-read FASTQ files and regular FASTQ files
is the presence of the BX:Z and VX:i SAM tags. The format also uses /1 and /2 (the older CASAVA format)
to denote a forward/reverse read. Full description here.
Detailed Explanation
No one wins if everyone is using their own platform-specific file formats. Regardless of the technology used to create the linked reads, Arachne accepts what is called the 'standard' format shown below. This format conforms to the FASTQ and SAM file specs, which are internationally-agreed upon formats, meaning the reads can be used anywhere and doesn't distinguish between barcode formats. This also means it is future-proofed against yet-to-be-invented linked-read technologies, barcode encodings, etc. The trick is the inclusion of two specific SAM-compliant tags: the `BX:Z` tag to denote the barcode and the `VX:i` tag to denote whether the barcode is considered valid for whatever the encoding design is. This means the **location** and **meaning** of the barcodes are always consistent across formats. For example, in TELLseq data, an `N` in a barcode (e.g. `ATGGAGANAA`) indicates the barcode is invalid, so it would inherit a `VX:i` tag of `0` (e.g. `VX:i:0`). For completeness, the 'standard' linked-read FASTQ format follows:| record line | what's in it |
|---|---|
| 1 | Read ID starting with @ and ending with /1 (R1) or /2 (R2). After the read ID, there is TAB followed by any number of tab-delimited SAM tags, but must include BX:Z and VX:i tags |
| 2 | Sequence as ATCGN nucleotides |
| 3 | + sign |
| 4 | PHRED quality scores for nucleotides in line 2 |
BX:Zis the barcode, which is any combination of non-space characters- e.g.
BX:Z:1_2_3,BX:Z:A03C55B49D19,BX:Z:ATTTAGGGAGAGAGA
- e.g.
VX:iis the validation tagVX:i:0= invalid |VX:i:1= valid
@SEQID/1 BX:Z:BARCODE VX:i:0
ATGCGTA.......................
+
FFFFIII.......................
Using a TELLseq-style barcode ATGGAGANAA, where an N indicates it's invalid, the first line of a FASTQ record in the forward read would look like (SAM tag order doesn't matter):
@SEQID/1 BX:Z:ATGGAGANAA VX:i:0
ATGCGTA.......................
+
FFFFIII.......................

