⚠️ Archived research project — no longer maintained. SCX was built as a research exploration and is not maintained: no further development, bug fixes, or releases should be expected, and issues/PRs will not be reviewed. It is shared as-is for reference — do not adopt it for production or long-term use. Install from GitHub Releases for the last published build.
A purpose-built binary file format for single-cell RNA-seq data. SCX replaces h5ad with 3-7× smaller files (vs uncompressed h5ad), fastest reads at census scale (1.5× faster than Zarr on 1M cells, up to 7× with parallel decode, up to 3.2× parallel write scaling), up to 10× less memory, a GPU-saturating training loader, and a lazy query engine — with native bindings for both Python and R, fully compatible with the scverse ecosystem (scanpy, scVI, AnnData) and Seurat v5.
Python — works with scanpy, scVI, and any scverse tool:
import pyscx
adata = pyscx.open("experiment.scx").to_anndata()
import scanpy as sc
sc.pp.normalize_total(adata)
sc.tl.pca(adata)
sc.tl.leiden(adata)R — works with Seurat v5 and SingleCellExperiment:
library(rscx)
exp <- scx_open("experiment.scx")
seurat_obj <- exp$to_seurat() # Seurat v5 assay
sce <- exp$to_sce() # SingleCellExperimentSee docs/quickstart.md for a runnable end-to-end pipeline
(install → convert → QC → normalize → HVG → PCA → neighbors → UMAP → leiden → markers).
This repo ships a skills/scx-usage/ skill for agent-assisted
work with pyscx and scx-cli. It is task-oriented (not a format spec): install
troubleshooting, conversion recipes, backed/lazy processing, pyscx.accel.*, and ML
loaders. Start at skills/scx-usage/SKILL.md; deeper
reference lives alongside it (reference/installation.md, conversion.md,
processing.md, ml-loading.md).
Helpful when Claude Code is helping you install pyscx (GitHub Release wheels vs source builds, optional extras, common rpath/HDF5/GPU failures) or write usage code without guessing API shapes and gotchas.
To load it in Claude Code from a clone of this repository:
# From the repository root (where .claude-plugin/ and skills/ live)
claude --plugin-dir .The manifest at .claude-plugin/plugin.json registers
the plugin as scx-usage. You can also symlink or copy skills/scx-usage/ into
.claude/skills/ (project-shared) or ~/.claude/skills/ (personal)
so it auto-loads on startup. Other coding agents in this repo are pointed at the same
skill via AGENTS.md.
-
Fast at every scale — on 1M cells SCX is 17× faster than the gzipped h5ad most researchers ship, 1.5× faster than Zarr, and produces a file 4–5× smaller than anndata's default uncompressed h5ad. Single file, BLAKE3-checksummed, mmap-friendly, and HPC-safe: no
HDF5_USE_FILE_LOCKING=FALSEworkaround on NFS / Lustre / GPFS. - Scales on CPU and GPU — shard-level parallelism via rayon delivers up to 7× read and 3.2× write scaling. The GPU path (rapids-singlecell for PCA · kNN · UMAP · preprocessing, cuGraph for Leiden, plus native CUDA kernels for HVG · DE · Harmony) gives 3.8× end-to-end on PCA → kNN → UMAP → Leiden at 1M cells; the training loader hits 1,405 batches/s — 82× faster than TileDB-SOMA-ML.
-
Atlas-scale memory footprint — backed mode +
MADV_DONTNEEDstreaming. A full 1M-cell preprocess-to-cluster pipeline (open → QC → normalize → log1p → HVG → PCA → kNN → UMAP → Leiden) runs at ~11 GB peak RSS vs ~22 GB materialised (51% less; lazy preprocessing alone peaks at ~3.5 GB). Backed mode lets you open a 10M-cell atlas without allocating the full matrix. -
Rust-native analysis accelerators — drop-in replacements for
sc.pp.*/sc.tl.*: PCA, kNN, UMAP, Leiden, differential expression (Wilcoxon and DESeq2-style replicate-aware negative-binomial GLM), pseudobulk, Harmony2 batch integration, cell-by-cell exact LISI, and PFlog (v4) shifted-log normalization with out-of-core baseline-aware PCA. Same scanpy-shaped API, 3–68× faster; every op has adevice="auto"switch that picks GPU when available. -
Foundation model training and tokenisation engine — triple-buffered streaming loader, multi-set batch executor for cell-set models (7.76× faster, 33× fewer allocations), reuse-signal row-group LRU caching, compiled transformer tokenisation kernels (
pyscx.tokenizefor Geneformer, scGPT, UCE, STATE3, CellFM), and spatial neighbourhood plan builders (15,000 plans/s). -
Mutable without rewriting — append new cells, mark-delete doublets, attach cell and gene annotations in-place (
attach_obs_columns,attach_var_columns), compact, merge, or roll back in milliseconds. Append writes new matrix shards in O(new cells); obs metadata is rewritten as a merged Arrow IPC covering all cells (see docs/operations.md). Delete is a logical mask via compressed Roaring Bitmaps, not a data rewrite. -
Multimodal native — CITE-seq, 10x Multiome, and TEA-seq carried in a single v2 file on a shared cell axis. Per-modality codec selection (Scx1 for RNA UMI, Zstd for ADT, Lz4Shuffle/Zstd for ATAC), per-modality
scx merge/compact/subset --modality, and a synchronizedMultimodalTrainingDatasetthat yields cell-aligned RNA+ADT+ATAC batches with zero-copy NumPy handoff. Round-trips withmudata.MuData(Python), Seurat v5 multi-assay, and BioconductorMultiAssayExperiment(R). See docs/multimodal.md. -
Lazy query engine & detection indexing — predicate pushdown skips ~55% of shards on realistic queries; selective reads land in under 5 ms. Detection bitmap sidecars (
BitmapShard) evaluate boolean presence queries and non-zero counts instantly in$O(\text{cardinality})$ without reading CSR values. Condition-grouped sharding (GroupIndex) packs perturbation cohorts contiguously into single shards. -
Drop-in for scverse and Seurat — works with AnnData, scanpy, and scVI (Python), and Seurat v5 + SingleCellExperiment (native R via
rscx, compiled withextendrwithout Python/reticulate). Round-trips cleanly with h5ad, 10x HDF5, and Cell Ranger MTX. -
Cloud-native — streaming push/pull to S3, GCS, and Azure with selective download (only the shards you need) and a direct
open_cloud()path that skips the full download.
h5ad stores integer UMI counts as 32-bit floats. SCX detects this and uses the narrowest integer type that fits (uint8/uint16), then applies domain-specific codecs designed for the statistical properties of count data. Result:
| Dataset | Cells | h5ad | SCX (best) | Best ratio | Read time (SCX vs Zarr) |
|---|---|---|---|---|---|
| PBMC 3K | 2,700 | 21.5 MB | 4.4 MB | 4.9× | 0.04s vs 0.01s |
| Smart-seq2 | 50,000 | 1.07 GB | 350 MB | 3.1× | 0.99s vs 0.37s |
| Tabula Sapiens | 100,000 | 1.59 GB | 322 MB | 4.9× | 0.62s vs 0.56s |
| CELLxGENE Census 1M | 1,000,000 | 11.4 GB | 2.35 GB | 4.8× | 2.9s vs 4.0s |
| CELLxGENE Census 5M | 5,000,000 | 91.4 GB | 12.5 GB | 7.3× | 35s vs 41s |
SCX uses memory-mapped I/O with MADV_DONTNEED page release. During streaming
aggregation (row_sums, col_sums), only one decoded shard is resident at a time.
On a 1M-cell dataset (2.7 GB SCX file):
| h5ad (full load) | SCX (streaming) | Savings | |
|---|---|---|---|
| Peak memory | 11.6 GB | 1.1 GB | 10× |
This means you can stream through atlas-scale datasets without materializing.
SCX supports backed mode — data stays on disk and loads on demand, one shard at a time. Open a 10M-cell atlas without allocating the full matrix:
adata = pyscx.open("atlas.scx").to_anndata(backed=True)
# X stays on disk — only the accessed shard is decoded
# Row slicing leverages whole-shard skipping, bypassing shards with 0 selected cells:
subset = adata[adata.obs["cell_type"] == "T cell"].copy()
# Now subset is a regular AnnData — preprocess normally
# Or use SCX accelerators for compute-heavy steps (3-68× faster at scale)
sc.pp.normalize_total(subset, target_sum=1e4)
sc.pp.log1p(subset)
sc.pp.pca(subset)
# In-flight column projection during eager load (8.3× less peak RAM):
# Extracts HVG/marker panels directly without full-width matrix expansion in RAM
hvg_adata = pyscx.open("atlas.scx").to_anndata(var_names=hvg_genes, obsp=[])Layers, deletion vectors, and anndata.abc.CSRDataset registration all work
transparently. See docs/scanpy/backed-mode.md
for details.
SCX includes a triple-buffered Rust pipeline (tokio I/O → rayon decode → GPU) that keeps your GPU fed. Zero Python on the hot path — all I/O, decompression, shuffling, sparse-to-dense conversion, and normalization happen in compiled Rust.
| SCX | AnnData | TileDB-SOMA-ML | scDataLoader | |
|---|---|---|---|---|
| 1M cells (batches/sec) | 1,405 | 16.3 | 17.1 | 4.4 |
| vs SCX | — | 86× slower | 82× slower | 319× slower |
batch_size=1024, HVG=2000, normalize+log1p (hvg_norm scenario). See detailed results below.
# GPU-saturating training loader — no num_workers needed
dataset = pyscx.TrainingDataset(
"atlas.scx",
batch_size=1024,
hvg_indices=hvg_array, # decode only HVGs → 15× less data
normalize=True,
log1p=True,
)
for batch in dataset:
x = torch.from_numpy(batch["X"]).to(device)
# model.forward(), loss.backward(), ...Single-cell foundation models (Geneformer, scGPT, UCE, STATE3, CellFM) bottleneck heavily on interpreted Python tokenisation loops. SCX provides compiled, lock-free Rust kernels that operate directly on batch tensors with the GIL released:
import pyscx.tokenize as tok
# Rank-based gene tokenisation (Geneformer, TranscriptFormer):
tokens = tok.rank_tokens(batch["X"], max_len=2048)
# Non-zero quantile expression binning (scGPT, CellFM):
binned = tok.bin_values(batch["X"], n_bins=51)
# Expression-weighted stochastic gene sampling (UCE):
sampled = tok.sample_genes(batch["X"], n_sample=1024)
# Top-k feature cropping with sentinel padding:
cropped = tok.top_k(batch["X"], k=1024, pad_token_id=0)For cell-set transformers and spatial graph neural networks, pyscx.SparseCellSetDataset
uses a multi-set batch executor that pre-allocates exact CSR batches, caches recurrent
control groups, and parallelizes row-level transformations across Rayon threads — yielding
a 7.76× throughput gain and slashing heap allocations from 2,698 to 81 per batch:
# Gather variable-length cell sets across files with zero-copy NumPy handoff
dataset = pyscx.SparseCellSetDataset("atlas.scx", plan=cell_set_plan)
# Synthesize 15,000 spatial microenvironment plans/sec directly from coordinates:
spatial_plans = pyscx.neighborhood_plans_from_coords(coords, k=15)For ML workloads where each batch is a list of (perturbed_cell, control_cell)
pairs (perturbation training, contrastive learning, donor-matched designs),
pyscx.IndexPlanDataset is the sibling row-source: consumer-supplied plans,
paired dense {X, X_paired, pairs, obs, obs_paired} batches, optional
shard prefetch lookahead. Plan-driven access is intentionally random — at
1M cells it reaches 20K cells/s (3.7× slower than the sequential
TrainingDataset ceiling) but is 106× faster than the cell-load-scx
ScxBackedSparseDataset Python-loop baseline. See
docs/api/python-training.md § IndexPlanDataset.
Fork-safe under DataLoader(num_workers > 0) when the dataset is
constructed lazily inside the worker's __iter__ — see
docs/api/python-training.md § Fork safety under PyTorch DataLoader
for the recommended IterableDataset wrapper, do/don't list, and the
rayon-pool gotcha (the durable regression lives in
pyscx/tests/test_fork_safety.py).
SCX includes a lazy query engine with two-level predicate pushdown. Filter by cell type, tissue, donor — SCX skips entire shards that can't match, reading only the data you need.
result = (pyscx.open("atlas.scx")
.query()
.filter_obs("cell_type == 'T cell' and tissue == 'lung'")
.select_genes(hvg_indices)
.with_normalize(1e4)
.with_log1p()
.collect())
adata = result.to_anndata() # only matching cells, zero-copy
print(f"Skipped {result.skipped_shards}/{result.total_shards} shards")
# Instant O(cardinality) boolean queries via Detection Bitmaps (Roaring Bitmaps):
# Evaluates presence queries without scanning CSR index or value arrays
exp = pyscx.open("atlas.scx")
expressing_cells = exp.cells_expressing("CD3D")
counts = exp.detection_counts(["CD3D", "MS4A1", "GNLY"])
# Condition-grouped physical locality (scx sort --group-by):
# Reads an entire perturbation cohort in 1 contiguous disk read
cohort = exp.read_group("BRCA1_sg1")
# Lossless typed queries (avoids float32 mantissa loss for counts > 2^24):
exact_counts = exp.query().collect(data_dtype="uint32")Average shard skip rate: 55% on realistic queries. Selective query latency: <5 ms.
SCX supports in-place annotation, append, delete, compact, merge, and rollback — no need to rewrite multi-gigabyte expression matrices when adding cell annotations or removing doublets.
# In-place annotation attachment (instant, zero matrix rewrite, automatic index updates)
pyscx.attach_obs_columns("atlas.scx", cell_annotations_df)
pyscx.attach_var_columns("atlas.scx", gene_annotations_df)
# Append new cells (new CSR shards at EOF; obs metadata rewritten for all cells)
pyscx.append("atlas.scx", "new_batch.scx")
# Logical deletion via compressed Roaring Bitmaps (instant O(1), no data rewrite)
pyscx.mark_deleted("atlas.scx", doublet_indices)
# Reclaim space when convenient (drops orphaned sections, canonicalizes shards)
pyscx.compact("atlas.scx", "atlas_clean.scx")
# Oops? Roll back to the previous version instantly (header pointer update)
pyscx.rollback("atlas.scx")Categorical columns retain strict dictionary fidelity, ordering, and unused factor levels
through all mutations — never demoted to plain strings. Carry invariant enforcement
(carry.rs) guarantees that auxiliary mappings (obsm, obsp,
varm, varp) and sidecars are never silently lost during mutating operations.
A single SCX v2 file carries multiple modalities (RNA + ADT + ATAC + …) on a shared cell axis. Each modality picks its own codec — Scx1 for RNA UMI counts, Zstd for ADT, Lz4Shuffle/Zstd for ATAC peaks — instead of forcing one compression scheme across feature spaces with very different statistics.
import mudata
import pyscx
# CITE-seq: write a MuData(rna, adt) into one .scx file
mu = mudata.read_h5mu("citeseq.h5mu")
pyscx.from_mudata(mu, "citeseq.scx") # codec="auto" picks per modality
# Read back as MuData, or extract one modality as AnnData
reader = pyscx.open("citeseq.scx")
mu = reader.to_mudata() # full MuData
rna = reader.to_anndata(modality="rna") # single modality
# Train cell-aligned RNA+ADT batches in one pass
ds = pyscx.MultimodalTrainingDataset(
"citeseq.scx", modalities=["rna", "adt"], batch_size=1024,
)All file-mutation ops route per modality: scx merge, scx compact,
scx subset --modality NAME --filter "...", and scx append (which
preserves per-modality CSC sidecars). Streaming h5mu ↔ SCX conversion is
the default in both directions — pyscx.from_h5mu / pyscx.to_h5mu and
scx convert --from h5mu / --to h5mu all bound peak RSS to one shard
per matrix. R support covers Seurat v5 multi-assay (exp$to_seurat())
and Bioconductor MultiAssayExperiment round-trips.
See docs/multimodal.md for the full Python / R / CLI surface, the format model (format.md § 13), and per-modality codec defaults (codec.md § Per-modality codec defaults).
SCX provides streaming push/pull with S3, GCS, and Azure — no intermediate files. Selective pull downloads only matching shards, saving bandwidth on atlas-scale data.
# Stream from GCS → local packed file (parallel downloads)
pyscx.pull("gs://bucket/atlas.scxd/", "atlas.scx")
# Selective pull — only download shards containing T cells
# (shard-granular: output may include extra cells from partially matching shards)
pyscx.pull("gs://bucket/atlas.scxd/", "t_cells.scx",
filter="cell_type == 'T cell'")
# Direct cloud reads (no full download needed)
exp = pyscx.open_cloud("gs://bucket/atlas.scxd/")
print(exp.n_obs, exp.n_vars)At >1M cells, even optimized CPU code for PCA, kNN, and UMAP takes minutes, and batch integration or replicate-aware differential expression can stall for hours.
SCX provides GPU-accelerated analysis via rapids-singlecell — PCA (rsc.pp.pca), kNN (rsc.pp.neighbors), and UMAP (rsc.tl.umap) run in-VRAM on the GPU, with native CUDA and Rust kernels for HVG, DE, Harmony2, LISI, and PFlog.
All accessed through the same Python API with a single device="auto"|"cpu"|"gpu" parameter:
import pyscx
adata = pyscx.open("atlas.scx").to_anndata(backed=True)
# GPU-accelerated pipeline — up to 16× faster per-op, 3.8× end-to-end on 1M cells
pyscx.accel.pca(adata, n_comps=50, device="gpu")
pyscx.accel.neighbors(adata, n_neighbors=15, device="gpu")
pyscx.accel.umap(adata, device="gpu")
# Clean-room Harmony2 batch integration (17× CPU / 68× GPU speedup):
pyscx.accel.harmony_integrate(adata, key="batch")
# Cell-by-cell exact LISI diversity scores (14×–114× faster, exact harmonypy parity):
pyscx.accel.compute_lisi(adata, keys=["batch", "cell_type"])
# Replicate-aware negative-binomial GLM DE (native Rust DESeq2 equivalent):
res = pyscx.accel.pseudobulk_dex(adata, groupby="cell_type", condition="disease", backend="nb_glm")
# PFlog (v4) shifted-log normalization on raw counts (avoids depth-division bias):
# Runs out-of-core PCA on 1M cells using 1.4 GB RAM instead of 246 GB dense OOM:
pyscx.accel.pflog_pca(adata, n_comps=50)
import scanpy as sc
sc.tl.leiden(adata) # downstream scanpy works identically
sc.pl.umap(adata, color="leiden")SCX provides to_gpu_anndata() for minimal-copy device handoff to a
GPU-resident AnnData with cupyx.scipy.sparse.csr_matrix X, enabling
full in-VRAM pipelines via rapids-singlecell. When no GPU is available,
every operation falls back to CPU automatically with a warning.
SCX ships Rust-accelerated equivalents of the metrics in
cell-eval — pseudobulk means,
bundled bulk metrics (pearson_delta / mse / mae / mse_delta / mae_delta),
discrimination score, energy distance, knockdown efficiency, and clustering
agreement. Output is numerically equivalent to the Python reference
(32/32 parity tests pass), so you can swap in pyscx.accel.* without changing
the rest of your pipeline.
# One call replaces cell-eval's pearson_delta + mse + mae + mse_delta + mae_delta
results = pyscx.accel.perturbation_metrics(adata_real, adata_pred)
# Per-cell knockdown efficiency vs control
pyscx.accel.knockdown_efficiency(adata, pert_col="perturbation", control="control")
# → adata.obs["KnockDownEfficiency"], adata.obs["KnockDownGeneFC"]
# Energy distance: faer-gemm + f32 default (or bounded Gram-block GPU kernel):
# Evaluates >50,000 control cells on GPU without CUDA Out-of-Memory (<512 MB VRAM):
corr = pyscx.accel.energy_distance(adata_real, adata_pred, device="gpu")
# Clustering agreement: native-Rust HNSW + Leiden (no scanpy under the hood).
score = pyscx.accel.clustering_agreement(adata_real, adata_pred, metric="ami")| Operation | 20K × 2K × 50 ² | 100K cells | 500K cells | 1M cells |
|---|---|---|---|---|
| Pseudobulk means | 11.8× | 11.6× | 13.8× | 19.4× |
| Bulk metrics (5 bundled) | 10.5× | 12.1× | 13.6× | 21.9× |
| Discrimination score | 11.8× | 12.0× | 12.9× | 20.1× |
| Energy distance (gemm + f32, default) | 52.1× | — ² | — ¹ | — ¹ |
| Energy distance (scalar + f64, legacy) | 10.2× | 14.4× | — ¹ | — ¹ |
| Clustering agreement (native Rust Leiden) | 3.0× ³ | 10–13× | 24.6× | 10.0× |
¹ Reference's sklearn.metrics.pairwise_distances doesn't scale above 100K.
² 20K column is the canonical measurement (captured 2026-04-27 with the Phase 1+2 (backend, dtype) matrix). 100K–1M columns are pre-Phase-1 historical baselines.
³ Speedup grows with n_perts × embedding-dim; at 200 perts × 300 genes the ratio is 12.8×.
See docs/scanpy/accel-perturbation-metrics.md
for the full API and docs/performance/perturbation-metrics.md
for the benchmark methodology.
HDF5 acquires mandatory POSIX file locks on every open — even for reads. On shared and network filesystems (NFS, Lustre, GPFS) this causes the dreaded:
OSError: Unable to open file (unable to lock file, errno = 37, error message = 'No locks available')
Common workarounds (HDF5_USE_FILE_LOCKING=FALSE, rebuilding with --disable-file-locking)
disable data integrity checks entirely. Jupyter notebooks that hold an h5ad open will
block other processes from reading the same file.
SCX takes a different approach:
| h5ad (HDF5) | SCX | |
|---|---|---|
| Read locking | Mandatory — blocks other readers/writers | None — reads never lock |
| Write locking | Mandatory — blocks all other access | Advisory flock() — only during append/delete/compact |
| Concurrent reads | ❌ Blocked if any writer is active | ✅ Always allowed, even during writes |
| Network filesystems | Frequently broken (errno 37) |
Works — advisory locks degrade gracefully |
| Workaround needed? | HDF5_USE_FILE_LOCKING=FALSE |
No workaround needed |
SCX's immutable-fragment design (append-only sections + atomic header update) means readers
always see a consistent snapshot without any locking. Multiple notebooks, pipeline stages,
or training jobs can read the same .scx file simultaneously — no coordination required.
SCX splits the expression matrix into CSR shards — fixed-size, independently
decompressible chunks of rows (default: 10,000–16,384 cells per shard). Sharding
enables parallel decoding, memory-bounded reads, shard-level predicate pushdown,
append-without-rewrite, and selective cloud downloads. Control shard size via
--shard-size (CLI) or shard_size= (Python/R).
See docs/sharding.md for a full guide including sizing
guidelines, CLI/Python/R commands, and how sharding powers each SCX feature.
SCX exploits multicore CPUs at every stage. Rayon decodes shards in parallel
during reads (up to 7× at 32 threads) and encodes shards in parallel during
writes (up to 3.2× at 32 threads). The training loader runs a triple-buffered
pipeline (tokio I/O → rayon decode → Python/GPU) so that I/O, decompression,
and GPU transfer all overlap. Cloud downloads run as concurrent async tasks.
File mutations are serialized with advisory flock() locks — reads never lock.
| Component | Threading model | Runtime |
|---|---|---|
| Shard decode (read) | Data parallelism | Rayon par_iter |
| Shard encoding (write) | Data parallelism | Rayon par_iter |
| Query engine | Parallel shard decode + filter | Rayon |
| Training loader | Triple-buffered pipeline | tokio + rayon + std::thread |
| Cloud I/O | Parallel async downloads/uploads | tokio |
| File mutations | Advisory file locks | fs4 flock() |
| GPU decode | Massively parallel kernels | CUDA |
See docs/multithreading.md for a full guide
including pipeline architecture, thread safety of key types, and how to
control parallelism.
Every existing single-cell format has significant trade-offs that SCX was designed to avoid:
The scverse standard. Ubiquitous but showing its age at atlas scale.
- File locking breaks on HPC. HDF5 uses mandatory POSIX
flock(), which fails on NFS, Lustre, and GPFS — the most common HPC parallel filesystems. Users must setHDF5_USE_FILE_LOCKING=FALSE, disabling integrity checks entirely. - Single-threaded reads in Python. h5py holds a global lock and does not release the GIL during HDF5 calls, so multithreaded reads gain zero parallelism.
- Integer counts stored as float32. AnnData stores UMI counts as 32-bit floats by default, wasting 2-4× space for data that fits in uint8/uint16.
- Slow or weak compression. Default gzip is slow to decompress; lzf is fast but achieves poor ratios. Zstd requires the third-party
hdf5pluginand is not natively supported. - No cloud-native access. HDF5 metadata is scattered throughout the file, requiring many small range requests on S3/GCS. The S3 VFD is read-only and limited.
- No append without rewrite. Adding cells or modifying obs/var requires rewriting the entire file. Backed mode (
r+) only supports updating X values in-place.
Cloud-native array storage. Good for object stores, problematic on local/HPC filesystems.
- Multi-file directory structure. Each chunk is a separate file on disk. A 1M-cell dataset produces tens of thousands of files, hitting inode quotas on Lustre/GPFS and causing heavy metadata server load.
- No atomic writes. A Zarr store is a directory tree — interrupted writes leave partial/corrupt state with no rollback mechanism.
- No integrity verification. No built-in checksums, tree hashing, or corruption detection. There is no way to validate a Zarr store's integrity after transfer or filesystem errors.
- v2/v3 ecosystem fragmentation. Zarr v3 is a breaking change (new metadata format, new codec pipeline). The anndata Zarr backend still defaults to v2; v3 sharding support is immature and significantly slower in practice.
- No query or filter capability. Zarr provides array-level chunk access only — no predicate pushdown, no cell/gene filtering without reading full chunks.
- Filesystem overhead on HPC. The one-file-per-chunk design suits object stores (S3) but penalizes local and HPC filesystems with per-file open/close syscall costs, directory traversal, and block alignment waste.
CELLxGENE Census standard. Powerful for cloud queries, heavy for everything else.
- Multi-file directory structure. Like Zarr, each TileDB array is a directory with many internal fragment files. On HPC shared filesystems, metadata operations (open, stat, list) are slow due to inode pressure. Fragment proliferation after repeated writes requires periodic consolidation — an operational burden absent from single-file formats.
- Deep dependency stack. TileDB-SOMA depends on TileDB Core (C++), libtiledbsoma, PyArrow, and multiple SOMA API layers (~5 layers deep). Building from source requires CMake and C++17. Conda packages frequently lag or conflict with RAPIDS/CUDA environments.
- Slow for simple operations. Opening a SOMA experiment requires listing and reading all fragment metadata — cold opens on networked storage take 5-30 seconds for large experiments. Simple "read all X into memory" is slower than h5ad for datasets under ~500K cells.
- Complex API. Reading a matrix requires navigating Experiment → Collection → Measurement → X["raw"] with Arrow table intermediaries, vs a single
sc.read_h5ad(path)call. - Storage bloat before consolidation. Fragment-based writes cause 1.5-3× storage bloat until consolidated. Each mutation creates a new fragment rather than updating in place.
- Limited scanpy integration. Converting SOMA to AnnData for scanpy/scVI typically materializes the full dataset, negating lazy-read benefits.
SQL-native lazy format backed by Lance columnar files + a DuckDB query layer. Compelling design idea for obs-driven analytics, but heavy at realistic read sizes in our benchmarks.
- Slow full-materialize path.
LazyAnnData.compute()goes through Polars fragment processors and builds the CSR in Python-owned memory. Full read of Census 1M takes 53 s vs 2.7 s for SCX and 4.0 s for Zarr (lz4); peak RSS during the materialization reached 35 GB, an order of magnitude more than any other format we tested. - String
cell_id/gene_idis load-bearing. The publicget_submatrixAPI returns long-form(cell_id, gene_id, value)with string ids. In real-world datasets (census, cell barcodes, etc.) those strings are not unique across the file's Lance fragments, so joining back to integer positions explodes row counts — consumers must go through the internalexpression(cell_integer_id, gene_integer_id, value)SQL table to get a scipy CSR correctly. - Non-standard selector semantics.
get_submatrix(cell_selector=[…])treats an integer list as positional but returns string ids that don't always match the list (e.g. cell at position 201469 hascell_id = "1469"in census_1m). Integer numpy arrays, string lists, and mixed-type selectors all fail in different ways. - PyTorch loader stalls at atlas scale.
SLAFDataLoader's Mixture-of-Scanners prefetcher produced 0 batches on Census 10M in our benchmark with the default config (90 s TTFB, then timeout). On Census 1M it ran cleanly at ~4 batches/sec across raw/hvg/norm/hvg_norm — roughly 340× slower than SCX, and about 4× slower thanAnnDatain-memory iteration. - Heavy on-disk, small-file directory layout. A 1M-cell dataset expands
to 4.0 GB vs 2.4 GB for SCX, and the
.slafdirectory holds hundreds of Lance fragment + statistics files — same HPC-filesystem inode pressure problem as Zarr / TileDB-SOMA. - SQL predicate pushdown is a real strength. For
cell_type == 'T cell'on Census 1M, SLAF's SQL path returns the matching expression records in 10 s — same order of magnitude as SCX's catalog pushdown.
| Issue | h5ad | Zarr | TileDB-SOMA | SLAF | SCX |
|---|---|---|---|---|---|
| Single file | Yes | No (directory) | No (directory) | No (directory) | Yes |
| HPC filesystem friendly | No (flock) | No (inode flood) | No (inode flood) | No (inode flood) | Yes (mmap, advisory locks) |
| Atomic writes | No | No | Fragment-based | Fragment-based | Yes (atomic rename) |
| Integrity verification | Partial | None | Per-fragment | Per-fragment |
Full (BLAKE3: catalog verified on open; validate() re-hashes all section payloads) |
| Parallel reads | No (GIL) | Chunk-level | Tile-level | Fragment-level | Shard-level (rayon) |
| Append without rewrite | No | No | Yes (fragments) | Yes (fragments) | Yes (append sections) |
| In-place annotation attachment | No (full rewrite) | Fragmented | Fragment bloat | Fragment bloat |
Yes (attach_obs/var_columns) |
| Transactional rollback | No | No | No | No |
Yes (instant |
| Cloud-native access | No | Yes | Yes | Yes | Yes (single-GET root, selective pull) |
| Built-in query engine | No | No | Yes | Yes (SQL) | Yes (predicate pushdown + Roaring bitmaps) |
| Domain-specific compression | No | No | No | No | Yes (Scx1 & ShufDeltaZstd) |
| Integer-aware storage | No (float32) | No (float32) | No (float64) | u16 per-cell | Yes (lossless uint8/16/32 auto-detect) |
| ML training loader | No | No | Yes (tiledbsoma-ml) | Yes (slow) | Yes (1,405 batches/s, zero-copy) |
| Foundation model tokenisation | No | No | No | No |
Yes (compiled pyscx.tokenize) |
| Out-of-core / backed reads | Partial (r+) | Partial | Partial | No (full materialize) | Yes (streaming + backed) |
| Normalization without densifying | No | No | No | No | Yes (PFlog baseline-aware PCA) |
| Native dual Python + R stack | Python only | Python only | C++ wrapped | Python only |
Yes (native pyscx + rscx) |
# Create a virtual environment
uv venv .venv
# Install pyscx and dependencies
uv pip install maturin numpy scipy pyarrow anndata
# Build from source
cd pyscx && ../.venv/bin/maturin develop --release
# With cloud support (S3/GCS/Azure):
cd pyscx && ../.venv/bin/maturin develop --release --features hdf5,cloudGPU support requires the CUDA Toolkit (≥ 12.0) and rapids-singlecell (which brings cuML, cuGraph, and cupy). There are three ways to set this up — conda is recommended as it handles the full CUDA + RAPIDS dependency tree.
Option A: conda (recommended) — resolves CUDA version matching automatically:
# Create a dedicated GPU environment
conda create -n scx-gpu python=3.13
conda activate scx-gpu
# Install rapids-singlecell + RAPIDS (cuGraph for Leiden) — pin cuda-version to match your driver
# Run `nvidia-smi` to check your driver's max CUDA version
conda install -c rapidsai -c conda-forge rapids-singlecell cugraph cuda-version=12.2
# Install Python deps + build pyscx with GPU support
pip install maturin numpy scipy pyarrow anndata scanpy scikit-learn leidenalg
cd pyscx && maturin develop --release --features hdf5,gpuOption B: system CUDA Toolkit — for native-only GPU ops (HVG, DE, Harmony, Leiden via cuGraph):
# 1. Install CUDA Toolkit ≥ 12.0
# Ubuntu/Debian:
# wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
# sudo dpkg -i cuda-keyring_1.1-1_all.deb
# sudo apt update && sudo apt install cuda-toolkit-12-2
# Or see: https://developer.nvidia.com/cuda-downloads
# 2. Ensure nvcc is on PATH
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version # should print CUDA 12.x
# 3. Build pyscx with GPU support
uv venv .venv
uv pip install maturin numpy scipy pyarrow anndata
cd pyscx && ../.venv/bin/maturin develop --release --features hdf5,gpuWithout rapids-singlecell, PCA/kNN/UMAP will fall back to CPU. Native CUDA kernels (HVG, DE, Harmony) still run on GPU.
Option C: container — for reproducible environments or CI:
# Build the GPU image (multi-stage: compiles Rust + CUDA kernels, then slim runtime)
docker build -f Dockerfile.gpu -t scx-gpu .
# Run with GPU access
docker run --gpus all -it scx-gpu
docker run --gpus all -v /data:/data scx-gpu python my_analysis.pyVerifying the installation:
import pyscx
# Check GPU availability
print(pyscx.accel.gpu_available()) # True if CUDA device found
# Run with explicit GPU — warns and falls back to CPU if unavailable
pyscx.accel.pca(adata, n_comps=50, device="gpu")See docs/gpu-setup.md for troubleshooting, SLURM
configuration, and driver compatibility details. For install pitfalls and
GitHub-Release-wheel-vs-source guidance aimed at agents, see
skills/scx-usage/reference/installation.md.
Pre-built binaries (recommended) — published on each scx-cli-v* tag at GitHub Releases. Linux only (x86_64 and arm64), glibc ≥ 2.35 (Ubuntu 22.04+, Debian 13+, RHEL 10+). Bundles hdf5 (h5ad conversion) and cloud (S3/GCS/Azure); libhdf5 is statically linked so no system libraries are required at runtime.
# Last published release — see https://github.com/ArcInstitute/scx/releases
# for whether a newer one ever shipped
VERSION=0.20.0
# Pick the matching target for your platform:
# linux x86_64 → x86_64-unknown-linux-gnu
# linux arm64 → aarch64-unknown-linux-gnu
TARGET=x86_64-unknown-linux-gnu
# Requires the GitHub CLI (https://cli.github.com/) and `gh auth login`
# while the repo is private.
gh release download "scx-cli-v${VERSION}" -R ArcInstitute/scx \
-p "scx-cli-${VERSION}-${TARGET}.tar.gz"
tar xzf "scx-cli-${VERSION}-${TARGET}.tar.gz"
./scx-cli-${VERSION}-${TARGET}/scx --version
# Move onto your PATH:
install -m 0755 "scx-cli-${VERSION}-${TARGET}/scx" ~/.local/bin/scxIf the repository is public, the asset can also be fetched without gh:
curl -L "https://github.com/ArcInstitute/scx/releases/download/scx-cli-v${VERSION}/scx-cli-${VERSION}-${TARGET}.tar.gz" | tar xzBuild from source — for macOS, Windows, musl, or custom feature sets:
# Build the CLI tool (binary is named `scx`; crate is `scx-cli`).
cargo build -p scx-cli --release
# End-user install from a clone — bundles h5ad/h5mu/10x conversion.
# (The crate is not on crates.io yet, so `cargo install scx-cli` does not work;
# install from the local checkout with --path.)
cargo install --path scx-cli --features default-bin
# With h5ad conversion support (requires libhdf5-dev):
cargo build -p scx-cli --release --features hdf5
# With cloud operations:
cargo build -p scx-cli --release --features hdf5,cloudRequires a Rust toolchain (rustc ≥ 1.78 and cargo). Install via rustup.
# From the repository root:
R CMD INSTALL rscx/
# Or, from within R:
devtools::install_local("rscx/")The package compiles the Rust workspace during installation (handled by src/Makevars).
No pre-built binaries are distributed — Cargo builds all SCX crates from source.
System requirements:
- Rust toolchain:
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh - R ≥ 4.2.0
- R packages:
Matrix,methods(required);Seurat≥ 5.0,SingleCellExperiment(optional)
For an end-to-end walkthrough, see the scanpy tutorial notebook.
SCX supports roundtrip conversion with h5ad, 10x HDF5, and Cell Ranger MTX formats — convert in, work with SCX, convert back out.
The h5ad ↔ SCX and h5mu ↔ SCX paths stream by default — peak RSS is
bounded by one shard's worth of CSR per matrix (and one shard's worth
of obs/var per column when the source carries ObsMetadataShard /
VarMetadataShard sections) regardless of total file size. Pass
--stream=false (CLI) or stream=False (Python) to opt into the
legacy materialising paths. 10x → SCX also streams by default. MTX ↔ SCX
has a single path in each direction, with nothing for --stream to select:
SCX → MTX always streams (one decoded shard at a time), MTX → SCX always
materialises. Just omit --stream and every direction does the right thing.
# h5ad ↔ SCX (roundtrip, streaming by default)
scx convert experiment.h5ad experiment.scx
scx convert --to h5ad experiment.scx experiment.h5ad
# h5mu ↔ SCX (multimodal, streaming by default)
scx convert experiment.h5mu experiment.scx
scx convert --to h5mu experiment.scx experiment.h5mu
scx convert --to h5ad experiment.scx rna.h5ad --modality rna # extract one modality
# Cell Ranger MTX ↔ SCX (roundtrip)
scx convert /path/to/filtered_feature_bc_matrix/ experiment.scx
scx convert --to mtx experiment.scx /path/to/output_dir/
# …and one modality of a multimodal file (required there — an MTX directory
# holds one matrix over one feature space)
scx convert --to mtx multiome.scx /path/to/rna_dir/ --modality rna
# 10x HDF5 → SCX
scx convert filtered_feature_bc_matrix.h5 experiment.scximport pyscx
# From AnnData object
pyscx.from_anndata(adata, "experiment.scx")
# From 10x HDF5
pyscx.from_10x("filtered_feature_bc_matrix.h5", "experiment.scx")
# From Cell Ranger MTX directory
pyscx.from_mtx("/path/to/filtered_feature_bc_matrix", "experiment.scx")
# Streaming export back to h5ad / h5mu
pyscx.to_h5ad("experiment.scx", "experiment.h5ad")
pyscx.to_h5mu("experiment.scx", "experiment.h5mu")
pyscx.to_h5ad("experiment.scx", "rna.h5ad", modality="rna") # extract one modality
# Export back to MTX (streams shard by shard)
pyscx.to_mtx("experiment.scx", "/path/to/output_dir")
pyscx.to_mtx("multiome.scx", "/path/to/rna_dir", modality="rna")adata = pyscx.open("experiment.scx").to_anndata()
# → standard AnnData with X, obs, var, obsm, layers, unsimport pyscx
adata = pyscx.open("atlas.scx").to_anndata()
# Auto-detect GPU (falls back to CPU if unavailable)
pyscx.accel.pca(adata, n_comps=50) # device="auto" by default
pyscx.accel.neighbors(adata, n_neighbors=15) # device="auto" by default
pyscx.accel.umap(adata) # device="auto" by default
# Force GPU
pyscx.accel.pca(adata, n_comps=50, device="gpu")
# Multi-GPU selection
pyscx.accel.pca(adata, n_comps=50, device="gpu:1")
# → standard AnnData with X, obs, var, obsm, layers, unsresult = (pyscx.open("experiment.scx")
.query()
.filter_obs("cell_type == 'B cell'")
.collect())
adata = result.to_anndata()scx info experiment.scx # file metadata
scx validate experiment.scx # verify checksums
scx query experiment.scx --filter "tissue == 'lung'" --count # or positionally: scx query experiment.scx "tissue == 'lung'"
scx append atlas.scx batch2.scx
scx merge batch1.scx batch2.scx --output atlas.scxHeadline numbers at Census 1M (CELLxGENE Census, 1M cells):
| Area | SCX | Next best | SCX advantage |
|---|---|---|---|
| File size vs uncompressed h5ad | 2.35 GB | 11.4 GB | 4.8× smaller |
| Read (full load to AnnData) | 2.74 s | 3.99 s (Zarr lz4) | 1.5× faster |
| Column projection (2K HVGs) | 3.53 s | 7.24 s (Zarr lz4) | 2.0× faster |
| Parallel read (32 threads) | 3.0 s | — (no other format scales) | 6.1× vs 1 thread |
| Parallel write (32 threads, pcodec) | 11.4 s | — | 3.2× vs 1 thread |
| Out-of-core pipeline peak RSS | ~11 GB | ~22 GB (materialized) | 51% reduction |
| Training loader (batches/s) | 1,405 | 17.1 (TileDB-SOMA-ML) | 82× faster |
| GPU end-to-end pipeline (H100) | 286 s | 1,077 s (CPU) | 3.8× faster |
| Selective query (55% shard skip) | 4.2 ms | — | — |
| Append 10K cells | 1 ms | — | — |
Full benchmark suite in docs/performance/: compression and read/write/conversion timings across h5ad / Zarr / TileDB-SOMA / SLAF, parallel read and write scaling, column projection, memory (peak RSS and out-of-core), CPU analysis accelerators (PCA / DE / Leiden), Harmony2 + LISI scaling, perturbation metrics (cell-eval parity), GPU codec and pipeline breakdowns, training loader across datasets, query engine, and file operations. Every number is backed by a manifest entry in benchmarks/comprehensive/results/ — see docs/benchmark_manifest.md for the schema and verification workflow.
SCX is a Rust workspace with 16 crates:
| Crate | Purpose |
|---|---|
scx-format |
Pure on-disk layout/spec: header, catalog, shard structs, checksums (no I/O) |
scx-format-io |
Runtime reader/writer, backed/streaming access, shard codec dispatch, sidecars |
scx-codec |
Domain-specific codecs: Rice, FOR-BP, Delta-Golomb, Zstd, LZ4+shuffle, Pcodec, byte-shuffle |
scx-sparse |
CSR matrix type (scipy-compatible) |
scx-ops |
Append, delete, compact, merge, rollback |
scx-engine |
Lazy query engine with predicate pushdown |
scx-loader |
Triple-buffered ML training data loader |
scx-accel |
Rust-native analysis accelerators: PCA, kNN, UMAP, DE, pseudobulk |
scx-gpu |
CUDA-accelerated codec decoding, cuSPARSE interop, GPU sparse-to-dense |
scx-cloud |
Cloud access: push, pull, explode, pack, CloudReader |
scx-convert |
External format ↔ SCX conversion (h5ad, h5mu, MTX, 10x) |
scx-mtx |
Matrix Market (MTX) I/O: Cell Ranger directory read/write |
scx-cli |
CLI tool |
pyscx |
Python bindings (PyO3) |
rscx |
R bindings (extendr) |
scx-integration-tests |
Cross-crate integration tests: golden files, conformance vectors, lifecycle |
For technical details, see docs/architecture.md, docs/format.md, docs/codec.md, docs/api/, docs/sharding.md, docs/multithreading.md, docs/cloud.md, and docs/scanpy/README.md. For agent-oriented install and usage workflows in Claude Code, see skills/scx-usage/SKILL.md.
MIT