Coevolutionary discovery using genome language models
Minerva predicts coevolution using genome language models. Powered by Minerva-MLM, it delivers database-scale, alignment-free, interaction-specific predictions across prokaryotic genomes. Through adaptation on homologous loci, Minerva can discover additional interactions.
Try it in the browser, no install needed: Minerva on Hugging Face Spaces.
Demo · Install · Checkpoints · Quick start · Preparing inputs · Interaction heads · Jacobian fingerprinting · RNA structure · Eukaryotic RNA · Finetuning · Examples · Citation
To install Minerva, use:
pip install minerva-dnaMinerva uses flash-attn automatically when it is installed. Otherwise it falls
back to PyTorch SDPA. Please install flash attention first for faster inference.
Minerva-MLM is a 650M parameter transformer trained for over 1.3 trillion tokens (~3.4 Terabases). Minerva-MLM is initialized from gLM2 650M and adopts the mixed-modality tokenization, and was trained at 4096 and 8192 context lengths.
Checkpoints are hosted on Hugging Face:
| Model | Context | Hugging Face repo |
|---|---|---|
| Minerva-MLM | 4096 | gbrixi/minerva-mlm |
| Minerva-MLM-8k | 8192 | gbrixi/minerva-mlm-8k |
Minerva-MLM checkpoints include three interaction heads and Jacobian fingerprint types:
- base_pairing — RNA base-pairing contacts
- protein — protein contact prediction
- repeat — repeat element detection
from transformers import AutoTokenizer
from minerva import MinervaForMaskedLM
import torch
model = MinervaForMaskedLM.from_pretrained(
"gbrixi/minerva-mlm", torch_dtype=torch.bfloat16,
).cuda().eval()
tokenizer = AutoTokenizer.from_pretrained("gbrixi/minerva-mlm")
tokens = tokenizer(
"<+>cgcggggtggagcagcctggtagctcgtcgggctcataacccgaagatcgtcggttcaaatccggcccccgcaacca",
return_tensors="pt",
).to(model.device)
with torch.no_grad():
outputs = model(**tokens, output_interactions=True)
base_pairing = outputs.interactions["base_pairing"] # [batch, L, L]
protein = outputs.interactions["protein"] # [batch, L, L]
repeat = outputs.interactions["repeat"] # [batch, L, L]Importing the class directly needs no
trust_remote_code. Without the package installed, useAutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True).
Minerva reads a mixed protein + DNA sequence: coding regions are upper-case
amino acids, intergenic regions are lower-case nucleotides, and <+> / <->
markers denote strand.
<+>MALTKVEKRNRIKRRVRGK<+>aatttaaggaa<->MLGIDNIERVKPGGLELVDRLV
└── CDS (protein) ──┘└ intergenic ┘└──── CDS on - strand ────┘
There are three ways to produce this format depending on what you start with:
| You have | Use | What happens |
|---|---|---|
| Annotated GenBank (CDS features) | minerva.data.extract_and_tokenize_gb(path) |
CDS translated, intergenic kept as DNA, strand markers inserted. One sequence per LOCUS. |
| Unannotated sequence (FASTA / raw DNA) | minerva.gene_calling.build_minerva_input(seq) |
Genes called with Pyrodigal, then packaged. |
| Raw genome + external CDS calls | minerva.sequence_utils.build_prodigal_mixed_sequence(seq, cds) |
Your own [{start, end, strand}] calls packaged, with genome↔token maps. |
If you only have a FASTA file or a raw nucleotide string, we use Pyrodigal to automatically call genes:
from minerva.gene_calling import build_minerva_input, fasta_to_minerva_inputs
# From a single nucleotide string
out = build_minerva_input(sequence) # dict: token_string + coord maps
token_string = out["token_string"]
# From a FASTA file (one result per record)
inputs = fasta_to_minerva_inputs("contigs.fasta")Pyrodigal needs ≥ 20 kb to estimate gene-scoring statistics from a sequence;
shorter contigs use its pre-trained profiles, which meta=True forces
for metagenomic assemblies. A CDS token is one amino acid and an intergenic
token one base, so token_to_genome / genome_to_token map token index to
genome position.
Minerva's context is 4096 (gbrixi/minerva-mlm) or 8192 tokens
(gbrixi/minerva-mlm-8k). One token is one amino acid, one nucleotide, or one
strand marker, so a typical (~88 % coding) bacterial genome packs to ~10 kb per
4096 tokens (~20 kb for the 8k model).
Pass max_tokens to the builders to cap a sequence. It truncates at a gene
boundary, keeps the 5′ end, and keeps the coordinate maps consistent.
out = build_minerva_input(sequence, max_tokens=4096) # <= 4096 tokensTo cover a whole genome, tile it instead with
minerva.sequence_utils.chunk_sequence_with_stride.
CDS translate with NCBI table 11 by default, but a GenBank feature's own
/transl_table takes precedence. Override with translation_table= on the
builders or --translation_table on scripts/finetune.py.
See examples/ for runnable, end-to-end walkthroughs.
These continue from the quick start, with model, tokenizer, tokens and
outputs already defined.
To plot the interaction-head outputs:
from minerva.visualization import plot_contacts
token_list = tokenizer.convert_ids_to_tokens(tokens["input_ids"][0].tolist())
plot_contacts(outputs, tokens=token_list, track=True) # track=True adds a CDS / intergenic stripplot_contacts takes the model outputs, a dict of head maps, or a Jacobian
FingerprintResult (below), and renders them all in the same palette. For a
zoomable view with hover values, plot_contacts_interactive takes the same
arguments and returns a Bokeh layout (pip install "minerva-dna[viz]").
Set interaction_layers=6 to use the six-layer interaction heads:
outputs = model(**tokens, output_interactions=True, interaction_layers=6)Raw attention tensors follow the Hugging Face convention:
outputs = model(
**tokens,
output_interactions=True,
output_attentions=True,
attention_layers=[31, 32],
)Use get_fingerprints when you want named interaction-pattern channels from a
sequence. It computes the required full categorical Jacobian internally and returns
a FingerprintResult; the large raw Jacobian is not kept unless requested.
sequence = "<+>cgcggggtggagcagcctggtagctcgtcgggctcataacccgaagatcgtcggttcaaatccggcccccgcaacca"
window = (0, min(len(sequence), 96))
fp = model.get_fingerprints(
sequence,
tokenizer,
position_range=window,
max_batch_size=64,
)
fp.channel_names # ["basepairing", "repeat", "protein", "other"]
basepairing = fp["basepairing"] # [L, L]
protein = fp["protein"] # [L, L]
repeat = fp["repeat"] # [L, L]To plot the result:
from minerva.visualization import plot_contacts
plot_contacts(fp, tokens=fp.tokens, title="Minerva multimodal fingerprint")The base_pairing head returns a dense contact map which can be converted to an RNA structure using minerva.rna_structure:
from minerva.rna_structure import call_structure, call_structures
token_list = tokenizer.convert_ids_to_tokens(tokens["input_ids"][0].tolist())
s = call_structure(outputs.interactions["base_pairing"], tokens=token_list)
s.dot_bracket # '(((((((..((((........)))).(((((.......)))))...'
s.to_vienna("t.fa") # read by RNAfold, forna, VARNA, R2R
s.to_ct("t.ct") # connect table, keeps pseudoknots
s.plot() # matplotlib FigureEach intergenic region of a mixed locus is a separate molecule, so
call_structures returns one structure per region:
structures = call_structures(outputs.interactions["base_pairing"], token_list)An interactive viewer is in
examples/notebooks/rna_structure.ipynb.
For researchers studying eukaryotic RNAs, Minerva provides a RiNALMo-based checkpoint with base-pairing and repeat interaction heads. See eukaryotic RNA support for usage and finetuning.
scripts/finetune.py wraps HF Trainer + accelerate with Minerva's MLM loss.
It supports full finetuning and LoRA, and ingests GenBank files directly.
# LoRA finetune
accelerate launch --num_processes=8 scripts/finetune.py \
--output_dir ./output \
--genbank_file genome.gb \
--tokenizer_name gbrixi/minerva-mlm \
--model_name_or_path gbrixi/minerva-mlm \
--use_lora --lora_r 1 --lora_alpha 2 \
--learning_rate 1e-4 --bf16
# Full finetune on a GenBank file across 8 GPUs
accelerate launch --num_processes=8 scripts/finetune.py \
--output_dir ./output \
--genbank_file genome.gb \
--tokenizer_name gbrixi/minerva-mlm \
--model_name_or_path gbrixi/minerva-mlm \
--per_device_train_batch_size 4 \
--bf16Minerva-MLM LoRA checkpoints loaded with PEFT:
from peft import PeftModel
from minerva import MinervaForMaskedLM
base = MinervaForMaskedLM.from_pretrained("gbrixi/minerva-mlm")
model = PeftModel.from_pretrained(base, "path/to/lora_ckpt")A LOCUS longer than --max_seq_length is split into non-overlapping blocks,
each its own training example, see
minerva.finetuning.load_genbank_dataset.
minerva/
modeling_minerva.py # MinervaConfig / MinervaForMaskedLM (custom transformer + heads)
modeling_rinalmo.py # RiNALMoMinervaForMaskedLM (RNA backbone + heads)
tokenization_rinalmo.py # RiNALMo nucleotide tokenizer
backbones.py # Backbone-specific training configuration
interaction_heads.py # Shared interaction heads
data.py # GenBank parsing + tokenization
gene_calling.py # Pyrodigal gene calling: FASTA/raw DNA -> mixed tokens
sequence_utils.py # reverse-complement + external-CDS -> mixed tokens
masking.py # DataCollatorForMinervaMLM
losses.py # grouped_mlm_loss
example_data/ # sample GenBank loci (minerva.data.example_path)
scripts/
finetune.py # HF Trainer / accelerate wrapper
examples/ # end-to-end tutorials & notebooks
call_genes_from_fasta.py # FASTA/raw DNA -> Minerva input walkthrough
notebooks/ # interactive Colab-ready notebooks
tests/ # package unit + smoke tests
If you use Minerva, please cite Li et al., bioRxiv 2026:
@article{li2026minerva,
title = {Coevolutionary mining of prokaryotic non-coding elements with a genome language model},
author = {Li, David B. and Brixi, Garyk and Kim, Alexandra S. and Fiamenghi, Mateus B. and Driscoll, Claudia L. and Evans, Simone A. and Gao, Alex and Ivanova, Natalia N. and Kyrpides, Nikos C. and Deisseroth, Karl and Wilkinson, Max E. and Fischbach, Michael A. and Hie, Brian L.},
journal = {bioRxiv},
year = {2026},
doi = {10.64898/2026.09.22.753630},
url = {https://www.biorxiv.org/content/10.64898/2026.09.22.753630},
publisher = {Cold Spring Harbor Laboratory}
}If you use the Jacobian fingerprints, please also cite the original categorical Jacobian, Zhang et al., PNAS 2024:
@article{zhang2024categoricaljacobian,
title = {Protein language models learn evolutionary statistics of interacting sequence motifs},
author = {Zhang, Z. and Wayment-Steele, H. K. and Brixi, G. and Wang, H. and Kern, D. and Ovchinnikov, S.},
journal = {Proceedings of the National Academy of Sciences},
volume = {121},
number = {45},
pages = {e2406285121},
year = {2024},
doi = {10.1073/pnas.2406285121}
}Apache 2.0 — see LICENSE. Minerva-MLM is initialized from gLM2 650M (Tatta Bio, Apache 2.0).