tcricpy is a tool used to send requests to TcRictionary: a T-cell receptor
database. Think of TcRictionary as an easy to use gateway to a number of other
T-cell related databases, with a focus on making data acquisition simple.
TcRictionary is currently in alpha and so some features may be incomplete or not quite working as intended. Please create an issue on this repo if you discover any unexpected behaviour. We may possibly be aware of it, but knowing what issues are actually causing problems in your workflow is very helpful.
There is a web app with similar functionality available at: www.tcrictionary.org.
An important note when using TcRictionary is that the underlying structure is a graph. This means when making queries, it is worth quickly consulting the schema to observe the path you wish to take and the conotations this has upon the semantics of your query. The current schema (as of 27/06/25) is:
You can read more about how to use and format your queries at: docs.tcrictionary.org.
To install tcricpy, clone the repo, activate your desired environment, and
then from the top level of the repo run:
pip install .To use tcricpy in scripts, you must first initialise a TcrictionaryClient,
and then call one of the associated methods via asyncio.run() as they are
asynchronous functions.
There are two main categories of request that tcricpy can make to
TcRictionary.
The TcrictionaryClient.annotate() method queries TcRictionary with your
supplied information (annotatees) and attempts to find any database entries that
match. It then searches for the requested information (annotations) connected to
those database entries.
# ./examples/auto-annotation.py
import asyncio
import pandas as pd
from tcricpy import client
db = client.TcrictionaryClient()
input_df = pd.DataFrame(
{
"VBeta.gene": [
"TRBV14",
"TRBV13-2",
"TRBV6-3",
],
"VAlpha.gene": ["TRAV26-1", "TRAV6D-3", "TRAV29/DV5"],
"JAlpha.gene": ["TRAJ37", "TRAJ13", "TRAJ52"],
"Cdr3Alpha.id": [
"CIVVRSSNTGKLIF",
"CAANSGTYQRF",
"CAASVYAGGTSYGKLTF",
],
}
)
print("Starting annotation...")
# Directly run the coroutine returned by db.annotate
result = asyncio.run(db.annotate(annotatees=input_df, annotations="Epitope"))
print("Annotation finished.")
print(result)As under the hood, TcRictionary uses a graph database, it is possible to supply a manual path through the structure to further refine the semantics of your query.
Additional information about how pathing works can be found at docs.tcrictionary.org.
# ./examples/path-annotation.py
import asyncio
import pandas as pd
from tcricpy import client
db = client.TcrictionaryClient()
input_df = pd.DataFrame(
{
"Epitope.id": ["VMAPRTLIL"],
}
)
print("Starting annotation...")
# Directly run the coroutine returned by db.annotate
result = asyncio.run(
db.annotate(
annotatees=input_df,
annotations="Cdr3Alpha",
module_path=[
"PMhc",
"Study_PMhc_Bridge",
"Study",
"Tcr_Study_Bridge",
"Tcr",
],
)
)
print("Annotation finished.")
print(result)The second way to use TcRictionary is to query against the entire database for a given set of database entries. Use this method if you wish to obtain full lists of TCR-pMHC pairs.
⚠️ Note: This functionality is currently slow. Depending on the categories selected it may not complete in a reasonable time frame. We are investigating emailing you the result once they query is finished. At present your query will timeout at 10 minutes.
# ./examples/discovery.py
import asyncio
from tcricpy import client
db = client.TcrictionaryClient()
print("Starting discovery...")
# Directly run the coroutine returned by db.annotate
result = asyncio.run(
db.discover(
to_discover=["Cdr3Alpha"],
annotations="Epitope",
module_path=[
"Tcr",
"Tcr_Study_Bridge",
"Study",
"Study_PMhc_Bridge",
"PMhc",
],
)
)
print("Discovery finished.")
print(result)A command-line tool for annotating TCR sequences produced by Decombinator using TcRictionary.
Install dependencies alongside tcricpy:
pip install . pandaspython annotate_tcr.py --chain {alpha,beta} [OPTIONS] FILES...FILES accepts one or more .tsv or .tsv.gz paths, including shell-expanded
globs:
# Single file
python annotate_tcr.py --chain beta sample_beta.tsv.gz
# Shell glob (expanded by the shell)
python annotate_tcr.py --chain beta data/dcr_*_beta.tsv.gz
# Quoted glob (expanded by the script)
python annotate_tcr.py --chain beta "data/**/*_beta.tsv.gz"| Flag | Default | Description |
|---|---|---|
--chain |
(required) | TCR chain: alpha or beta |
--output |
<first_file>_annotated.csv |
Path for the combined output CSV |
--chunk-size |
1000 |
Rows per TcRictionary API call |
--skip-errors |
False |
Skip files with errors instead of aborting |
The tool expects Decombinator's translated output files — the TSV files produced
by the Decombinator pipeline's translation step, which contain columns including
v_call, j_call, junction_aa, and productive. Both uncompressed (.tsv)
and gzip-compressed (.tsv.gz) files are supported.
Rows where productive != "T" are filtered out before annotation.
All input files are processed and their results concatenated into a single CSV
annotated with Epitope, Species, and Epitope_PRODUCES_Species from
TcRictionary. The first column of the output, source_file, contains the stem
of the originating filename (all extensions stripped, e.g.
dcr_PROJECT_0001_1_beta), allowing rows to be traced back to their source when
multiple files are combined.