Skip to content
innate2adaptivePublic

About

Send requests programmatically to TcRictionary: a T-cell receptor database

Resources

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

38 Commits

Folders and files

Repository files navigation

TcRictionary Python Client (tcricpy)

tcricpy is a tool used to send requests to TcRictionary: a T-cell receptor database. Think of TcRictionary as an easy to use gateway to a number of other T-cell related databases, with a focus on making data acquisition simple.

TcRictionary is currently in alpha and so some features may be incomplete or not quite working as intended. Please create an issue on this repo if you discover any unexpected behaviour. We may possibly be aware of it, but knowing what issues are actually causing problems in your workflow is very helpful.

There is a web app with similar functionality available at: www.tcrictionary.org.

An important note when using TcRictionary is that the underlying structure is a graph. This means when making queries, it is worth quickly consulting the schema to observe the path you wish to take and the conotations this has upon the semantics of your query. The current schema (as of 27/06/25) is:

Graph Schema

You can read more about how to use and format your queries at: docs.tcrictionary.org.

Installation

To install tcricpy, clone the repo, activate your desired environment, and then from the top level of the repo run:

pip install .

Script Usage

To use tcricpy in scripts, you must first initialise a TcrictionaryClient, and then call one of the associated methods via asyncio.run() as they are asynchronous functions.

There are two main categories of request that tcricpy can make to TcRictionary.

1. Annotate

The TcrictionaryClient.annotate() method queries TcRictionary with your supplied information (annotatees) and attempts to find any database entries that match. It then searches for the requested information (annotations) connected to those database entries.

# ./examples/auto-annotation.py
import asyncio

import pandas as pd

from tcricpy import client

db = client.TcrictionaryClient()

input_df = pd.DataFrame(
    {
        "VBeta.gene": [
            "TRBV14",
            "TRBV13-2",
            "TRBV6-3",
        ],
        "VAlpha.gene": ["TRAV26-1", "TRAV6D-3", "TRAV29/DV5"],
        "JAlpha.gene": ["TRAJ37", "TRAJ13", "TRAJ52"],
        "Cdr3Alpha.id": [
            "CIVVRSSNTGKLIF",
            "CAANSGTYQRF",
            "CAASVYAGGTSYGKLTF",
        ],
    }
)

print("Starting annotation...")

# Directly run the coroutine returned by db.annotate
result = asyncio.run(db.annotate(annotatees=input_df, annotations="Epitope"))

print("Annotation finished.")
print(result)

As under the hood, TcRictionary uses a graph database, it is possible to supply a manual path through the structure to further refine the semantics of your query.

Additional information about how pathing works can be found at docs.tcrictionary.org.

# ./examples/path-annotation.py
import asyncio

import pandas as pd

from tcricpy import client

db = client.TcrictionaryClient()

input_df = pd.DataFrame(
    {
        "Epitope.id": ["VMAPRTLIL"],
    }
)

print("Starting annotation...")

# Directly run the coroutine returned by db.annotate
result = asyncio.run(
    db.annotate(
        annotatees=input_df,
        annotations="Cdr3Alpha",
        module_path=[
            "PMhc",
            "Study_PMhc_Bridge",
            "Study",
            "Tcr_Study_Bridge",
            "Tcr",
        ],
    )
)

print("Annotation finished.")
print(result)

2. Discover

The second way to use TcRictionary is to query against the entire database for a given set of database entries. Use this method if you wish to obtain full lists of TCR-pMHC pairs.

⚠️ Note: This functionality is currently slow. Depending on the categories selected it may not complete in a reasonable time frame. We are investigating emailing you the result once they query is finished. At present your query will timeout at 10 minutes.

# ./examples/discovery.py
import asyncio

from tcricpy import client

db = client.TcrictionaryClient()

print("Starting discovery...")

# Directly run the coroutine returned by db.annotate
result = asyncio.run(
    db.discover(
        to_discover=["Cdr3Alpha"],
        annotations="Epitope",
        module_path=[
            "Tcr",
            "Tcr_Study_Bridge",
            "Study",
            "Study_PMhc_Bridge",
            "PMhc",
        ],
    )
)

print("Discovery finished.")
print(result)

CLI Usage

Decombinator Annotation CLI

A command-line tool for annotating TCR sequences produced by Decombinator using TcRictionary.

Installation

Install dependencies alongside tcricpy:

pip install . pandas

Usage

python annotate_tcr.py --chain {alpha,beta} [OPTIONS] FILES...

FILES accepts one or more .tsv or .tsv.gz paths, including shell-expanded globs:

# Single file
python annotate_tcr.py --chain beta sample_beta.tsv.gz

# Shell glob (expanded by the shell)
python annotate_tcr.py --chain beta data/dcr_*_beta.tsv.gz

# Quoted glob (expanded by the script)
python annotate_tcr.py --chain beta "data/**/*_beta.tsv.gz"

Options

Flag Default Description
--chain (required) TCR chain: alpha or beta
--output <first_file>_annotated.csv Path for the combined output CSV
--chunk-size 1000 Rows per TcRictionary API call
--skip-errors False Skip files with errors instead of aborting

Input Format

The tool expects Decombinator's translated output files — the TSV files produced by the Decombinator pipeline's translation step, which contain columns including v_call, j_call, junction_aa, and productive. Both uncompressed (.tsv) and gzip-compressed (.tsv.gz) files are supported.

Rows where productive != "T" are filtered out before annotation.

Output

All input files are processed and their results concatenated into a single CSV annotated with Epitope, Species, and Epitope_PRODUCES_Species from TcRictionary. The first column of the output, source_file, contains the stem of the originating filename (all extensions stripped, e.g. dcr_PROJECT_0001_1_beta), allowing rows to be traced back to their source when multiple files are combined.

About

Send requests programmatically to TcRictionary: a T-cell receptor database

Resources

Stars

1 star

Watchers

1 watching

Forks

Contributors

Languages