Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion modules/pipeline/EXTERNAL_PATHS.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ cp .env.example .env
| 04 | `04_vep/raw_vep.tsv` | 原始 VEP TSV |
| 04 | `04_vep/vep.log` | VEP 日志 |
| 05 | `05_vcf_info_to_csv/vep_output.with_info.csv` | 合并 VCF INFO/FORMAT |
| 06 | `06_genos_evee_annotation/vep_output.with_genos_evee.csv` | GENOS-EVEE 宽表 |
| 06 | `06_genos_evee_annotation/vep_output.with_genos_evee.csv` | GENOS-VarRisk 宽表 |
| 07 | `07_hla_filter/vep_output.no_hla.csv` | **最终输出**(默认 `hla_filter=yes`) |
| 汇总 | `full_pipeline.outputs.tsv` | 各步路径索引 |
| 日志 | `logs/full_pipeline.log` | 全流程日志 |
Expand Down
10 changes: 5 additions & 5 deletions modules/pipeline/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# OpenRare Pipeline — 罕见病变异注释与宽表输出全流程

从患者 VCF 出发,串联 Beagle phasing、VCF 前置处理、假基因注释、VEP 多插件注释、INFO 回填、GENOS-EVEE 注释与 HLA 过滤,一键产出可解读的宽表 CSV;支持命令行与 FastAPI 两种调用方式。
从患者 VCF 出发,串联 Beagle phasing、VCF 前置处理、假基因注释、VEP 多插件注释、INFO 回填、GENOS-VarRisk 注释与 HLA 过滤,一键产出可解读的宽表 CSV;支持命令行与 FastAPI 两种调用方式。

## 目录

Expand Down Expand Up @@ -123,7 +123,7 @@ full_log /path/to/output/logs/full_pipeline.log
最终 CSV 表头示例(列较多,此处仅示意):

```text
#CHROM,POS,REF,ALT,...,Gene,Consequence,CLIN_SIG,...,GENOS-EVEE,...
#CHROM,POS,REF,ALT,...,Gene,Consequence,CLIN_SIG,...,GENOS-VarRisk,...
chr1,1197557,G,A,...,TTLL10,missense_variant,...,0.42,...
```

Expand Down Expand Up @@ -315,7 +315,7 @@ pixi run bash scripts/run_full_pipeline.sh --help
| 03 | `03_pseudogene_annotation/` | 假基因 INFO 注释(可跳过) | `preprocessed.pseudogene_annotated.vcf.gz` |
| 04 | `04_vep/` | VEP + CADD/SpliceAI/dbNSFP/LoFTEE/ClinVar 等 | `vep_output.base.csv` |
| 05 | `05_vcf_info_to_csv/` | VCF INFO/FORMAT 回填到 CSV | `vep_output.with_info.csv` |
| 06 | `06_genos_evee_annotation/` | GENOS-EVEE 疾病预测分数 | `vep_output.with_genos_evee.csv` |
| 06 | `06_genos_evee_annotation/` | GENOS-VarRisk 疾病预测分数 | `vep_output.with_genos_evee.csv` |
| 07 | `07_hla_filter/` | 删除 GRCh38 HLA/MHC 区行(可 `--hla-filter no` 跳过) | **`vep_output.no_hla.csv`** |

---
Expand All @@ -338,7 +338,7 @@ pixi run bash scripts/run_full_pipeline.sh --help
|------|------|
| `<out-dir>/04_vep/vep_output.base.csv` | VEP 基础 CSV |
| `<out-dir>/04_vep/raw_vep.tsv` | VEP 原始 TSV(默认保留) |
| `<out-dir>/06_genos_evee_annotation/vep_output.with_genos_evee.csv` | GENOS-EVEE 宽表(HLA 过滤前) |
| `<out-dir>/06_genos_evee_annotation/vep_output.with_genos_evee.csv` | GENOS-VarRisk 宽表(HLA 过滤前) |
| `<out-dir>/07_hla_filter/vep_output.no_hla.csv` | **最终宽表**(默认) |
| `<out-dir>/full_pipeline.outputs.tsv` | 各步产物路径索引 |
| `<out-dir>/logs/full_pipeline.log` | 全流程日志 |
Expand Down Expand Up @@ -407,7 +407,7 @@ VEP 配置 [`modules/vep_runner/config/vep_runner_config.json`](modules/vep_runn
|------|------|
| [complete_pipeline/README.md](complete_pipeline/README.md) | API 与 CLI 详细参数 |
| [modules/vep_runner/README.md](modules/vep_runner/README.md) | VEP 插件与注释库安装 |
| [modules/genos_evee_annotation/README.md](modules/genos_evee_annotation/README.md) | GENOS-EVEE 注释与数据库构建 |
| [modules/genos_evee_annotation/README.md](modules/genos_evee_annotation/README.md) | GENOS-VarRisk 注释与数据库构建 |
| [modules/hla_filter/README.md](modules/hla_filter/README.md) | HLA/MHC 区行过滤 |
| [modules/result_sorting/README.md](modules/result_sorting/README.md) | 可选致病性排序(主流程默认不启用) |
| [resource_mock/README.md](resource_mock/README.md) | 联调用迷你 mock 库 |
Expand Down
4 changes: 2 additions & 2 deletions modules/pipeline/complete_pipeline/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
> 完整架构和数据资源下载说明见上一层 `../README.md`。本文件保留主程序/API 的详细调用说明。


本目录包含 FastAPI 服务(`full_pipeline_api.py`)与 API 任务目录,负责把 liftover(可选)、phasing、VCF 前处理、假基因注释、VEP runner、VCF INFO 回填、GENOS-EVEE 注释、HLA 过滤串成一个完整流程。
本目录包含 FastAPI 服务(`full_pipeline_api.py`)与 API 任务目录,负责把 liftover(可选)、phasing、VCF 前处理、假基因注释、VEP runner、VCF INFO 回填、GENOS-VarRisk 注释、HLA 过滤串成一个完整流程。

主程序:

Expand Down Expand Up @@ -307,7 +307,7 @@ API 的字段和主程序参数一一对应。常规字段和高级覆盖字段
| `top_k_transcripts` | `--top-k-transcripts` | 高级覆盖 | 每个 variant-gene 保留的转录本数量。 |
| `clinical_tissue` | `--clinical-tissue` | 高级覆盖 | 手动传 GTEx tissue。 |
| `keep_raw_vep` | `--keep-raw-vep` | 高级覆盖 | `true` 对应 `yes`,`false` 对应 `no`。 |
| `genos_evee_db` | `--genos-evee-db` | 高级覆盖 | GENOS-EVEE CPRA 数据库路径。 |
| `genos_evee_db` | `--GENOS-VarRisk-db` | 高级覆盖 | GENOS-VarRisk CPRA 数据库路径。 |
| `dry_run` | `--dry-run` | 调试 | 只打印命令,不实际运行。 |

示例:API 只跑 chr1:
Expand Down
4 changes: 2 additions & 2 deletions modules/pipeline/complete_pipeline/full_pipeline_api.py
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ class RunRequest(BaseModel):
java_bin: Optional[str] = Field(None, description="Advanced override: Java executable for Beagle")
top_k_transcripts: Optional[int] = Field(None, description="Advanced override: transcript selection count")
clinical_tissue: str = Field("", description="Advanced override: GTEx tissue name for phenotype-aware transcript expression")
genos_evee_db: Optional[str] = Field(None, description="Advanced override: indexed GENOS-EVEE CPRA TSV.GZ")
genos_evee_db: Optional[str] = Field(None, description="Advanced override: indexed GENOS-VarRisk CPRA TSV.GZ")
keep_raw_vep: Optional[bool] = Field(None, description="Advanced override: keep raw VEP TSV")
dry_run: bool = False

Expand Down Expand Up @@ -298,7 +298,7 @@ def add_option(name: str, value) -> None:
add_option("--java-bin", resolve_path(req.java_bin) if req.java_bin else None)
add_option("--top-k-transcripts", req.top_k_transcripts)
add_option("--clinical-tissue", req.clinical_tissue)
add_option("--genos-evee-db", resolve_path(req.genos_evee_db) if req.genos_evee_db else None)
add_option("--GENOS-VarRisk-db", resolve_path(req.genos_evee_db) if req.genos_evee_db else None)
if req.keep_raw_vep is not None:
cmd.extend(["--keep-raw-vep", "yes" if req.keep_raw_vep else "no"])
if req.dry_run:
Expand Down
8 changes: 4 additions & 4 deletions modules/pipeline/modules/genos_evee_annotation/README.md
Original file line number Diff line number Diff line change
@@ -1,21 +1,21 @@
# GENOS-EVEE 疾病预测数据库注释模块
# GENOS-VarRisk 疾病预测数据库注释模块

将疾病预测模型输出中的 `p_fusion` 按 CPRA 匹配到宽表,列名固定为 `GENOS-EVEE`。对应主流程第 **06** 步。
将疾病预测模型输出中的 `p_fusion` 按 CPRA 匹配到宽表,列名固定为 `GENOS-VarRisk`。对应主流程第 **06** 步。

## 脚本

| 脚本 | 用途 |
|------|------|
| `scripts/build_genos_evee_db.py` | 将 prediction TSV shard 构建为 bgzip + tabix 数据库 |
| `scripts/add_genos_evee_to_csv.py` | 在 05 宽表中加入 `GENOS-EVEE` 列 |
| `scripts/add_genos_evee_to_csv.py` | 在 05 宽表中加入 `GENOS-VarRisk` 列 |

## 主流程

```bash
bash scripts/run_full_pipeline.sh \
--input-vcf /path/to/input.vcf.gz \
--out-dir /path/to/output \
--genos-evee-db /path/to/genos_evee.cpra.tsv.gz
--GENOS-VarRisk-db /path/to/genos_evee.cpra.tsv.gz
```

默认数据库路径由 `.env` 的 `FULL_PIPELINE_GENOS_EVEE_DB` 或 `config/path_utils.py` 解析。数据库不存在时流程仍可运行,该列全部填 `-`。
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
import tempfile
from pathlib import Path

OUTPUT_COLUMN = "GENOS-EVEE"
OUTPUT_COLUMN = "GENOS-VarRisk"


def normalize_chrom(value: str) -> str:
Expand Down Expand Up @@ -70,7 +70,7 @@ def annotate(input_csv: Path, database: Path | None, output_csv: Path, log_json:
output_csv.parent.mkdir(parents=True, exist_ok=True)
database_available = database is not None and database.is_file() and Path(f"{database}.tbi").is_file()
if database is not None and database.exists() and not database_available:
raise FileNotFoundError(f"GENOS-EVEE database index not found: {database}.tbi")
raise FileNotFoundError(f"GENOS-VarRisk database index not found: {database}.tbi")
if database_available and shutil.which("tabix") is None:
raise RuntimeError("tabix command not found")

Expand Down Expand Up @@ -153,9 +153,9 @@ def annotate(input_csv: Path, database: Path | None, output_csv: Path, log_json:


def main() -> None:
parser = argparse.ArgumentParser(description="Add GENOS-EVEE p_fusion scores to a V3 wide CSV before sorting.")
parser = argparse.ArgumentParser(description="Add GENOS-VarRisk p_fusion scores to a V3 wide CSV before sorting.")
parser.add_argument("--input-csv", required=True)
parser.add_argument("--database", help="Indexed GENOS-EVEE .tsv.gz; missing/empty means fill '-'")
parser.add_argument("--database", help="Indexed GENOS-VarRisk .tsv.gz; missing/empty means fill '-'")
parser.add_argument("--output-csv", required=True)
parser.add_argument("--log-json")
args = parser.parse_args()
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -165,7 +165,7 @@ def build_database(inputs: list[Path], output: Path, threads: int, log_json: Pat
)
unique = 0
with sorted_path.open("r", encoding="utf-8") as src, merged.open("w", encoding="utf-8") as dst:
dst.write("#CHROM\tPOS\tREF\tALT\tGENOS-EVEE\n")
dst.write("#CHROM\tPOS\tREF\tALT\tGENOS-VarRisk\n")
previous: tuple[str, str, str, str] | None = None
for line in src:
parts = line.rstrip("\n").split("\t")
Expand All @@ -190,7 +190,7 @@ def build_database(inputs: list[Path], output: Path, threads: int, log_json: Pat


def main() -> None:
parser = argparse.ArgumentParser(description="Build indexed GENOS-EVEE CPRA database from prediction TSV shards.")
parser = argparse.ArgumentParser(description="Build indexed GENOS-VarRisk CPRA database from prediction TSV shards.")
parser.add_argument("--input", action="append", required=True, help="TSV file, directory, or glob; repeatable")
parser.add_argument("--output", required=True, help="Output .tsv.gz database")
parser.add_argument("--threads", type=int, default=4)
Expand Down
2 changes: 1 addition & 1 deletion modules/pipeline/modules/hla_filter/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# HLA/MHC 区域行过滤模块

从 GENOS-EVEE 宽表中删除 GRCh38 HLA/MHC 区域的变异行。对应主流程第 **07** 步(默认开启)。
从 GENOS-VarRisk 宽表中删除 GRCh38 HLA/MHC 区域的变异行。对应主流程第 **07** 步(默认开启)。

## 过滤区间

Expand Down
4 changes: 2 additions & 2 deletions modules/pipeline/modules/result_sorting/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

## 与主流程的关系

当前 `scripts/run_full_pipeline.sh` **默认不调用**本模块。主流程在 06 GENOS-EVEE 注释之后,由 07 HLA 过滤产出最终宽表:
当前 `scripts/run_full_pipeline.sh` **默认不调用**本模块。主流程在 06 GENOS-VarRisk 注释之后,由 07 HLA 过滤产出最终宽表:

```text
05 vcf_info_to_csv → 06 genos_evee_annotation → 07 hla_filter(默认最终 CSV)
Expand All @@ -14,7 +14,7 @@

## 输入

- `--input-csv`:带 `vcf_info_*`(及可选 `GENOS-EVEE`)列的宽表,通常来自 `05_vcf_info_to_csv/` 或 `06_genos_evee_annotation/`。
- `--input-csv`:带 `vcf_info_*`(及可选 `GENOS-VarRisk`)列的宽表,通常来自 `05_vcf_info_to_csv/` 或 `06_genos_evee_annotation/`。
- `--output-csv`:排序后的 CSV。

## 单独运行示例
Expand Down
2 changes: 1 addition & 1 deletion modules/pipeline/modules/vcf_info_to_csv/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

功能:读取进入 04 VEP runner 的 VCF,把 VCF 中所有 INFO 字段展开为 `vcf_info_<INFO_ID>` 列,并把样本 FORMAT 展开为 `vcf_format_*` 列,追加到 04 生成的 VEP CSV 中。

主流程中本模块为第 **05** 步;输出 `vep_output.with_info.csv` 进入 **06** GENOS-EVEE 注释,再经 **07** HLA 过滤(默认)得到最终宽表。
主流程中本模块为第 **05** 步;输出 `vep_output.with_info.csv` 进入 **06** GENOS-VarRisk 注释,再经 **07** HLA 过滤(默认)得到最终宽表。

## 输入

Expand Down
97 changes: 92 additions & 5 deletions modules/pipeline/modules/vep_runner/scripts/run_vep_to_csv.py
Original file line number Diff line number Diff line change
Expand Up @@ -805,6 +805,50 @@ def parse_uploaded_variation(value: str) -> dict[str, str]:
return parsed


def unique_ordered(values: list[str]) -> list[str]:
seen: set[str] = set()
result: list[str] = []
for value in values:
if not value or value in seen:
continue
seen.add(value)
result.append(value)
return result


def split_variant_ids(variant_id: str) -> list[str]:
variant_id = (variant_id or "").strip()
if not variant_id or variant_id in {".", "-"}:
return []
return unique_ordered([variant_id, *[item.strip() for item in variant_id.split(";")]])


def parse_vep_location(location: str) -> tuple[str, str] | None:
location = (location or "").strip()
if not location or ":" not in location:
return None
chrom, _, interval = location.partition(":")
start = interval.split("-", 1)[0].strip()
if not chrom or not start:
return None
return chrom, start


def location_lookup_keys(chrom: str, pos: str, ref: str, alt: str) -> list[str]:
keys = [f"loc:{chrom}:{pos}:{alt}"]
try:
shifted_pos = str(int(pos) + 1)
except ValueError:
return keys

keys.append(f"loc:{chrom}:{shifted_pos}:{alt}")
if len(ref) < len(alt) and alt.startswith(ref):
keys.append(f"loc:{chrom}:{shifted_pos}:{alt[len(ref):] or '-'}")
elif len(ref) > len(alt) and ref.startswith(alt):
keys.append(f"loc:{chrom}:{shifted_pos}:-")
return keys


def vep_uploaded_variation_keys(chrom: str, pos: str, ref: str, alt: str) -> list[str]:
keys = [f"{chrom}_{pos}_{ref}/{alt}"]
if len(ref) == len(alt):
Expand Down Expand Up @@ -838,9 +882,52 @@ def original_variant_lookup_keys(
back to the original VCF allele.
"""
keys = vep_uploaded_variation_keys(chrom, pos, ref, alt)
if variant_id and variant_id not in {".", "-"}:
keys.append(variant_id)
return keys
keys.extend(location_lookup_keys(chrom, pos, ref, alt))
for item in split_variant_ids(variant_id):
keys.append(f"id:{item}")
keys.append(f"id:{item}|allele:{alt}")
return unique_ordered(keys)


def vep_row_candidate_keys(row: dict[str, str]) -> list[str]:
uploaded = row.get("Uploaded_variation", "")
allele = row.get("Allele", "")
keys: list[str] = []
if uploaded:
keys.append(uploaded)

parsed = parse_uploaded_variation(uploaded)
chrom = parsed.get("chrom", "")
pos = parsed.get("pos", "")
ref = parsed.get("ref", "")
alt = parsed.get("alt", "")
if chrom and pos and ref and alt:
keys.extend(vep_uploaded_variation_keys(chrom, pos, ref, alt))
if allele and allele != alt:
keys.extend(vep_uploaded_variation_keys(chrom, pos, ref, allele))

location = parse_vep_location(row.get("Location", ""))
if location and allele:
loc_chrom, loc_pos = location
keys.append(f"loc:{loc_chrom}:{loc_pos}:{allele}")

for item in split_variant_ids(uploaded):
if allele:
keys.append(f"id:{item}|allele:{allele}")
if alt:
keys.append(f"id:{item}|allele:{alt}")
keys.append(f"id:{item}")
keys.append(item)

return unique_ordered(keys)


def lookup_original_variant(row: dict[str, str], original_variants) -> dict[str, str] | None:
for key in vep_row_candidate_keys(row):
value = original_variants.get(key)
if value is not None:
return value
return None


def vcf_info_column_name(info_id: str) -> str:
Expand Down Expand Up @@ -1915,7 +2002,7 @@ def convert_vep_table_to_csv(
row = dict(zip(header, parts))
extra = parse_extra(row.pop("Extra", ""))
row.update(extra)
original_variant = original_variants.get(row.get("Uploaded_variation", ""))
original_variant = lookup_original_variant(row, original_variants)
add_requested_columns(row, extra, original_variant)
extra_keys.update(extra)
variant_id = row.get("Uploaded_variation", "")
Expand Down Expand Up @@ -2026,7 +2113,7 @@ def parse_vep_table_groups(
extra = parse_extra(row.pop("Extra", ""))
row.update(extra)
variant_id = row.get("Uploaded_variation", "")
original_variant = original_variants.get(variant_id)
original_variant = lookup_original_variant(row, original_variants)
add_requested_columns(row, extra, original_variant)
if current_key is not None and variant_id != current_key:
yield current_key, current_rows
Expand Down
2 changes: 1 addition & 1 deletion modules/pipeline/resource_mock/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ bash scripts/run_full_pipeline.sh \
--phasing no \
--ccre-bed resource_mock/regulatory/hg38/encode_screen_v4_grch38_ccre.slim.mock.bed.gz \
--ncrna-bed resource_mock/ncrna/hg38/gencode.v49.ncrna_gene.slim.mock.bed.gz \
--genos-evee-db resource_mock/genos_evee/genos_evee.cpra.mock.tsv.gz \
--GENOS-VarRisk-db resource_mock/genos_evee/genos_evee.cpra.mock.tsv.gz \
--dry-run
```

Expand Down
Loading