Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
1e9b218
Add Logistic PCA (LPCA) analysis
Jeebjean Jul 10, 2026
bb338a8
Add LPCA Docker wrapper
Jeebjean Jul 10, 2026
efd1005
Add unit tests for LPCA analysis
Jeebjean Jul 21, 2026
8f89119
Add LPCA visualization: generate lpca.png and lpca-coordinates.txt
Jeebjean Jul 21, 2026
3d9a38f
Increase Docker timeout to 600s for large LPCA matrices
Jeebjean Jul 21, 2026
e4b4679
Restore config.yaml to main defaults, keep only lpca block
Jeebjean Jul 21, 2026
41b64b7
Skip LPCA tests if Docker image not available
Jeebjean Jul 21, 2026
c49362f
Add rARPACK for memory-efficient LPCA with partial_decomp=TRUE
Jeebjean Jul 23, 2026
7536d8c
Address PR review: palette, CV check, KDE note, config docs, rename t…
Jeebjean Jul 27, 2026
63d941d
Remove transpose option, fix m doc, update README, rename party variable
Jeebjean Jul 27, 2026
a44bb50
Restore deleted comments in config.yaml
Jeebjean Jul 27, 2026
58ceac3
Pass summarize_networks output to run_lpca, matching PCA convention
Jeebjean Jul 27, 2026
33f0435
Move binary matrix to Snakefile output, pass path to run_lpca
Jeebjean Jul 27, 2026
90b5094
Update tests: remove skipif, use summarize_networks, remove transpose
Jeebjean Jul 27, 2026
20ff901
Add lpca_analysis_all rule for cross-algorithm LPCA
Jeebjean Jul 27, 2026
a16ba19
Add dedicated LPCA test inputs and stronger tests
Jeebjean Jul 27, 2026
b063bbd
Add LPCA to Docker image build CI workflow
Jeebjean Jul 27, 2026
b1905c0
Fix test input files: use tab-separated format
Jeebjean Jul 27, 2026
ec14dc8
Add lpca aggregate_per_algorithm option and per-algorithm rule; enabl…
Jeebjean Jul 30, 2026
1704c47
Set lpca include back to false after confirming CI run
Jeebjean Jul 30, 2026
5b64162
Add stronger LPCA test comparing scores to known outputs (sign-robust)
Jeebjean Jul 30, 2026
2ced8af
Add robustness to LPCA R wrapper and test code
agitter Sep 29, 2026
e1c0be2
Refactor LPCA wrapper and tests
agitter Oct 1, 2026
e2c91d2
Add tests, refactor config and Snakefile
agitter Oct 1, 2026
2d7d91c
Formatting and doc cleanup
agitter Oct 2, 2026
894643d
Config whitespace
agitter Oct 2, 2026
f2a24f4
Update comments
agitter Oct 2, 2026
415c799
Memory discussions
agitter Oct 4, 2026
7100ada
Disable LPCA by default
agitter Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .github/workflows/build-containers.yml
Original file line number Diff line number Diff line change
Expand Up @@ -73,3 +73,8 @@ jobs:
# of detected changes.
always_build: true
context: ./
build-and-remove-lpca:
uses: "./.github/workflows/build-and-remove-template.yml"
with:
path: docker-wrappers/LPCA
container: reedcompbio/lpca
71 changes: 70 additions & 1 deletion Snakefile
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import shutil
import yaml
from spras.dataset import Dataset
from spras.evaluation import Evaluation
from spras.analysis import ml, summary, cytoscape
from spras.analysis import ml, summary, cytoscape, lpca
from spras.config.revision import detach_spras_revision
import spras.config.config as _config

Expand All @@ -24,6 +24,7 @@ _config.init_global(config)
out_dir = _config.config.out_dir
algorithm_params = _config.config.algorithm_params
pca_params = _config.config.pca_params
lpca_params = _config.config.lpca_params
hac_params = _config.config.hac_params
container_settings = _config.config.container_settings
include_aggregate_algo_eval = _config.config.analysis_include_evaluation_aggregate_algo
Expand All @@ -43,6 +44,8 @@ def algo_has_mult_param_combos(algo):
return len(algorithm_params.get(algo, {})) > 1

algorithms_mult_param_combos = [algo for algo in algorithms if algo_has_mult_param_combos(algo)]
# LPCA requires at least three runs; keep the PCA eligibility threshold unchanged.
algorithms_lpca = [algo for algo in algorithms if len(algorithm_params[algo]) >= 3]

# Get the parameter dictionary for the specified
# algorithm and parameter combination hash
Expand Down Expand Up @@ -92,6 +95,19 @@ def make_final_input(wildcards):
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}jaccard-matrix.txt',out_dir=out_dir,sep=SEP,dataset=dataset_labels,algorithm_params=algorithms_with_params))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}jaccard-heatmap.png',out_dir=out_dir,sep=SEP,dataset=dataset_labels,algorithm_params=algorithms_with_params))

if _config.config.analysis_include_lpca:
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}lpca.png',out_dir=out_dir, sep=SEP, dataset=dataset_labels))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}lpca-deviance.txt',out_dir=out_dir, sep=SEP, dataset=dataset_labels))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}lpca-coordinates.txt',out_dir=out_dir, sep=SEP, dataset=dataset_labels))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}lpca-binary-matrix.csv',out_dir=out_dir, sep=SEP, dataset=dataset_labels))

# Only run LPCA per algorithm for the algorithms collected in algorithms_lpca
if _config.config.analysis_include_lpca_aggregate_algo:
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-lpca-deviance.txt',out_dir=out_dir, sep=SEP, dataset=dataset_labels, algorithm=algorithms_lpca))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-lpca.png',out_dir=out_dir, sep=SEP,dataset=dataset_labels,algorithm=algorithms_lpca))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-lpca-coordinates.txt',out_dir=out_dir, sep=SEP, dataset=dataset_labels,algorithm=algorithms_lpca))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-lpca-binary-matrix.csv',out_dir=out_dir, sep=SEP, dataset=dataset_labels,algorithm=algorithms_lpca))

if _config.config.analysis_include_ml_aggregate_algo:
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-pca.png',out_dir=out_dir,sep=SEP,dataset=dataset_labels,algorithm=algorithms_mult_param_combos))
final_input.extend(expand('{out_dir}{sep}{dataset}-ml{sep}{algorithm}-pca-variance.txt',out_dir=out_dir,sep=SEP,dataset=dataset_labels,algorithm=algorithms_mult_param_combos))
Expand Down Expand Up @@ -403,6 +419,59 @@ rule ml_analysis_aggregate_algo:
ml.hac_horizontal(summary_df, output.hac_image_horizontal, output.hac_clusters_horizontal, **hac_params)
ml.pca(summary_df, output.pca_image, output.pca_variance, output.pca_coordinates, **pca_params)

# Track fit and plot settings so changes schedule a new fit.
rule lpca_analysis_all:
input:
pathways = expand('{out_dir}{sep}{{dataset}}-{algorithm_params}{sep}pathway.txt', out_dir=out_dir, sep=SEP, algorithm_params=algorithms_with_params)
output:
lpca_png = SEP.join([out_dir, '{dataset}-ml', 'lpca.png']),
lpca_deviance = SEP.join([out_dir, '{dataset}-ml', 'lpca-deviance.txt']),
lpca_coord = SEP.join([out_dir, '{dataset}-ml', 'lpca-coordinates.txt']),
lpca_matrix = SEP.join([out_dir, '{dataset}-ml', 'lpca-binary-matrix.csv'])
params:
k = lpca_params.k,
m = lpca_params.m,
labels = lpca_params.labels,
run:
summary_df = ml.summarize_networks(input.pathways)
lpca.run_lpca(
summary_df,
output_png=output.lpca_png,
output_deviance=output.lpca_deviance,
output_coord=output.lpca_coord,
output_matrix=output.lpca_matrix,
k=params.k,
m=params.m,
labels=params.labels,
container_settings=container_settings
)

rule lpca_analysis_aggregate_algo:
input:
pathways = collect_pathways_per_algo
output:
lpca_png = SEP.join([out_dir, '{dataset}-ml', '{algorithm}-lpca.png']),
lpca_deviance = SEP.join([out_dir, '{dataset}-ml', '{algorithm}-lpca-deviance.txt']),
lpca_coord = SEP.join([out_dir, '{dataset}-ml', '{algorithm}-lpca-coordinates.txt']),
lpca_matrix = SEP.join([out_dir, '{dataset}-ml', '{algorithm}-lpca-binary-matrix.csv'])
params:
k = lpca_params.k,
m = lpca_params.m,
labels = lpca_params.labels,
run:
summary_df = ml.summarize_networks(input.pathways)
lpca.run_lpca(
summary_df,
output_png=output.lpca_png,
output_deviance=output.lpca_deviance,
output_coord=output.lpca_coord,
output_matrix=output.lpca_matrix,
k=params.k,
m=params.m,
labels=params.labels,
container_settings=container_settings
)

# Ensemble the output pathways for each dataset per algorithm
rule ensemble_per_algo:
input:
Expand Down
61 changes: 38 additions & 23 deletions config/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -50,29 +50,29 @@ containers:
# requirements = versionGE(split(Target.CondorVersion)[1], "24.8.0") && (isenforcingdiskusage =!= true)
enable_profiling: false

# Override the default container image for specific algorithms.
# Keys are algorithm names (as they appear in the algorithms list below).
# Values are interpreted based on the container framework:
#
# Image reference (e.g., "pathlinker:v3"):
# Prepends the registry prefix. Works with both Docker and Apptainer.
#
# Full image reference with registry (e.g., "ghcr.io/myorg/pathlinker:v3"):
# Used as-is (prefix NOT prepended). Works with both Docker and Apptainer.
#
# Local .sif file path (e.g., "images/pathlinker_v2.sif"):
# Apptainer/Singularity only. Skips pulling from registry and uses the
# pre-built .sif directly. When running via HTCondor with shared-fs-usage: none (set
# via the spras_profile config when running SPRAS against HTCondor), .sif paths listed
# here are automatically included in htcondor_transfer_input_files.
# Ignored with a warning if the framework is Docker.
#
# Example (one of each type):
# images:
# omicsintegrator1: "images/omics-integrator-1_v2.sif" # local .sif (Apptainer only)
# pathlinker: "pathlinker:v1234" # image name only (base_url/owner prepended)
# omicsintegrator2: "some-other-owner/oi2:latest" # owner/image (base_url prepended)
# mincostflow: "ghcr.io/reed-compbio/mincostflow:v2" # full registry reference (used as-is)
# Override the default container image for specific algorithms.
# Keys are algorithm names (as they appear in the algorithms list below).
# Values are interpreted based on the container framework:
#
# Image reference (e.g., "pathlinker:v3"):
# Prepends the registry prefix. Works with both Docker and Apptainer.
#
# Full image reference with registry (e.g., "ghcr.io/myorg/pathlinker:v3"):
# Used as-is (prefix NOT prepended). Works with both Docker and Apptainer.
#
# Local .sif file path (e.g., "images/pathlinker_v2.sif"):
# Apptainer/Singularity only. Skips pulling from registry and uses the
# pre-built .sif directly. When running via HTCondor with shared-fs-usage: none (set
# via the spras_profile config when running SPRAS against HTCondor), .sif paths listed
# here are automatically included in htcondor_transfer_input_files.
# Ignored with a warning if the framework is Docker.
#
# Example (one of each type):
# images:
# omicsintegrator1: "images/omics-integrator-1_v2.sif" # local .sif (Apptainer only)
# pathlinker: "pathlinker:v1234" # image name only (base_url/owner prepended)
# omicsintegrator2: "some-other-owner/oi2:latest" # owner/image (base_url prepended)
# mincostflow: "ghcr.io/reed-compbio/mincostflow:v2" # full registry reference (used as-is)

# This list of algorithms should be generated by a script which checks the filesystem for installs.
# It shouldn't be changed by mere mortals. (alternatively, we could add a path to executable for each algorithm
Expand Down Expand Up @@ -272,3 +272,18 @@ analysis:
# adds evaluation per algorithm per dataset-goldstandard pair
# evaluation per algorithm will not run unless ml include and ml aggregate_per_algorithm are set to true
aggregate_per_algorithm: true
lpca:
# Compare pathway graphs across all algorithms, independently of ml.include
# LPCA allocates dense edge-by-edge matrices, requiring quadratic memory.
# Partial decomposition does not remove them.
# LPCA fits may fail due to out-of-memory errors.
include: false
# Also fit each algorithm that has at least three configured parameter combinations.
# Requires lpca.include: true; algorithms with fewer combinations are omitted.
aggregate_per_algorithm: false
# Only a two-component fit is supported. Keep k=2.
k: 2
# Fixed, finite, strictly positive tuning parameter; no automatic selection
m: 6
# Show run identifiers on the LPCA plot
labels: true
6 changes: 6 additions & 0 deletions config/egfr.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,12 @@ analysis:
labels: true
kde: true
remove_empty_pathways: true
lpca:
include: false
aggregate_per_algorithm: false
k: 2
m: 6
labels: true
evaluation:
include: true
aggregate_per_algorithm: true
40 changes: 40 additions & 0 deletions docker-wrappers/LPCA/Dockerfile
Comment thread
Jeebjean marked this conversation as resolved.
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# Logistic PCA (logisticPCA) wrapper for SPRAS.
# The caller supplies Rscript /app/run_lpca.R and its arguments.
# Cross-validation is not supported, and no ENTRYPOINT is set.

# Pinned R version for reproducibility. Bump deliberately, not to :latest.
FROM rocker/r-base:4.4.2

LABEL org.opencontainers.image.source="https://github.com/Reed-CompBio/spras"
LABEL org.opencontainers.image.description="Logistic PCA (logisticPCA) wrapper for SPRAS"

# System libraries needed to compile ggplot2 (a hard Import of logisticPCA)
# and its dependency stack from source on Debian.
RUN apt-get update && apt-get install -y --no-install-recommends \
libcurl4-openssl-dev \
libssl-dev \
libxml2-dev \
libfontconfig1-dev \
libfreetype6-dev \
libpng-dev \
libtiff5-dev \
libjpeg-dev \
&& rm -rf /var/lib/apt/lists/*

# Install tested versions of required R packages including rARPACK's RSpectra backend.
# remotes finds these versions in CRAN or its archive.
# Install only required dependencies and do not upgrade previously installed ones.
# Transitive and system dependencies are not fully pinned by this Dockerfile.
RUN Rscript -e "options(repos = c(CRAN = 'https://cran.r-project.org')); \
install.packages('remotes'); \
versions <- c(RSpectra = '0.16-2', rARPACK = '0.11-0', logisticPCA = '0.2'); \
for (pkg in names(versions)) { \
remotes::install_version(pkg, version = versions[[pkg]], \
dependencies = NA, upgrade = 'never'); \
stopifnot(packageVersion(pkg) == package_version(versions[[pkg]])); \
}; \
library(logisticPCA); library(rARPACK); library(RSpectra)"

COPY run_lpca.R /app/run_lpca.R

WORKDIR /app
99 changes: 99 additions & 0 deletions docker-wrappers/LPCA/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
# LPCA (Logistic PCA) wrapper

Comment thread
Jeebjean marked this conversation as resolved.
Docker image: https://hub.docker.com/r/reedcompbio/lpca

This wrapper uses [logisticPCA](https://github.com/andland/logisticPCA)
([Landgraf & Lee, 2020](https://doi.org/10.1016/j.jmva.2020.104668)) for an
exploratory, two-dimensional comparison of reconstructed networks. It calls
`logisticPCA`, not `logisticSVD`. It supports only `k = 2` components and a fixed,
finite, strictly positive `m`. Cross-validation for automatic selection of `m`
and modifying `k` are not supported.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Extra linebreak here

## Container interface

The image contains one R script, invoked explicitly by the caller:

```text
Rscript /app/run_lpca.R <input.csv> <scores.csv> <k=2> <m> <deviance.txt>
```

### Input

The CSV has a header row, a first column of nonempty, unique run identifiers,
and one column per binary edge feature. Rows are runs, not edges. The caller
must transpose the edge-by-run matrix returned by `summarize_networks` before
writing it.

This wrapper deliberately requires at least three runs, three edge features,
and three distinct binary network profiles. All feature values must be finite
numeric zeros or ones. Missing, nonnumeric, and nonbinary entries are errors.
Duplicate profiles and constant features are
retained when the input otherwise satisfies these requirements.

`k` must equal 2. `m` must be finite and strictly positive.

The default `m=6` is based on the exploratory
[LPCA-SPRAS experiments](https://github.com/Jeebjean/lpca-spras).
No single value of `m` was universally best.

### Outputs and numerical checks

The raw scores CSV contains `datapoint_labels`, `PC1`, and `PC2`, with one row
per input run and no synthetic centroid. It is an intermediate for Python's
PCA-compatible plotting and coordinate formatting, not a second public
coordinate table.

The deviance summary is a small text file:

```text
components: 2
m: <fixed positive value>
percent_deviance_explained: <100 times the package's whole-model statistic>
```

## SPRAS configuration

The `analysis.lpca` config settings are:

```yaml
analysis:
lpca:
include: false
aggregate_per_algorithm: false
k: 2
m: 6
labels: true
```

## Decomposition

`partial_decomp=TRUE` is always set. It uses
`rARPACK` (backed by `RSpectra`) for partial decompositions where supported. The
package can fall back to full eigendecomposition. It still constructs dense
edge-by-edge matrices, so memory use can grow quadratically with the number of
edge features. For example, the `egfr.yaml` configuration (19 runs, 18,851 edge features, `k=2`, `m=6`)
requires nearly 20GB memory. Completing this run may require modifying the memory
provided to the running container.

## Building and publishing the image

Replace `<tag>` below with the tag from `LPCA_CONTAINER_SUFFIX` in
`spras/analysis/lpca.py`. Publishing to the default registry requires access
to the `reedcompbio` Docker Hub organization:

docker build -t "reedcompbio/lpca:<tag>" docker-wrappers/LPCA/
docker push "reedcompbio/lpca:<tag>"

## Testing
Run the test Python script to test the Docker image
```commandline
python docker-wrappers/LPCA/test_container.py
```
Expected output will end with
```commandline
=== RESULT: 28 passed; 0 failed ===
Container exit status: 0
```

## AI
GPT 6 Astra was used to refactor these files and write the test code.
Loading
Loading