Skip to content

Make the pydeseq2 Milo solver reproduce R Milo - #1110

Merged
Zethson merged 1 commit into
mainfrom
milo-pydeseq2-lrt
Sep 23, 2026
Merged

Zethson merged 1 commit into
mainfrom
milo-pydeseq2-lrt

Conversation

@Zethson

@Zethson Zethson commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

The pydeseq2 solver of da_nhoods now normalises like R Milo (library size times TMM factors, ported from edgeR), keeps the dispersions of pydeseq2 and tests with a likelihood ratio test instead of the Wald test, reporting the log fold change with the prior count of edgeR.
The Wald test and the floor of 0.5 that DESeq2 puts on fitted means lost almost all power in neighbourhoods a group barely populates, so the solver missed every Asymptomatic neighbourhood miloR calls on the tutorial data.
Follows #1109; the tables compare with miloR 2.8.1 on identical neighbourhood counts (Stephenson COVID data, 113 donors, 4307 neighbourhoods) and on a null and planted-depletion benchmark built from the same neighbourhoods.

Design logFC r with miloR p-value ρ with miloR SpatialFDR < 0.1 (shared with miloR) miloR
~Status 0.975 → 1.000 0.976 → 0.994 858 (743) → 817 (737) 803
~Site+Status 0.959 → 0.992 0.965 → 0.982 915 (810) → 912 (846) 926
~Site+COVID_severity, COVID_severityAsymptomatic 0.851 → 0.995 0.967 → 0.982 1 (0) → 72 (71) 144
~Site+COVID_severity_continuous 0.745 → 0.994 0.987 → 0.991 795 (679) → 803 (726) 747
Benchmark miloR pydeseq2 before → after
Planted 90% depletion, 10 donors per group: recall 0.50 0.31 → 0.48
Same: false discovery proportion 0.098 0.038 → 0.045
Permuted labels, 113 donors: false calls per run 14.8 3.2 → 3.8
10 against 10 healthy donors: false calls per run (20 runs) 5.85 1.10 → 3.05
5 against 5 healthy donors: false calls per run (20 runs) 11.45 0 → 6.30

@Zethson
Zethson changed the base branch from milo-match-r to main September 23, 2026 12:12
@Zethson Zethson closed this Sep 23, 2026
@Zethson Zethson reopened this Sep 23, 2026
@codecov-commenter

codecov-commenter commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.92308% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.10%. Comparing base (c7e0f02) to head (370d1cd).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/pertpy/tools/_milo.py 96.72% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #1110      +/-   ##
==========================================
+ Coverage   79.98%   80.10%   +0.12%     
==========================================
  Files          55       55              
  Lines        7518     7571      +53     
==========================================
+ Hits         6013     6065      +52     
- Misses       1505     1506       +1     
Files with missing lines Coverage Δ
src/pertpy/tools/_milo_glmm.py 99.46% <100.00%> (+<0.01%) ⬆️
src/pertpy/tools/_milo.py 79.85% <96.72%> (+1.26%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The pydeseq2 solver tested with the Wald test of DESeq2, whose fitted means
are floored at 0.5. Both cost almost all power in neighbourhoods that one
group barely populates, which are common in Milo, so on the Stephenson data
the solver found none of the 144 Asymptomatic neighbourhoods miloR calls.

The solver now normalises like R Milo with library size times TMM factors,
ported from calcNormFactors of edgeR and matching it to machine precision,
keeps the MAP dispersions of pydeseq2 and tests with a likelihood ratio test
fitted without the floor. The reported log fold change comes from a refit
with the prior count of edgeR, so it stays finite when a group has no cells.
Cook's outlier handling is dropped because edgeR has none. logCPM now is a
log CPM instead of the baseMean of DESeq2, shared with the mixed model.

On identical neighbourhood counts of the Stephenson data the overlap of
significant neighbourhoods with miloR rises from 0.79-0.81 to 0.84-0.88
(Jaccard) and to 71 of miloR's 144 Asymptomatic calls. On a benchmark built
from the same neighbourhoods its power matches miloR (recall 0.48 against
0.50) while it makes fewer false calls under null data (3.8 against 14.8
with permuted labels, 3.1 against 5.9 with ten against ten healthy donors).
The remaining gap is the dispersion estimate, which would need edgeR itself.

test_da_nhoods_default_contrast compares p-values instead of SpatialFDR,
which is constant on its six sample fixture with the new test.
@Zethson
Zethson merged commit 794bb43 into main Sep 23, 2026
22 checks passed
@Zethson
Zethson deleted the milo-pydeseq2-lrt branch September 23, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants