Make the pydeseq2 Milo solver reproduce R Milo - #1110
Merged
Merged
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1110 +/- ##
==========================================
+ Coverage 79.98% 80.10% +0.12%
==========================================
Files 55 55
Lines 7518 7571 +53
==========================================
+ Hits 6013 6065 +52
- Misses 1505 1506 +1
🚀 New features to boost your workflow:
|
The pydeseq2 solver tested with the Wald test of DESeq2, whose fitted means are floored at 0.5. Both cost almost all power in neighbourhoods that one group barely populates, which are common in Milo, so on the Stephenson data the solver found none of the 144 Asymptomatic neighbourhoods miloR calls. The solver now normalises like R Milo with library size times TMM factors, ported from calcNormFactors of edgeR and matching it to machine precision, keeps the MAP dispersions of pydeseq2 and tests with a likelihood ratio test fitted without the floor. The reported log fold change comes from a refit with the prior count of edgeR, so it stays finite when a group has no cells. Cook's outlier handling is dropped because edgeR has none. logCPM now is a log CPM instead of the baseMean of DESeq2, shared with the mixed model. On identical neighbourhood counts of the Stephenson data the overlap of significant neighbourhoods with miloR rises from 0.79-0.81 to 0.84-0.88 (Jaccard) and to 71 of miloR's 144 Asymptomatic calls. On a benchmark built from the same neighbourhoods its power matches miloR (recall 0.48 against 0.50) while it makes fewer false calls under null data (3.8 against 14.8 with permuted labels, 3.1 against 5.9 with ten against ten healthy donors). The remaining gap is the dispersion estimate, which would need edgeR itself. test_da_nhoods_default_contrast compares p-values instead of SpatialFDR, which is constant on its six sample fixture with the new test.
Zethson
force-pushed
the
milo-pydeseq2-lrt
branch
from
September 23, 2026 12:52
1b7961e to
370d1cd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The pydeseq2 solver of
da_nhoodsnow normalises like R Milo (library size times TMM factors, ported from edgeR), keeps the dispersions of pydeseq2 and tests with a likelihood ratio test instead of the Wald test, reporting the log fold change with the prior count of edgeR.The Wald test and the floor of 0.5 that DESeq2 puts on fitted means lost almost all power in neighbourhoods a group barely populates, so the solver missed every Asymptomatic neighbourhood miloR calls on the tutorial data.
Follows #1109; the tables compare with miloR 2.8.1 on identical neighbourhood counts (Stephenson COVID data, 113 donors, 4307 neighbourhoods) and on a null and planted-depletion benchmark built from the same neighbourhoods.
~Status~Site+Status~Site+COVID_severity,COVID_severityAsymptomatic~Site+COVID_severity_continuous