Skip to content

Warn about repeated triplets only when chance cannot explain them - #16

Open
snarles wants to merge 1 commit into
LukasMut:mainfrom
snarles:duplicate-triplet-warning
Open

snarles wants to merge 1 commit into
LukasMut:mainfrom
snarles:duplicate-triplet-warning

Conversation

@snarles

@snarles snarles commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

load_data now issues warnings when there are duplicated rows in excess of what would be be typically seen in uniformly sampled triplets.

The check reports the observed repeat count, the count expected under an explicit null, and the probability of seeing at least that many, and warns only when that probability is below 0.2.

Null: the N trials are drawn uniformly from the K = C(M,3) possible triplets, with trial identity the unordered set {i,j,k}. The statistic is W, the number of colliding pairs. The collision indicators are pairwise uncorrelated, so W matches Binomial(C(N,2), 1/K) in both moments exactly, which is what makes the binomial reference right rather than merely convenient. One-sided: a deficit of repeats just means sampling without replacement, which main_tripletize does deliberately.

The expected counts are written with log1p/expm1 throughout. This is not cosmetic: the textbook form of the expected number of repeated triplets cancels catastrophically, returning a negative count at M=5000/N=100 and a value five orders of magnitude too large at M=20000/N=50. The p-value likewise underflows float64 well before the evidence stops being readable, and binom.logsf returns -inf there, so a log-scale Poisson tail keeps the reported log10 p finite.

Data is never altered or de-duplicated, and the check never raises.

load_data warned on any duplicated row, which on realistic data is mostly noise: a clean uniform sample of 200k triplets over the 1854 THINGS objects contains ~38 repeated rows by the birthday paradox alone. Warning there teaches users to ignore the warning.

The check now reports the observed repeat count, the count expected under an explicit null, and the probability of seeing at least that many, and warns only when that probability is below 0.2.

Null: the N trials are drawn uniformly from the K = C(M,3) possible triplets, with trial identity the unordered set {i,j,k}. The statistic is W, the number of colliding pairs. The collision indicators are pairwise uncorrelated, so W matches Binomial(C(N,2), 1/K) in both moments exactly, which is what makes the binomial reference right rather than merely convenient. One-sided: a deficit of repeats just means sampling without replacement, which main_tripletize does deliberately.

The expected counts are written with log1p/expm1 throughout. This is not cosmetic: the textbook form of the expected number of repeated triplets cancels catastrophically, returning a negative count at M=5000/N=100 and a value five orders of magnitude too large at M=20000/N=50. A fractions.Fraction reference pins this in the tests. The p-value likewise underflows float64 well before the evidence stops being readable, and binom.logsf returns -inf there, so a log-scale Poisson tail keeps the reported log10 p finite.

Data is never altered or de-duplicated, and the check never raises.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant