Conversation
load_data warned on any duplicated row, which on realistic data is mostly noise: a clean uniform sample of 200k triplets over the 1854 THINGS objects contains ~38 repeated rows by the birthday paradox alone. Warning there teaches users to ignore the warning.
The check now reports the observed repeat count, the count expected under an explicit null, and the probability of seeing at least that many, and warns only when that probability is below 0.2.
Null: the N trials are drawn uniformly from the K = C(M,3) possible triplets, with trial identity the unordered set {i,j,k}. The statistic is W, the number of colliding pairs. The collision indicators are pairwise uncorrelated, so W matches Binomial(C(N,2), 1/K) in both moments exactly, which is what makes the binomial reference right rather than merely convenient. One-sided: a deficit of repeats just means sampling without replacement, which main_tripletize does deliberately.
The expected counts are written with log1p/expm1 throughout. This is not cosmetic: the textbook form of the expected number of repeated triplets cancels catastrophically, returning a negative count at M=5000/N=100 and a value five orders of magnitude too large at M=20000/N=50. A fractions.Fraction reference pins this in the tests. The p-value likewise underflows float64 well before the evidence stops being readable, and binom.logsf returns -inf there, so a log-scale Poisson tail keeps the reported log10 p finite.
Data is never altered or de-duplicated, and the check never raises.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
load_data now issues warnings when there are duplicated rows in excess of what would be be typically seen in uniformly sampled triplets.
The check reports the observed repeat count, the count expected under an explicit null, and the probability of seeing at least that many, and warns only when that probability is below 0.2.
Null: the N trials are drawn uniformly from the K = C(M,3) possible triplets, with trial identity the unordered set {i,j,k}. The statistic is W, the number of colliding pairs. The collision indicators are pairwise uncorrelated, so W matches Binomial(C(N,2), 1/K) in both moments exactly, which is what makes the binomial reference right rather than merely convenient. One-sided: a deficit of repeats just means sampling without replacement, which main_tripletize does deliberately.
The expected counts are written with log1p/expm1 throughout. This is not cosmetic: the textbook form of the expected number of repeated triplets cancels catastrophically, returning a negative count at M=5000/N=100 and a value five orders of magnitude too large at M=20000/N=50. The p-value likewise underflows float64 well before the evidence stops being readable, and binom.logsf returns -inf there, so a log-scale Poisson tail keeps the reported log10 p finite.
Data is never altered or de-duplicated, and the check never raises.