test: deflake the two intermittently-failing 1.19 CI tests - #137
Merged
Merged
Conversation
Two pre-existing flaky tests were intermittently failing on the loaded 1.19 CI runners, blocking unrelated PRs. Both are fixed at the root, not by loosening blindly: - OtelAdversarialFuzzTest "linear, not quadratic": the single-shot :timer.tc ratio flaked when a GC pause inflated one sample. Now takes the MIN of 5 runs per size — the fastest run is the least-contended measurement, so noise can't inflate the ratio, while a genuinely quadratic accumulation still shows ~16x (the algorithmic factor dominates the min). - MathOracleTest float sum/mean vs Polars: the tolerance was relative to the RESULT, which collapses under catastrophic cancellation — naive left-to-right summation (pyex/CPython) and Polars' pairwise/SIMD summation legitimately diverge by ~N·ε·Σ|xᵢ| when signed terms cancel (the result is near zero while the partial sums are large). New `assert_sum_close/3` scales the tolerance to the INPUT magnitude (Σ|xᵢ|), bounding the algorithmic divergence regardless of cancellation while still catching a genuinely wrong sum. Test-only; no library changes. Full suite green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019NokzcR7BiAigPgC78zpk9
ivarvong
added a commit
that referenced
this pull request
Jul 2, 2026
…nhenge The fixture's default 5s wall-clock timeout flaked on a starved 1.19 CI runner (passed on rerun, and on both sibling matrix jobs from the same commit). Same deflake shape as #137 and the existing options.json users. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KuxYKoh8pXEna5ohZUgJsn
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Two pre-existing flaky tests were intermittently failing on the loaded Elixir 1.19 CI runners — they cost #136 two re-runs (a different one flaked each time), and they tax every PR. Fixed at the root, not by blindly loosening.
Fixes
OtelAdversarialFuzzTest"attribute/event accumulation is linear, not quadratic" — a single-shot:timer.tcratio (time(40k)/time(10k) < 8.0). A GC pause or scheduler hiccup on one sample inflated the ratio. Now takes the min of 5 runs per size: the fastest run is the least-contended measurement, so noise can't push the ratio up — while a genuinely quadratic accumulation still shows ~16× (the algorithmic factor dominates the min, so the test keeps its teeth).MathOracleTestfloatsum/meanvs Polars — the tolerance was relative to the result, which collapses under catastrophic cancellation. pyex'ssum()is naive left-to-right (like CPython); Polars is pairwise/SIMD — they legitimately diverge by ~N·ε·Σ|xᵢ|when signed terms cancel (the result is near zero while the partial sums are large), so a result-relative tolerance goes to ~1e-6and occasionally trips. Newassert_sum_close/3scales the tolerance to the input magnitude (Σ|xᵢ|), which bounds the algorithmic divergence regardless of cancellation while still catching a genuinely wrong sum.Scope / verification
Test-only; no library changes. Hammered locally — the OTel test passes repeatedly, and
math_oracle_test(42 properties) is clean across seeds. Full suite green.🤖 Generated with Claude Code
https://claude.ai/code/session_019NokzcR7BiAigPgC78zpk9