Fix cosine normalization for small-norm float32 vectors - #688
Open
shoaib0657 wants to merge 1 commit into
Open
shoaib0657 wants to merge 1 commit into
shoaib0657 wants to merge 1 commit into
Conversation
shoaib0657
marked this pull request as ready for review
October 2, 2026 00:26
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Addresses #685.
For a component around
1e-25, squaring it in float underflows to zero. The current normalization then produces cosine distances around-1e10and+1e10instead of0and2for identical and opposite vectors.This fixes normalization in both
IndexandBFIndex. It keeps float arithmetic for ordinary vectors and recomputes tiny or overflowing norms in double. Zero vectors keep their existing behavior, and normalized vectors are still stored as float32.Tests
Added regression tests for tiny and large vectors, different stored/query scales, float boundaries, and zero vectors. The tiny-vector test reproduces the original bug.
Local checks on Ubuntu WSL2 with GCC 13.3, Python 3.12.3, and NumPy 2.5.3:
make test: 25 tests passed.make examples: all five examples passed.git diff --check: passed.Performance
Measured
BFIndex.add_itemson 50,000 random float32 vectors of dimension 128: seed 685, one thread, one warmup, and 15 timed rounds per version. Hardware: Intel Core i5-11320H under WSL2. Both builds used-O3/C++11 without-march=native; the order of the versions was randomized each round.Insertion time is close to upstream on these ordinary inputs. The timing includes normalization and copying. HNSW build/query performance and workloads dominated by the double fallback were not measured.
Previously saved indexes affected by this bug need to be rebuilt from the original vectors.