All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog.
- Server-scale benchmark profile (#52): seeded 100K-vector dataset generator (
scripts/gen_dataset.py) and aserver_scaleharness measuring plain /working_dim=256/ cascade modes against exact f32 ground truth, with results, a 2K/10K/100K scale curve, and positioning in the new "Server scale" section ofdocs/BENCHMARK.md. Headline: compression holds at 4.78x/100K; recall is N-dependent (5-bit r@10 0.974 → 0.850) because true-neighbor margins collapse as N grows — vecq's measured sweet spot is the local/on-device profile up to ~10K vectors (r@10 0.932 @ 18 ms/q single-thread); the 2-bit cascade is not a single-threaded throughput win at server N.
- Crate metadata honesty pass (#49): new
vecq-coredescription — "Training-free vector quantization (4/5/6-bit) and search — the SQLite profile for edge vector storage" — plushomepage,keywords, andcategories, so the crates.io page reflects the current scope; README gains badges and benchmark-reproduce folds. - Dependency/CI bumps:
softprops/action-gh-release2 → 3 (#48), usearch 2.26.1 → 2.26.2 in the bench harness (#50). No library code changes since 0.3.0.
- Zero-copy read-only views:
VecqView::from_bytesparses any byte owner (mmap,Vec<u8>, …) without copying payloads — map+parse ~64 µs vs 4.9 ms full load at 12k vectors, results bit-identical to the loaded index (#25, #43) - Configurable Lloyd-Max width:
VecqIndex::set_bits(4|5|6)with 5-bit default — the compression/recall sweet spot (4.78x, recall@10 0.979 on real data). 4-bit stays available for maximum squeeze + cascade; 6-bit reaches residual-class recall at 25% less storage. File format v1.5 (width byte, plain non-4-bit only); 4-bit and residual outputs stay byte-identical (#39) - Opt-in residual quantization:
VecqIndex::with_residual, two-pass 4-bit codes, exact-norm two-term scoring, format v1.4 (#23) - Cascade search:
search_cascaderuns a 2-bit prefilter followed by 4-bit rescore — up to 3.6x faster than a plain 4-bit scan at iso-recall on 100k vectors (#22, #37) - Keyed API:
add_keyed(insert-or-replace under a stableu64key),remove_keyed(tombstones),search_keyed,compact, pluskey_of/contains_key/slots/tombstonesintrospection (#10, #16) - Multi-vectors-per-key (
add_keyed_multi,remove_keyed_at) andrelabel— keyed parity with usearch (#26, #33) - x86_64 AVX2 scoring path with runtime detection, bit-identical to the scalar/NEON paths (enforced by tests) and 4-vector batching (#11, #17)
- Matryoshka-aware
working_dimtruncation:VecqIndex::with_working_dim(dim, working_dim, seed)quantizes only the leading dims of Matryoshka-trained embeddings (#24, #30) - SQLite BLOB storage guide (
docs/SQLITE.md): schema shapes, save/load pattern, atomicity, measured latencies, pitfalls (#12, #18) - Head-to-head benchmark vs TurboQuant-MSE and RaBitQ at 4 bits (
vs_quantizersharness + results indocs/BENCHMARK.md) (#28, #34) - Per-architecture scoring-path table in
docs/BENCHMARK.md - mmap cold-start harness (
view_mmap): quantifies time-to-first-query for full load vs zero-copy view
VecqIndex::newnow defaults to 5-bit width (was 4-bit): ~19% more storage per vector in exchange for substantially higher recall; callset_bits(4)to restore the old default. This is the one behavioral change in the release- File format v1.2: the reserved header field now carries
working_dim(0 = full dim); payload layout identical to v1.1, readers still accept v1 and v1.1 - File format v1.3: keyed-slot table persisted so the key→slot map survives save/reload
- NEON batch kernel for 5/6-bit scoring (u64-window extraction + LUT gather), bit-identical to the scalar reference; 5-bit scan 4.27 → 3.21 ms/q
len()/is_empty()report live (non-tombstoned) vector counts;slots()reports total
- Keyed API now survives save/reload: file format v1.3 stores a keyed-slot table, and
from_bytesrestores the full key→slot map (#32)
- README rewritten for the width/view era: modes table, residual +
VecqView/mmap examples, persistence & serving guidance - Benchmark doc updated with the width matrix and mmap cold-start numbers
Note: the crates.io
vecq-core0.2.0 artifact contains only the version bump and CI fixes frommain; the feature entries below landed ondevelopafter the tag was cut and are therefore part of 0.3.0. Kept here for history.
- Keyed API:
add_keyed(insert-or-replace under a stableu64key),remove_keyed(tombstones),search_keyed,compact, pluskey_of/contains_key/slots/tombstonesintrospection (#10, #16) - Multi-vectors-per-key (
add_keyed_multi,remove_keyed_at) andrelabel— keyed parity with usearch (#26, #33) - x86_64 AVX2 scoring path with runtime detection, bit-identical to the scalar/NEON paths (enforced by tests) and 4-vector batching (#11, #17)
- Matryoshka-aware
working_dimtruncation:VecqIndex::with_working_dim(dim, working_dim, seed)quantizes only the leading dims of Matryoshka-trained embeddings (#24, #30) - SQLite BLOB storage guide (
docs/SQLITE.md): schema shapes, save/load pattern, atomicity, measured latencies, pitfalls (#12, #18) - Head-to-head benchmark vs TurboQuant-MSE and RaBitQ at 4 bits (
vs_quantizersharness + results indocs/BENCHMARK.md) (#28, #34) - Per-architecture scoring-path table in
docs/BENCHMARK.md
- File format v1.2: the reserved header field now carries
working_dim(0 = full dim); payload layout identical to v1.1, readers still accept v1 and v1.1 len()/is_empty()report live (non-tombstoned) vector counts;slots()reports total
- CLA bot exemption now matches actual bot logins (
dependabot[bot], notapp/dependabot) so dependabot PRs pass CI (#14)
- Release pipeline: push a
vX.Y.Ztag onmainto publishvecq-coreto crates.io and create the GitHub Release (adapted from the uteke release workflow).
First public release. Training-free 4-bit vector quantization and search.
- RHDH rotation (random diagonal sign + Walsh-Hadamard) with seeded header storage
- Lloyd-Max 4-bit scalar quantization, centroids embedded as constants
- Asymmetric scoring: f32 query against quantized database, per-vector scale correction
- Explicit NEON scoring path (aarch64), bit-identical to the scalar path (enforced by test)
- 4-vector batched NEON scoring with shared query loads
- Bounded min-heap top-k search (NaN-safe monotonic score keys, no per-query O(n) allocation)
- Single-file persistence, format v1.1 (f16 scales); readers accept v1 (f32 scales)
- Deterministic across platforms: fixed association order, no FMA contraction, seeded PRNG in header
- Benchmark suite vs usearch (
crates/vecq-bench): 0.89 ms/query, recall@10 0.958, 5.98x compression on real 768-dim embeddings
Initial spike: format v1, scalar scoring, first measurements.