Repository navigation
Conversation
hamersaw
force-pushed
the
feat/mem-wal-vector-prefilter-index
branch
2 times, most recently
from
October 10, 2026 13:22
77781be to
aff7b05
Compare
hamersaw
marked this pull request as ready for review
October 10, 2026 13:35
hamersaw
force-pushed
the
feat/mem-wal-vector-prefilter-index
branch
from
October 11, 2026 00:20
aff7b05 to
95f44d1
Compare
…red vector search A filtered vector search over the memtable runs `MemTableBruteForceVectorExec`, which on every query hashed every visible key to find each one's newest version, computed a distance for every visible row, applied the filter only afterwards, and did all of it synchronously inside `execute`. On 100,000 rows that is the whole cost of the search, however selective the filter. Narrow the rows before any distance is computed: - Ask the filter indexes for candidates, through the same split of the filter into index searches and a leftover that a filtered scan uses, and apply the filter only to candidates the indexes did not settle. - Keep each key's newest version without the pass over every key when no key was ever rewritten (the condition that already lets an unfiltered search use HNSW), and with one key-index seek per row when few rows pass the filter. - Compute distances only for the rows left, or for the whole batch when most of it is left. - Rank on the CPU pool from inside the stream rather than in `execute`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ected answers Ranks a prefiltered search over a keyed memtable and compares ids and distances with the nearest keys computed directly from the rows written: filters an index answers exactly, partly, or not at all, selective and broad, across rewrites, deletes, rows indexed but not yet visible, and distance bounds. An ignored test times a prefiltered search over 100,000 rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hamersaw
force-pushed
the
feat/mem-wal-vector-prefilter-index
branch
from
October 11, 2026 09:02
95f44d1 to
09d6061
Compare
Contributor
There was a problem hiding this comment.
✅ Gate recommendation: approve.
The whole-batch distance path removes the earlier gathering overhead. Selective searches retain substantial gains with correct newest-visible-version and prefilter semantics, while broad-filter latency stays close to the base.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With a prefilter, a memtable vector search skips the HNSW and runs
MemTableBruteForceVectorExec. For every visible row it:execute().On a 100k-row memtable with a 1/16 filter, this takes about 300ms in a debug build. In a release-build LanceDB server it adds about 100ms to every filtered vector query on a WAL table.
Fix
plan_vector_searchsplits the prefilter with feat(mem_wal): answer memtable filters with Lance's scalar expression pass #9804'splan_filter, the same split filtered scans use. At execution,evaluate_index_filternarrows the candidate rows. The predicate is re-applied only to candidates the indexes didn't settle exactly. If an index search fails, the exec logs a warning and filters every row.pk_has_overrides, read at execution time);spawn_cpu, inside the stream.Results
Debug build, 100k rows × 128 dims, median of the ignored
time_a_prefiltered_vector_search_over_a_large_memtable. Every filter took about 300ms before.category = 'cat03'(indexed, 1/17)id % 16 = 3(not indexed, 1/16)category < 'cat08'(about 47% of rows)Broad filters with rewrites are unchanged. Fixing them needs a mask on the memtable HNSW search (
index/hnsw.rs), left for a follow-up.Notes for the stack
use_indexoption, notmemtable_filter_indexes, which the LSM vector path (scanner/vector_search.rs) never sets.indexed_count, so every visible row is in the filter indexes.evaluate_index_filteris called without a match budget, so a broad match is listed in full.Tests
New
a_prefiltered_vector_search_ranks_only_the_newest_matching_versionschecks ids and distances against an independent reference. It covers:IN, broad, partly indexed and unindexed filters;It fails if the newest-version check is skipped, or if the filter isn't re-applied to unsettled candidates.
cargo test -p lance --lib dataset::mem_wal: 883 passed.cargo fmt --checkis clean.cargo clippy --all --tests --benches -- -D warningsreports onechunks_exactlint inindex/vector/ivf.rstest code that comes from the stack's base and is already gone onmain. These commits don't touch that file.🤖 Generated with Claude Code