Skip to content

perf(mem_wal): rank only prefiltered, newest memtable rows in a filtered vector search - #9812

Open
hamersaw wants to merge 2 commits into
lance-format:mainfrom
hamersaw:feat/mem-wal-vector-prefilter-index
Open

hamersaw wants to merge 2 commits into
lance-format:mainfrom
hamersaw:feat/mem-wal-vector-prefilter-index

Conversation

@hamersaw

@hamersaw hamersaw commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #9806 (#9804 → #9805 → #9806). Only the last two commits are new here: d2d62dbdf (perf) and 77781becf (test). Everything below them is the stack. I'll rebase onto main once it merges.

Problem

With a prefilter, a memtable vector search skips the HNSW and runs MemTableBruteForceVectorExec. For every visible row it:

  • hashes every primary key to find the newest versions, even when no key was ever rewritten;
  • computes a distance for every row, then applies the filter;
  • does all of this synchronously inside execute().

On a 100k-row memtable with a 1/16 filter, this takes about 300ms in a debug build. In a release-build LanceDB server it adds about 100ms to every filtered vector query on a WAL table.

Fix

  • Index candidates: plan_vector_search splits the prefilter with feat(mem_wal): answer memtable filters with Lance's scalar expression pass #9804's plan_filter, the same split filtered scans use. At execution, evaluate_index_filter narrows the candidate rows. The predicate is re-applied only to candidates the indexes didn't settle exactly. If an index search fails, the exec logs a warning and filters every row.
  • Newest version:
    • skipped when the key index reports no rewrites (pk_has_overrides, read at execution time);
    • one key-index seek per survivor when at most 1/8 of visible rows pass the filter, the same crossover the filter indexes' newest-only reads use;
    • otherwise, the existing pass that hashes every key.
  • Distances: computed only for surviving rows. When at least half of a batch survives, the whole vector column is used rather than copying rows out.
  • Threading: ranking and materialization run on the CPU pool via spawn_cpu, inside the stream.

Results

Debug build, 100k rows × 128 dims, median of the ignored time_a_prefiltered_vector_search_over_a_large_memtable. Every filter took about 300ms before.

filter no rewrites 1,000 rewrites
category = 'cat03' (indexed, 1/17) 24ms 28ms
id % 16 = 3 (not indexed, 1/16) 29ms 33ms
category < 'cat08' (about 47% of rows) 162ms 304ms

Broad filters with rewrites are unchanged. Fixing them needs a mask on the memtable HNSW search (index/hnsw.rs), left for a follow-up.

Notes for the stack

  • Index candidates are gated on the existing use_index option, not memtable_filter_indexes, which the LSM vector path (scanner/vector_search.rs) never sets.
  • There's no separate handling for unindexed rows. Visibility is derived from indexed_count, so every visible row is in the filter indexes.
  • evaluate_index_filter is called without a match budget, so a broad match is listed in full.

Tests

  • New a_prefiltered_vector_search_ranks_only_the_newest_matching_versions checks ids and distances against an independent reference. It covers:

    • indexed exact, IN, broad, partly indexed and unindexed filters;
    • rewrites and deletes;
    • rows indexed but not yet visible;
    • distance bounds.

    It fails if the newest-version check is skipped, or if the filter isn't re-applied to unsettled candidates.

  • cargo test -p lance --lib dataset::mem_wal: 883 passed. cargo fmt --check is clean. cargo clippy --all --tests --benches -- -D warnings reports one chunks_exact lint in index/vector/ivf.rs test code that comes from the stack's base and is already gone on main. These commits don't touch that file.

🤖 Generated with Claude Code

@hamersaw
hamersaw force-pushed the feat/mem-wal-vector-prefilter-index branch 2 times, most recently from 77781be to aff7b05 Compare October 10, 2026 13:22
@hamersaw
hamersaw marked this pull request as ready for review October 10, 2026 13:35

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added K-approved Latest Gatekeeper recommendation permits acceptance. K-risk Latest Gatekeeper recommendation includes a non-blocking risk. labels Oct 10, 2026
@hamersaw
hamersaw force-pushed the feat/mem-wal-vector-prefilter-index branch from aff7b05 to 95f44d1 Compare October 11, 2026 00:20
@lance-gatekeeper lance-gatekeeper Bot removed K-approved Latest Gatekeeper recommendation permits acceptance. K-risk Latest Gatekeeper recommendation includes a non-blocking risk. labels Oct 11, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 11, 2026
hamersaw and others added 2 commits October 11, 2026 04:02
…red vector search

A filtered vector search over the memtable runs `MemTableBruteForceVectorExec`,
which on every query hashed every visible key to find each one's newest
version, computed a distance for every visible row, applied the filter only
afterwards, and did all of it synchronously inside `execute`. On 100,000
rows that is the whole cost of the search, however selective the filter.

Narrow the rows before any distance is computed:

- Ask the filter indexes for candidates, through the same split of the
  filter into index searches and a leftover that a filtered scan uses, and
  apply the filter only to candidates the indexes did not settle.
- Keep each key's newest version without the pass over every key when no key
  was ever rewritten (the condition that already lets an unfiltered search use
  HNSW), and with one key-index seek per row when few rows pass the filter.
- Compute distances only for the rows left, or for the whole batch when most
  of it is left.
- Rank on the CPU pool from inside the stream rather than in `execute`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ected answers

Ranks a prefiltered search over a keyed memtable and compares ids and
distances with the nearest keys computed directly from the rows written:
filters an index answers exactly, partly, or not at all, selective and broad,
across rewrites, deletes, rows indexed but not yet visible, and distance
bounds. An ignored test times a prefiltered search over 100,000 rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@hamersaw
hamersaw force-pushed the feat/mem-wal-vector-prefilter-index branch from 95f44d1 to 09d6061 Compare October 11, 2026 09:02
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 11, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

The whole-batch distance path removes the earlier gathering overhead. Selective searches retain substantial gains with correct newest-visible-version and prefilter semantics, while broad-filter latency stays close to the base.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

K-approved Latest Gatekeeper recommendation permits acceptance. performance

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant