Skip to content
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,13 @@ The project publishes 0.x prerelease versions; a stable release line is not yet

### Added

- Cosine HNSW indexes on `embeddings_text` (768), `embeddings_visual` (512),
and `embeddings_face` (512) via migration 0025 (`#173`). Text search walks
`ORDER BY embedding <=> query LIMIT n` (planner-usable) and falls back to
exact per-file `DISTINCT ON` when a bounded scan underfills after
deduplication. Visual cosine-order queries can use the visual index. Face
clustering remains in-process; the face index is DDL only. Recall is not
claimed here; the live harness is `#175`.
- Advertise the model-free lexical file-search route in the MCP `mem_search`
schema and verify `tools/list` plus route/filter forwarding through `tools/call`.

Expand Down
5 changes: 4 additions & 1 deletion SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -537,7 +537,10 @@ embeddings_face (
- `files` FTS + trigram — 显式指定 `route=lexical` 的文件名无模型词法召回;
`auto` 只融合 text/visual,不自动回退到 lexical,worker 不可用时仍报错。
仅搜索 `files.name`,路径只用于筛选,不检索文件正文或路径片段。
- `embeddings_* (embedding)` — pgvector HNSW
- `embeddings_* (embedding)` — pgvector HNSW (`vector_cosine_ops`, migration 0025)
on `embeddings_text` (768), `embeddings_visual` (512), `embeddings_face` (512).
Text search uses cosine-order candidates plus exact per-file fallback.
Face clustering is still in-process. `index_generation_vectors` is not indexed.
- `file_entities (entity_id)` — 反查"和某人有关的所有文件"

### 6.3 文件夹一致性规则(重要)
Expand Down
21 changes: 11 additions & 10 deletions docs/MIGRATION_SEQUENCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@ These draft changes are cumulative, not independently deployable:

| Order | Draft / original PR | Migration | Required predecessor |
| --- | --- | --- | --- |
| 1 | #194 / #183 | 0024 file lexical lane | released/main schema 23 |
| 2 | #197 / #180 | 0025 HNSW indexes | #194, schema 24 |
| 3 | #195 / #185 | 0026 data-plane hygiene | #197, schema 25 |
| 1 | #194 / #183 | 0024 file lexical lane | released/main schema 23 (merged) |
| 2 | #173 HNSW completion (supersedes #197 HOLD) | 0025 HNSW indexes + text continuation | #194, schema 24 |
| 3 | #195 / #185 | 0026 data-plane hygiene | schema 25 |

The PR base chain is `main` → `codex/fix-pr-183` → `codex/fix-pr-180`
→ `codex/fix-pr-185`. Successor branches must include their predecessor schema and source. Local
Expand All @@ -19,13 +19,14 @@ tags `v0.1.0` / `v0.1.1` contain only migrations 0001–0023. This does not prov
that a private deployment never applied a draft. Consequently migration
numbers and SQL identities are retained, not renumbered on an assumption.

HNSW DDL and shipping text-query planner acceptance are separate evidence:
creating indexes does not prove that a text query uses them. An index-only
successor must describe that partial scope and leave the broader #173 acceptance
open. #195 still requires the predecessor schema 25 regardless of query strategy.
This document does not approve a product decision or query-strategy change,
waive a review gate, or authorize deployment. #176's model-free file-lane RFC
also requires a maintainer decision before this draft is made review-ready.
Migration 0025 creates the three cosine HNSW indexes. The shipping text route
no longer uses `DISTINCT ON (f.id) ORDER BY f.id` as its primary plan: it walks
cosine-ordered candidates and falls back to that exact query only when a bounded
scan underfills. Visual cosine-order already matched HNSW. Face DDL is not a
face-query speedup. #195 still requires predecessor schema 25. Recall and live
latency remain `#175`, not this migration.

This document does not waive a review gate or authorize deployment.

Goose startup remains strict: no `WithAllowMissing` or equivalent option is
enabled. A database that already applied 26 while omitting 24/25 will correctly
Expand Down
60 changes: 60 additions & 0 deletions docs/VALIDATION_HNSW.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# HNSW index + text continuation for #173

Migration `0025_ann_hnsw_indexes.sql` adds cosine HNSW indexes to text (768),
visual (512), and face (512) embeddings. Main already shipped lexical
migration 0024; this branch's head is 25.

## What this change proves

- Populated 24 → 25 → 24 → 25 preserves text/visual/face vectors and rebuilds
three `VALID` `vector_cosine_ops` HNSW indexes (`TestHNSWMigrationPostgres`).
- Post-index INSERT succeeds; an UPDATE to the wrong dimension is rejected by
the `vector(N)` column type (failure mode: PostgreSQL dimension error, not a
silent pad/truncate).
- `EXPLAIN (ANALYZE)` of the shipping text cosine-order query and the visual
cosine-order query names `idx_embeddings_text_embedding_hnsw` and
`idx_embeddings_visual_embedding_hnsw` on a 2,000-row corpus. Planner
settings are not forced.
- `TestTextANNFileSemanticsPostgres` keeps best-chunk-per-file top-k when one
file owns 101 nearest chunks, and still enforces owner, literal path,
allow-list, MIME, and time filters. Invalid allow-lists fail closed.

## Text continuation / fallback

A bounded `ORDER BY distance LIMIT n` scan can underfill after per-file
deduplication (`ef_search=40` returning 40 chunks of one file). The shipping
path:

1. Run a CTE `ORDER BY embedding <=> $1 LIMIT remaining` on `embeddings_text`
(HNSW-compatible; omit `ANY(exclude)` when the exclude list is empty).
2. Join those candidates to `files` and apply owner/path/MIME/time filters.
3. Keep the first sighting of each file (that chunk is the file's best).
4. Repeat, excluding selected files, until k files are collected.
5. If a round returns no new files, fill the remainder with the original
exact `DISTINCT ON (f.id) ORDER BY f.id, distance` query.

Step 1 is the planner-usable shape. Step 4 preserves the previous result
contract on pathological corpora. Iterative-scan GUC is not enabled.

## What this change does not prove

- Live embedding quality, production latency, index build time, or numerical
recall. Those belong to [#175](https://github.com/bytefolk/mem/issues/175)
(shipping search-path producer) and the closed producer attempt
[#184](https://github.com/bytefolk/mem/pull/184). Fixture scores are not
substituted.
- Face query speedup. `assignCluster` still averages centroids in Go.
- `index_generation_vectors` ANN. The column is undimensioned.

## Local gates

```bash
MEM_TEST_DB="$MEM_TEST_DB" ./scripts/verify.sh integration
```

`run_hnsw_migration` creates a fresh `_test` database and runs
`TestHNSWMigrationPostgres`, which records EXPLAIN ANALYZE. `scripts/verify_hnsw_indexes.sh`
is a manual `psql` helper; CI does not call it because libpq rejects some pgx URIs.

Face evidence is valid DDL and populated-table migration, not a measured
face-query speedup.
16 changes: 14 additions & 2 deletions scripts/verify.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ set -euo pipefail

REPO_ROOT="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
MODE="${1:-unit}"
EXPECTED_MIGRATION_HEAD=24
EXPECTED_MIGRATION_HEAD=25
MIGRATION_ROLLBACK_TARGET=11
MODEL_TEXT_CANONICAL_BASE=15
WORKSPACE_AI_PROFILE_BASE=16
Expand Down Expand Up @@ -349,6 +349,7 @@ run_postgres_tests() {
TestManagedAISettlementOutboxPostgres
TestReleasedFileStageRetryPostgres
TestDurableContextPostgres
TestTextANNFileSemanticsPostgres
)

integration_log="$(mktemp "${TMPDIR:-/tmp}/mem-integration.XXXXXX")"
Expand All @@ -359,7 +360,7 @@ run_postgres_tests() {
MEM_TEST_DB="$MEM_TEST_DB" go test \
${race_flag:+"$race_flag"} \
-v -count=1 -p 1 -timeout 20m \
-run '^(TestMemoryPostgres|TestHandoffPostgres|TestWorkspaceTransferPostgres|TestWorkspaceTransferMergeConservativePostgres|TestHandoffCrossAgentHTTPIntegration|TestRelocateHTTPPostgres|TestMemoryPathLifecycleIntegration|TestWorkspacePathLockingIntegration|TestFilePathLockingIntegration|TestAnnotationDecisionIntegration|TestIndexerEnrichmentIntegration|TestRecomputePerson|TestManagedEmbeddingEntitlementPostgres|TestManagedSearchReplayPostgres|TestManagedEmbeddingHTTPAuthorizationPostgres|TestAIProfilePostgres|TestIndexGenerationPostgres|TestManagedAISettlementOutboxPostgres|TestReleasedFileStageRetryPostgres|TestDurableContextPostgres)$' \
-run '^(TestMemoryPostgres|TestHandoffPostgres|TestWorkspaceTransferPostgres|TestWorkspaceTransferMergeConservativePostgres|TestHandoffCrossAgentHTTPIntegration|TestRelocateHTTPPostgres|TestMemoryPathLifecycleIntegration|TestWorkspacePathLockingIntegration|TestFilePathLockingIntegration|TestAnnotationDecisionIntegration|TestIndexerEnrichmentIntegration|TestRecomputePerson|TestManagedEmbeddingEntitlementPostgres|TestManagedSearchReplayPostgres|TestManagedEmbeddingHTTPAuthorizationPostgres|TestAIProfilePostgres|TestIndexGenerationPostgres|TestManagedAISettlementOutboxPostgres|TestReleasedFileStageRetryPostgres|TestDurableContextPostgres|TestTextANNFileSemanticsPostgres)$' \
./internal/memory \
./internal/handoff \
./internal/workspacetransfer \
Expand Down Expand Up @@ -397,9 +398,20 @@ run_integration() {
validate_test_database
with_fresh_test_database migration_sequence run_migration_sequence
with_fresh_test_database migration run_migration_round_trip
with_fresh_test_database hnsw_migration run_hnsw_migration
with_fresh_test_database integration run_postgres_integration
}

run_hnsw_migration() {
log "Populated HNSW migration, rollback, ingest, dimension rejection and EXPLAIN"
(
cd "${REPO_ROOT}/server"
MEM_HNSW_TEST_DB="$MEM_TEST_DB" go test -v -count=1 \
-run '^TestHNSWMigrationPostgres$' ./internal/db
)
log "Planner EXPLAIN is recorded by TestHNSWMigrationPostgres (psql URI script is manual)"
}

run_integration_race() {
validate_test_database
with_fresh_test_database integration_race run_postgres_integration_race
Expand Down
58 changes: 58 additions & 0 deletions scripts/verify_hnsw_indexes.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
#!/usr/bin/env bash
# Read-only planner verification for shipping text and visual query shapes.
# Requires a populated disposable database; does not force planner settings.
set -euo pipefail
trap 'echo "ERROR: HNSW verification aborted on an execution error; assertions are incomplete" >&2' ERR
DB_URL="${1:?Usage: $0 <database-url> <corpus-user-uuid>}"
CORPUS_USER="${2:?Supply the user UUID that owns the populated corpus}"
[[ "$CORPUS_USER" =~ ^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$ ]] || {
echo 'ERROR: corpus user must be a UUID' >&2; exit 1;
}
sql() { psql -X -A -t -v ON_ERROR_STOP=1 "$DB_URL" -c "$1"; }
db_name="$(sql 'SELECT current_database()')"
[[ "$db_name" == *_test ]] || { echo 'ERROR: database must end in _test' >&2; exit 1; }
pass=0
fail=0
for kind in text visual face; do
index="idx_embeddings_${kind}_embedding_hnsw"
valid="$(sql "SELECT count(*) FROM pg_index i
JOIN pg_class c ON c.oid = i.indexrelid JOIN pg_am a ON a.oid = c.relam
WHERE i.indrelid = 'embeddings_${kind}'::regclass
AND c.relname = '${index}' AND a.amname = 'hnsw' AND i.indisvalid
AND pg_get_indexdef(i.indexrelid) LIKE '%vector_cosine_ops%'")"
rows="$(sql "SELECT count(*) FROM embeddings_${kind} e JOIN files f ON f.id=e.file_id
WHERE f.user_id='${CORPUS_USER}'::uuid AND e.embedding IS NOT NULL")"
if [[ "$valid" == 1 && "$rows" -gt 0 ]]; then
echo "PASS: ${index} is valid; corpus contains ${rows} non-null vectors"
pass=$((pass + 1))
else
echo "FAIL: ${index}: valid=${valid}, corpus vectors=${rows}"
fail=$((fail + 1))
fi
done
assert_plan() {
local route="$1" query="$2" plan
plan="$(sql "EXPLAIN (ANALYZE, BUFFERS) ${query}")"
echo "$plan"
if grep -q "idx_embeddings_${route}_embedding_hnsw" <<<"$plan"; then
echo "PASS: shipping ${route} query uses HNSW"
pass=$((pass + 1))
else
echo "FAIL: shipping ${route} query does not use HNSW"
fail=$((fail + 1))
fi
}
# Match queryTextDistanceOrder: ANN CTE then file join. DISTINCT ON is fallback.
assert_plan text "WITH nearest AS (
SELECT e.id, e.file_id, e.embedding <=> array_fill(0.1::real, ARRAY[768])::vector AS dist
FROM embeddings_text e
ORDER BY e.embedding <=> array_fill(0.1::real, ARRAY[768])::vector ASC
LIMIT 10
)
SELECT e.id, f.id FROM nearest e JOIN files f ON f.id=e.file_id
WHERE f.user_id='${CORPUS_USER}'::uuid ORDER BY e.dist ASC"
assert_plan visual "SELECT e.file_id
FROM embeddings_visual e
ORDER BY e.embedding <=> array_fill(0.1::real, ARRAY[512])::vector ASC LIMIT 10"
echo "Results: ${pass} passed, ${fail} failed"
[[ "$fail" -eq 0 ]]
Loading
Loading