Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,6 +101,8 @@ The project publishes 0.x prerelease versions; a stable release line is not yet

### Fixed

- Add cosine HNSW indexes for legacy embedding tables while explicitly preserving the exact per-file text-search semantics and keeping the broader planner-use acceptance open; see `docs/VALIDATION_HNSW.md`.

- Follow the shared design language for reading, numeric and action alignment; generate the existing Web color variables from a pinned design-system token snapshot, and use a single consistent empty-state pattern. Refs #211.

- Improve web caption and status contrast in both themes, including tinted danger
Expand Down
7 changes: 6 additions & 1 deletion SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -537,7 +537,12 @@ embeddings_face (
- `files` FTS + trigram — 显式指定 `route=lexical` 的文件名无模型词法召回;
`auto` 只融合 text/visual,不自动回退到 lexical,worker 不可用时仍报错。
仅搜索 `files.name`,路径只用于筛选,不检索文件正文或路径片段。
- `embeddings_* (embedding)` — pgvector HNSW
- `embeddings_* (embedding)` — cosine pgvector HNSW (migration 0025).
Index availability does not guarantee route use: text search still selects
the best chunk per file exactly before top-k, visual planning depends on the
corpus and filters, and face clustering remains in Go. See
[HNSW validation](docs/VALIDATION_HNSW.md); #173 planner/recall acceptance is
not completed by the index migration.
- `file_entities (entity_id)` — 反查"和某人有关的所有文件"

### 6.3 文件夹一致性规则(重要)
Expand Down
144 changes: 144 additions & 0 deletions docs/VALIDATION_HNSW.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
# HNSW index foundation — partial delivery for #173

Migration `0025_ann_hnsw_indexes.sql` adds cosine HNSW indexes to text (768),
visual (512), and face (512) embeddings. Migration 0024 belongs to #183 and
0026 to #185. The integration runner declares this branch's head as 25.

## Scope and acceptance

This is the index foundation portion of #173, tracked with **Refs #173**.
It does not complete that issue. The shipping text route remains exact and is
not rewritten: a bounded ANN candidate set cannot establish best-chunk-per-file
results merely by returning k distinct files. Underfill fallback alone also
does not establish that unseen files or better chunks cannot outrank them.
A retrieval-policy change needs its own continuation/exhaustion/fallback and
quality acceptance. The strict planner gate below is retained unchanged in
meaning, and still fails for text.

The independently testable criteria for this partial change are:

- a populated 24 → 25 → 24 → 25 migration preserves text/visual/face vectors;
- all three valid indexes use HNSW and cosine operators;
- ingest and ordinary startup work after the migration; wrong dimensions fail;
- shipping text search still returns k distinct eligible files with the best
chunk even when one file holds 101 nearest chunks; owner, literal path,
allow-list, MIME and time boundaries remain enforced;
- the separate full-issue planner check cannot report success when text misses
HNSW. Face DDL is never reported as a shipping face search improvement.

This branch requires #194's lexical migration 0024; the
[cumulative migration sequence](MIGRATION_SEQUENCE.md) is included by that base.
Deployment order stays #194 → #197 → #195. Accepting this partial scope is a
review decision, not a claim that the remaining #173 planner criterion passed.

## Reproducible local gates

```bash
# Creates separate owned test databases, including a populated HNSW round trip,
# and runs the shipping text semantics regression. MEM_TEST_DB ends in _test.
MEM_TEST_DB="$MEM_TEST_DB" ./scripts/verify.sh integration
MEM_TEST_DB="$MEM_TEST_DB" ./scripts/verify.sh integration-race
```

`TestHNSWMigrationPostgres` requires a fresh database via `MEM_HNSW_TEST_DB`
and refuses existing Goose history. It seeds 2,000 vectors in each table before
index construction, checks rollback preservation, repeats the upgrade, writes
one more vector in each table, rejects wrong dimensions, and calls `DB.Migrate`.
`TestTextANNFileSemanticsPostgres` exercises the actual `runTextANN` method;
it is a result-semantics regression, not a planner-use assertion.

## Full #173 planner gate (still unmet)

Use PostgreSQL 16+ with pgvector, all migrations applied, and a populated
synthetic corpus in a disposable database whose name ends in `_test`:

```bash
bash scripts/verify_hnsw_indexes.sh "$MEM_TEST_DB" "$CORPUS_USER_UUID"
```

The read-only script validates the exact indexes and non-null vector counts
for the supplied corpus owner, and runs EXPLAIN ANALYZE for the shipping text
and visual ordering/deduplication shapes. It checks each route's index name,
propagates SQL errors, and exits nonzero if either route misses HNSW.
It does not disable sequential scans or claim a production latency threshold.

## Remaining acceptance boundary

The shipping text query already uses `DISTINCT ON (f.id)` ordered by file ID
before global top-k selection. A simplified `ORDER BY distance LIMIT` query
does not prove that this production query uses HNSW. This query shape is
unchanged from main: a failed text planner gate is an unmet optimization
criterion, not a regression introduced by the index DDL. The full-issue verification
must remain failed until that criterion is met. The partial index scope above
does not weaken that check. This is still an unmet feature acceptance criterion even
though the query predates this PR; an unchanged baseline is not a waiver.

### Historical bounded rewrite investigation (2026-09-10)

The smallest tested direct-distance rewrite filtered each chunk with a
`NOT EXISTS` peer having a smaller distance (UUID tie-break), then ordered by
`e.embedding <=> query_vector LIMIT 10`. PostgreSQL 17.10 / pgvector 0.8.3
selected `idx_embeddings_text_embedding_hnsw`, with the existing file-ID
index serving the peer lookup, on the 2,000-file synthetic fixture. Planner
selection alone did not establish equivalent results.

A rollback-only 768-dimensional counterexample separated one file's 101 near
chunks from the other 1,999 files: the first near vector was `[1,0,0,...]`,
the next 100 were `[1,i*0.001,0,...]`, and other files used `[0,1,0,...]`.
At unmodified defaults (`hnsw.ef_search=40`, `hnsw.iterative_scan=off`), the
shipping exact per-file query returned 10 files; the indexed anti-join returned
only 1. EXPLAIN ANALYZE showed the HNSW scan returning 40 candidate chunks,
39 then eliminated by per-file deduplication. All fixture mutations rolled
back. This is synthetic semantic/planner evidence, not latency evidence.

A separate exact-distance counterexample rules out a fixed oversampling cap:
81 closest chunks belonging to one file consume `8*k=80` candidates for
`k=10`, leaving one file after deduplication when ten files exist. Increasing
a fixed multiplier cannot guarantee k distinct files for unbounded chunk counts.

Concrete design blocker: the shipping contract chooses the best chunk per
eligible file before global top-k. A bounded approximate candidate scan can
underfill after deduplication or filtering. A safe rewrite needs a tested
candidate-exhaustion/continuation and fallback policy, including per-file
best-chunk selection and existing authorization/path/MIME/time filters.
Enabling iterative scans alone still needs an explicit scan-limit/exhaustion
policy; it is not proof of equivalence. That policy is not implemented or
accepted here. The promising anti-join is therefore not shipped, the original
query remains unchanged, and the text planner check must continue failing.
No planner settings or acceptance criteria were weakened.

Face indexing has no shipping SQL search route; its evidence is valid DDL and
populated-table migration, not a measured face-query speedup.

Recall, real embedding quality, production latency, index build time and index
size are NOT VERIFIED. #175 / #184 track live retrieval benchmark evidence;
fixture tests are not live quality results. No numerical improvement or recall
percentage is asserted here.

Fixed `vector(768)` / `vector(512)` column types reject wrong dimensions on
insert, before HNSW construction. NULL vectors can remain in the tables but
are not indexed. Down removes only the three new indexes and retains data.

## Operational limits

Migration 0025 remains transactional, consistent with startup migrations.
Regular `CREATE INDEX` blocks writes on the indexed tables until commit; plan
an ingestion maintenance window for a populated deployment. The rollback test
proves data preservation, not uninterrupted concurrent ingest or zero downtime.
`CREATE INDEX CONCURRENTLY` would need a separate nontransactional recovery
policy, including invalid/partially built indexes, and is not silently adopted.

The indexes retain pgvector defaults `m=16`, `ef_construction=64`. Follow
[pgvector's index/scan documentation](https://github.com/pgvector/pgvector#hnsw)
and measure representative recall, build time and ingest cost before tuning;
there is no asserted safe row-count threshold. Iterative scans still have
scan/memory limits and do not establish exact file-level result equivalence.
NULL and zero vectors are not indexed for cosine distance.

The strict verification script stops immediately on SQL/connection errors.
Its pass/fail counters summarize completed assertions, not infrastructure
errors; an interrupted run is never a pass. No `|| true` masks SQL failures.

The former MinIO lifecycle pull failure is fixed by the inherited main commit
`f1cc9eb` (#209), which uses pinned `quay.io/minio` images. No extra image or
workflow override belongs to this HNSW diff.
15 changes: 13 additions & 2 deletions scripts/verify.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ set -euo pipefail

REPO_ROOT="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
MODE="${1:-unit}"
EXPECTED_MIGRATION_HEAD=24
EXPECTED_MIGRATION_HEAD=25
MIGRATION_ROLLBACK_TARGET=11
MODEL_TEXT_CANONICAL_BASE=15
WORKSPACE_AI_PROFILE_BASE=16
Expand Down Expand Up @@ -349,6 +349,7 @@ run_postgres_tests() {
TestManagedAISettlementOutboxPostgres
TestReleasedFileStageRetryPostgres
TestDurableContextPostgres
TestTextANNFileSemanticsPostgres
)

integration_log="$(mktemp "${TMPDIR:-/tmp}/mem-integration.XXXXXX")"
Expand All @@ -359,7 +360,7 @@ run_postgres_tests() {
MEM_TEST_DB="$MEM_TEST_DB" go test \
${race_flag:+"$race_flag"} \
-v -count=1 -p 1 -timeout 20m \
-run '^(TestMemoryPostgres|TestHandoffPostgres|TestWorkspaceTransferPostgres|TestWorkspaceTransferMergeConservativePostgres|TestHandoffCrossAgentHTTPIntegration|TestRelocateHTTPPostgres|TestMemoryPathLifecycleIntegration|TestWorkspacePathLockingIntegration|TestFilePathLockingIntegration|TestAnnotationDecisionIntegration|TestIndexerEnrichmentIntegration|TestRecomputePerson|TestManagedEmbeddingEntitlementPostgres|TestManagedSearchReplayPostgres|TestManagedEmbeddingHTTPAuthorizationPostgres|TestAIProfilePostgres|TestIndexGenerationPostgres|TestManagedAISettlementOutboxPostgres|TestReleasedFileStageRetryPostgres|TestDurableContextPostgres)$' \
-run '^(TestMemoryPostgres|TestHandoffPostgres|TestWorkspaceTransferPostgres|TestWorkspaceTransferMergeConservativePostgres|TestHandoffCrossAgentHTTPIntegration|TestRelocateHTTPPostgres|TestMemoryPathLifecycleIntegration|TestWorkspacePathLockingIntegration|TestFilePathLockingIntegration|TestAnnotationDecisionIntegration|TestIndexerEnrichmentIntegration|TestRecomputePerson|TestManagedEmbeddingEntitlementPostgres|TestManagedSearchReplayPostgres|TestManagedEmbeddingHTTPAuthorizationPostgres|TestAIProfilePostgres|TestIndexGenerationPostgres|TestManagedAISettlementOutboxPostgres|TestReleasedFileStageRetryPostgres|TestDurableContextPostgres|TestTextANNFileSemanticsPostgres)$' \
./internal/memory \
./internal/handoff \
./internal/workspacetransfer \
Expand Down Expand Up @@ -397,9 +398,19 @@ run_integration() {
validate_test_database
with_fresh_test_database migration_sequence run_migration_sequence
with_fresh_test_database migration run_migration_round_trip
with_fresh_test_database hnsw_migration run_hnsw_migration
with_fresh_test_database integration run_postgres_integration
}

run_hnsw_migration() {
log "Populated HNSW migration, rollback and ingest regression"
(
cd "${REPO_ROOT}/server"
MEM_HNSW_TEST_DB="$MEM_TEST_DB" go test -v -count=1 \
-run '^TestHNSWMigrationPostgres$' ./internal/db
)
}

run_integration_race() {
validate_test_database
with_fresh_test_database integration_race run_postgres_integration_race
Expand Down
65 changes: 65 additions & 0 deletions scripts/verify_hnsw_indexes.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
#!/usr/bin/env bash
# Read-only planner verification for shipping text and visual query shapes.
# Requires a populated disposable database; does not force planner settings.
set -euo pipefail
# pass/fail counts are planner assertions only. SQL/connection errors abort
# immediately (rather than manufacture an assertion result or hide an error).
trap 'echo "ERROR: HNSW verification aborted on an execution error; assertions are incomplete" >&2' ERR
DB_URL="${1:?Usage: $0 <database-url> <corpus-user-uuid>}"
CORPUS_USER="${2:?Supply the user UUID that owns the populated corpus}"
[[ "$CORPUS_USER" =~ ^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$ ]] || {
echo 'ERROR: corpus user must be a UUID' >&2; exit 1;
}
# CORPUS_USER is interpolated below only after the strict UUID validation.
sql() { psql -X -A -t -v ON_ERROR_STOP=1 "$DB_URL" -c "$1"; }
db_name="$(sql 'SELECT current_database()')"
[[ "$db_name" == *_test ]] || { echo 'ERROR: database must end in _test' >&2; exit 1; }
pass=0
fail=0
for kind in text visual face; do
index="idx_embeddings_${kind}_embedding_hnsw"
valid="$(sql "SELECT count(*) FROM pg_index i
JOIN pg_class c ON c.oid = i.indexrelid JOIN pg_am a ON a.oid = c.relam
WHERE i.indrelid = 'embeddings_${kind}'::regclass
AND c.relname = '${index}' AND a.amname = 'hnsw' AND i.indisvalid
AND pg_get_indexdef(i.indexrelid) LIKE '%vector_cosine_ops%'")"
rows="$(sql "SELECT count(*) FROM embeddings_${kind} e JOIN files f ON f.id=e.file_id
WHERE f.user_id='${CORPUS_USER}'::uuid AND e.embedding IS NOT NULL")"
if [[ "$valid" == 1 && "$rows" -gt 0 ]]; then
echo "PASS: ${index} is valid; corpus contains ${rows} non-null vectors"
pass=$((pass + 1))
else
echo "FAIL: ${index}: valid=${valid}, corpus vectors=${rows}"
fail=$((fail + 1))
fi
done
assert_plan() {
local route="$1" query="$2" plan
# ANALYZE executes the read so vector/schema errors cannot hide behind EXPLAIN.
plan="$(sql "EXPLAIN (ANALYZE, BUFFERS) ${query}")"
echo "$plan"
if grep -q "Index Scan using idx_embeddings_${route}_embedding_hnsw" <<<"$plan"; then
echo "PASS: shipping ${route} query uses HNSW"
pass=$((pass + 1))
else
echo "FAIL: shipping ${route} query does not use HNSW"
fail=$((fail + 1))
fi
}
# Match runTextANN: per-file DISTINCT ON precedes global top-k. A simple
# ORDER BY distance LIMIT probe would not establish this query's index usage.
assert_plan text "SELECT evidence_id, file_id, score FROM (
SELECT DISTINCT ON (f.id) e.id::text AS evidence_id, f.id AS file_id,
1 - (e.embedding <=> array_fill(0.1::real, ARRAY[768])::vector) AS score
FROM embeddings_text e JOIN files f ON f.id=e.file_id
WHERE f.user_id='${CORPUS_USER}'::uuid
ORDER BY f.id, e.embedding <=> array_fill(0.1::real, ARRAY[768])::vector ASC
) hits ORDER BY score DESC LIMIT 10"
# Visual vectors have file_id (not e.id) and 512 dimensions.
assert_plan visual "SELECT e.file_id,
(1 - (e.embedding <=> array_fill(0.1::real, ARRAY[512])::vector))::real AS score
FROM embeddings_visual e JOIN files f ON f.id=e.file_id
WHERE f.user_id='${CORPUS_USER}'::uuid
ORDER BY e.embedding <=> array_fill(0.1::real, ARRAY[512])::vector ASC LIMIT 10"
echo "Results: ${pass} passed, ${fail} failed"
[[ "$fail" -eq 0 ]]
Loading
Loading