You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Text search over files has no model-free path. If the embedding worker is not
reachable, file search does not degrade — it refuses. Structured memory has a
deliberate lexical lane; the file corpus has none.
This is a design question first, not a defect: the model-free guarantee in the
repository's own decisions is consistently written about memories, and nobody on
this repository has ever applied it to files as far as I can find.
Source reproduction
On main @ 7a194f1eba4167d54bd46cf84cdbe86e00532319:
server/internal/search/search.go:220-221 — when the worker is nil or
disabled, search returns search disabled: worker not configured. There is no
branch that continues without an embedding.
SPEC.md:536 grants the deterministic no-model lane to memories
("memories FTS + trigram — 无模型的确定性立即召回"). files is not
mentioned.
The lexical index exists only on memories: search_tsv and the trigram GIN
are declared at server/internal/db/migrations/0008_agent_memories.sql:79-99
and :100-101; files (0001_init.sql:35-61) has no tsvector and no trigram
index, and server/internal/search/search.go contains no full-text or
trigram branch at all — its only retrieval routes are vector cosine, plus
mime/path/time filters.
CLI surface: the flag block is server/cmd/mem/cmds_search.go:103-114 and it
exposes --type --route --since --until --limit --format --idempotency-key. --route only selects text|visual|auto, all three of which are vector
routes — there is no text-only route to select. Separately, path scoping is not a service gap: search.Service already takes PathPrefix (search.go:190) and applies it via appendPathFilters at :321, :361, :641, :687. The CLI simply exposes no flag for it, so that
piece is CLI-only and smaller than the rest of this issue.
Precedent for the index shape already exists on the memories side: 0002_folders.sql:28 gives files(user_id, path text_pattern_ops), and 0008_agent_memories.sql:89/100-101 pair the same pattern index with a
trigram GIN. The file corpus has the first and not the second.
The stated product frame is an AI drive where a user must be able to find things
(GOAL.md §6, "找得到"). Two properties that are currently true by accident,
and that a decision would fix either way:
A self-hosted deployment that never configures a worker cannot search its own
file names' contents, but can browse by path. That asymmetry is not
documented as intentional anywhere I found.
Filename and path substring search — the thing that makes a drive usable with
no model at all — does not need a new idea; 0008 already shows the shape
for memories, and files.path already carries a text_pattern_ops index
used for subtree candidates.
Scope boundary
Explicitly not asked here: a Chinese tokenizer. I checked and the to_tsvector('simple', …) choice at 0008:79-80 is documented as intentional,
with the comment at :77-78 stating that pg_trgm below is the CJK and
typo-tolerant lane. Whether that trigram lane is good enough for Chinese is an
open empirical question, and it belongs to #175, which exists because
nothing currently measures it. This issue must not be merged into a tokenizer change.
Also out of scope: search result ranking, and anything that requires the
generation lifecycle.
Acceptance
A maintainer decides, recorded in this issue, whether a model-free
retrieval lane over file content and paths is intended for the file corpus
or explicitly not.
If intended: mem_search and the CLI can complete a query with the worker
stopped, the degradation is stated in SPEC.md §6.2 and in docs/mcp.md,
and a test proves the no-worker path returns results rather than an error.
If not intended: the current error becomes a documented, named limitation
in the deployment and MCP docs, and this issue closes with that reason.
Whatever is decided, the no-worker behavior of file search differs from
memory search in exactly the way the decision describes, and one test
covers each side of that difference.
Evidence level
E2 — source-level, including the negative claims (no lexical branch in search.go, no tsvector/trigram on files) which were checked by reading the
whole of search/search.go and every migration. No behavior was observed at
runtime.
Proposed triage
Applied on filing, per docs/maintainers/triage.md and the precedent in #135: type:rfc, area:server, area:recall, evidence:e2-source, status:needs-triage.
Left to a maintainer: priority, and any change to the set below.
type:rfc ("requires a reviewed product or architecture decision") is chosen over type:feature deliberately: the first acceptance item is a maintainer decision, and
the honest answer may be "not intended", in which case this closes with a documented
limitation instead of a change. triage.md says to change the type when the
classification changes, so flipping it to type:feature after a "yes" decision is
the expected path.
Summary
Text search over files has no model-free path. If the embedding worker is not
reachable, file search does not degrade — it refuses. Structured memory has a
deliberate lexical lane; the file corpus has none.
This is a design question first, not a defect: the model-free guarantee in the
repository's own decisions is consistently written about memories, and nobody on
this repository has ever applied it to files as far as I can find.
Source reproduction
On
main @ 7a194f1eba4167d54bd46cf84cdbe86e00532319:server/internal/search/search.go:220-221— when the worker is nil ordisabled, search returns
search disabled: worker not configured. There is nobranch that continues without an embedding.
SPEC.md:536grants the deterministic no-model lane tomemories("
memoriesFTS + trigram — 无模型的确定性立即召回").filesis notmentioned.
search_tsvand the trigram GINare declared at
server/internal/db/migrations/0008_agent_memories.sql:79-99and
:100-101;files(0001_init.sql:35-61) has no tsvector and no trigramindex, and
server/internal/search/search.gocontains no full-text ortrigram branch at all — its only retrieval routes are vector cosine, plus
mime/path/time filters.
server/cmd/mem/cmds_search.go:103-114and itexposes
--type --route --since --until --limit --format --idempotency-key.--routeonly selectstext|visual|auto, all three of which are vectorroutes — there is no text-only route to select. Separately, path scoping is
not a service gap:
search.Servicealready takesPathPrefix(search.go:190) and applies it viaappendPathFiltersat:321,:361,:641,:687. The CLI simply exposes no flag for it, so thatpiece is CLI-only and smaller than the rest of this issue.
0002_folders.sql:28givesfiles(user_id, path text_pattern_ops), and0008_agent_memories.sql:89/100-101pair the same pattern index with atrigram GIN. The file corpus has the first and not the second.
path"), feat(models): add a hardware-aware local embedding catalog and guided installer #47 ("keeps lexical/structured memory usable when the user skips this
step"), feat(platform): gate managed embeddings with workspace entitlements and idempotent usage #45 ("preserves model-free structured/lexical memory where
available"). feat(recall): orchestrate hybrid candidates, reranking, diversity, and budgets #71 was closed
not_planned.Why this is worth deciding rather than coding
The stated product frame is an AI drive where a user must be able to find things
(
GOAL.md§6, "找得到"). Two properties that are currently true by accident,and that a decision would fix either way:
file names' contents, but can browse by path. That asymmetry is not
documented as intentional anywhere I found.
no model at all — does not need a new idea;
0008already shows the shapefor
memories, andfiles.pathalready carries atext_pattern_opsindexused for subtree candidates.
Scope boundary
Explicitly not asked here: a Chinese tokenizer. I checked and the
to_tsvector('simple', …)choice at0008:79-80is documented as intentional,with the comment at
:77-78stating thatpg_trgmbelow is the CJK andtypo-tolerant lane. Whether that trigram lane is good enough for Chinese is an
open empirical question, and it belongs to #175, which exists because
nothing currently measures it. This issue must not be merged into a tokenizer change.
Also out of scope: search result ranking, and anything that requires the
generation lifecycle.
Acceptance
retrieval lane over file content and paths is intended for the file corpus
or explicitly not.
mem_searchand the CLI can complete a query with the workerstopped, the degradation is stated in
SPEC.md§6.2 and indocs/mcp.md,and a test proves the no-worker path returns results rather than an error.
in the deployment and MCP docs, and this issue closes with that reason.
memory search in exactly the way the decision describes, and one test
covers each side of that difference.
Evidence level
E2 — source-level, including the negative claims (no lexical branch in
search.go, no tsvector/trigram onfiles) which were checked by reading thewhole of
search/search.goand every migration. No behavior was observed atruntime.
Proposed triage
Applied on filing, per
docs/maintainers/triage.mdand the precedent in #135:type:rfc,area:server,area:recall,evidence:e2-source,status:needs-triage.Left to a maintainer: priority, and any change to the set below.
type:rfc("requires a reviewed product or architecture decision") is chosen overtype:featuredeliberately: the first acceptance item is a maintainer decision, andthe honest answer may be "not intended", in which case this closes with a documented
limitation instead of a change.
triage.mdsays to change the type when theclassification changes, so flipping it to
type:featureafter a "yes" decision isthe expected path.
Raised from the data/index audit on 2026-09-08.