Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,8 @@ The project publishes 0.x prerelease versions; a stable release line is not yet
- The release guard suites count manifest lines without `wc -l`, whose BSD
implementation pads the count with blanks, and no longer need GNU
`find -printf`.
- Add an opt-in file-search ranking producer with explicit request failures, conservative result identity mapping, and operator-declared configuration. Live provider quality remains separately unverified.

- The npm installer no longer aborts a concurrent first run on Windows. The
per-asset cache lock previously treated only `EEXIST` as contention, but a
contended `mkdir` on Windows may raise `EPERM` or `EACCES`, so a process
Expand Down
95 changes: 95 additions & 0 deletions benchmarks/recall/README.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,14 @@
# Multilingual recall benchmark

> **Evidence boundary: producer checks do not complete the live benchmark.**
>
> The producer corrections tracked in #196 do not complete #175's acceptance.
> The original #184 live-run requirement still needs a populated real memd,
> actual embedding-provider output and saved ranking artifacts. Unit or
> loopback HTTP fixtures establish transport behavior only. A real model-free
> lexical run also cannot establish vector/provider quality. See
> [Exact remaining live prerequisites](#exact-remaining-live-prerequisites).

This directory provides a small, repeatable retrieval benchmark. It is a
decision aid for comparing lexical, vector and hybrid configurations; it is
not evidence of production recall.
Expand Down Expand Up @@ -167,3 +176,89 @@ time and sanitized query failures. It does not copy query text, corpus text,
vectors or free-form provider errors. Credential-shaped configuration keys
such as `api_key`, `password`, `secret`, `token` and `authorization` are
rejected instead of being copied into an artifact.

## Live memd producer

The `produce` subcommand queries a running `memd` over file-search queries and
emits a `mem.recall-rankings.v1` file that the existing `run --rankings` path
consumes. Latency is measured client-side per request; the `0 ms` sentinel
warning above applies only to the offline lexical lane.

```bash
python3 -m benchmarks.recall produce \
--memd-url http://localhost:8080 \
--token "$MEM_TOKEN" \
--dataset benchmarks/recall/data/profile-text-v1 \
--output /tmp/live-rankings.json \
--dimension 768 --provider "$MEM_PROVIDER_LABEL" --model "$MEM_MODEL_LABEL" \
--mode vector
```

Then score the saved rankings (the default v1 baseline uses a different corpus):

```bash
python3 -m benchmarks.recall run \
--dataset benchmarks/recall/data/profile-text-v1 \
--rankings /tmp/live-rankings.json \
--output /tmp/live-artifact.json
```

Load the synthetic file corpus into an isolated test deployment first. The
producer does not ingest it. Supply a token bound to the dataset's workspace;
its labels do not establish the token's real workspace identity.

The producer maps each API result back to a dataset `doc_id` using the returned
folder `path` plus file `name`. Cross-workspace path collisions, unknown paths,
or ambiguous snippets fail the query instead of silently dropping evidence.
A same-workspace snippet-overlap tie also fails closed: document ID ordering
would invent identity evidence. Malformed paths, names, snippets and scores
fail with `invalid_result`, including malformed duplicate hits.
Query filters are translated where the API supports them: `path_prefix` becomes
`scope`, and `source_kind` becomes `type` (`image_caption` → `image`,
`text` → `text`). The `workspace` filter is not sent to the API because the
auth token determines workspace scope. A single token cannot select several
workspaces, so use a single-workspace fixture. Metadata filters are unsupported
and produce `unsupported_filter` without contacting the server; they are never
silently ignored.

Vector mode sends `route=text`; it does not claim a hybrid lexical/vector or
multimodal experiment. Lexical mode sends `route=lexical` and requires the
server capability from #183. It emits null provider/model/dimension as required
by the ranking schema. Provider, model, dimension and index configuration are
operator declarations, not discovered or verified server metadata. The
`hardware.host` value deliberately records only the producer client's
OS/architecture, never its hostname. It is not the server's hardware inventory
and cannot establish comparable performance conditions.

The full v1 corpus also contains structured-memory queries. `/v1/search` cannot
serve these; the producer records `unsupported_source_kind` and exits 2. Any
HTTP, mapping or response error also exits 2 while retaining an error artifact.
Use the existing `profile-text-v1` file-only fixture for this producer's bounded
acceptance. An HTTP fixture test proves transport and artifact handling only;
it does not establish live provider quality, production latency, or full-corpus
acceptance. Those remain NOT VERIFIED until a real populated memd run is saved.


### Exact remaining live prerequisites

1. Use an isolated authorized memd/Worker deployment and disposable PostgreSQL
database, and verify that its token is bound to the intended test workspace.
2. Load all five synthetic files from `data/profile-text-v1/corpus.jsonl`, keeping
paths and content intact. Verify indexing and actual file identities before
interpreting the producer's output. Direct database seeding can establish
retrieval/transport behavior but does not verify ingestion or Worker indexing.
3. For vector acceptance, select the same actual embedding model for corpus and
queries, record its dimension and profile, and independently inspect the
active generation/index and deployment revision. Producer labels alone do
not verify any of those properties.
4. Run the documented producer command, retain its output, score the resulting
rankings and record errors and measured client latencies. An empty/error run
or a fake embedding provider cannot establish vector quality. A lexical run
requires the separate #194 server capability and remains a distinct result.
5. The bounded file-only experiment does not cover #175's bilingual image-query
acceptance or the full v1 structured-memory corpus. The original issue and
live quality acceptance must not be described as complete on this evidence.

The producer remains opt-in; the normal recall CI gate runs deterministic unit
and fixture checks only. No real-model baseline is checked in until its actual
configuration and saved run are available for review.
41 changes: 41 additions & 0 deletions benchmarks/recall/__main__.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,9 @@
import sys
import tempfile

from .dataset import load_dataset
from .errors import BenchmarkError
from .live_producer import produce_rankings
from .runner import (
compare_artifacts,
comparison_summary,
Expand Down Expand Up @@ -66,6 +68,24 @@ def _parser() -> argparse.ArgumentParser:
type=Path,
default=PACKAGE_ROOT / "fixtures" / "external-rankings.leak.v1.json",
)

produce = subparsers.add_parser(
"produce",
help="query a live memd and emit mem.recall-rankings.v1",
)
produce.add_argument("--memd-url", required=True, help="base URL of memd")
produce.add_argument("--token", required=True, help="bearer token for auth")
produce.add_argument("--dataset", type=Path, default=DEFAULT_DATASET)
produce.add_argument("--output", type=Path, required=True)
produce.add_argument("--limit", type=int, default=10)
produce.add_argument("--timeout", type=float, default=30.0)
produce.add_argument("--engine", default="live-memd")
produce.add_argument("--dimension", type=int, default=768)
produce.add_argument(
"--mode", default="vector", choices=["lexical", "vector"]
)
produce.add_argument("--provider", default="operator-unspecified")
produce.add_argument("--model", default="operator-unspecified")
return parser


Expand Down Expand Up @@ -100,6 +120,27 @@ def main(argv: list[str] | None = None) -> int:
print(comparison_summary(comparison))
return 2 if candidate["metrics"]["overall"]["leakage_count"] else 0

if args.command == "produce":
dataset = load_dataset(args.dataset)
rankings = produce_rankings(
dataset,
base_url=args.memd_url,
token=args.token,
limit=args.limit,
timeout=args.timeout,
engine_label=args.engine,
dimension=args.dimension,
mode=args.mode,
provider=args.provider,
model=args.model,
)
write_json(args.output, rankings)
ok_count = sum(1 for q in rankings["queries"] if q["status"] == "ok")
err_count = sum(1 for q in rankings["queries"] if q["status"] == "error")
print(f"produced rankings: {ok_count} ok, {err_count} error")
print(f"rankings artifact: {args.output}")
return 2 if err_count else 0

first = run_benchmark(
dataset_dir=args.dataset,
generated_at="2000-01-01T00:00:00+00:00",
Expand Down
Loading
Loading