Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,6 +165,8 @@ The project publishes 0.x prerelease versions; a stable release line is not yet

### Fixed

- Enforce text-chunk uniqueness, index nullable memory references, cascade memory relations with workspace deletion, batch generation listings, and paginate relation listings with opaque cursors. Refs #178.

- Follow the shared design language for reading, numeric and action alignment; generate the existing Web color variables from a pinned design-system token snapshot, and use a single consistent empty-state pattern. Refs #211.

- Improve web caption and status contrast in both themes, including tinted danger
Expand Down
30 changes: 16 additions & 14 deletions docs/MIGRATION_SEQUENCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,30 +6,31 @@ These draft changes are cumulative, not independently deployable:
| --- | --- | --- | --- |
| 1 | #194 / #183 | 0024 file lexical lane | released/main schema 23 (merged) |
| 2 | #173 HNSW completion (supersedes #197 HOLD) | 0025 HNSW indexes + text continuation | #194, schema 24 |
| 3 | #195 / #185 | 0026 data-plane hygiene | schema 25 |
| 3 | #195 / #185 | 0027 data-plane hygiene | schema 26 |

The PR base chain is `main` → `codex/fix-pr-183` → `codex/fix-pr-180`
→ `codex/fix-pr-185`. Successor branches must include their predecessor schema and source. Local
repair branches are rebuilt on current main and replay the original authored
changes; published commit identities remain in the original PR history.
Keep this order when retargeting after a predecessor merges.
Migrations 0024–0026 are now on `main`. PR #195 must integrate that exact
history before adding data-plane hygiene as 0027. Local repair branches retain
the original authored commits and add a current-main integration commit rather
than rewriting published history.

On 2026-09-10, main `2986fe38175f54d99f15dd38a498708c6ecd88cd` and published
tags `v0.1.0` / `v0.1.1` contain only migrations 0001–0023. This does not prove
that a private deployment never applied a draft. Consequently migration
numbers and SQL identities are retained, not renumbered on an assumption.
The hygiene migration was originally reviewed as draft 0026, but main now owns
0026 for `memory_producer_agent`. The unmerged hygiene migration therefore
moves to 0027. An operator who privately applied the old draft under version 26
must stop and obtain a recovery plan; do not rewrite Goose history to make the
new main sequence appear valid.

Migration 0025 creates the three cosine HNSW indexes. The shipping text route
no longer uses `DISTINCT ON (f.id) ORDER BY f.id` as its primary plan: it walks
cosine-ordered candidates and falls back to that exact query only when a bounded
scan underfills. Visual cosine-order already matched HNSW. Face DDL is not a
face-query speedup. #195 still requires predecessor schema 25. Recall and live
latency remain `#175`, not this migration.
latency remain `#175`, not this migration. #195 now requires predecessor schema
26.

This document does not waive a review gate or authorize deployment.

Goose startup remains strict: no `WithAllowMissing` or equivalent option is
enabled. A database that already applied 26 while omitting 24/25 will correctly
enabled. A database that already applied 27 while omitting 24/25/26 will correctly
fail startup against the cumulative schema. Stop and obtain an operator-owned
recovery plan for such a database; do not edit its migration history, renumber
its SQL, or apply lower versions out of order to manufacture a pass.
Expand All @@ -55,9 +56,10 @@ an operator-owned recovery; do not mark an unverified partial schema applied.
`scripts/verify.sh integration` creates a separate, owned `_test` database and
runs `TestMigrationUpgradeSequence`. It applies real Goose migrations to 23,
seeds a file with duplicate text chunks, then advances one version at a time
to the branch's declared head (24, 25, or 26). Each step checks full applied
to the branch's declared head (24 through 27). Each step checks full applied
history and preserved data; subsequent steps check lexical backfill, valid
HNSW DDL, and deduplication/unique rejection. Finally the ordinary production
HNSW DDL, producer-agent indexing, and hygiene deduplication/unique rejection.
Finally the ordinary production
`DB.Migrate` startup path must accept the resulting history unchanged.

The dedicated test uses `MEM_MIGRATION_SEQUENCE_TEST_DB`, refuses a database
Expand Down
2 changes: 1 addition & 1 deletion scripts/verify.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ set -euo pipefail

REPO_ROOT="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
MODE="${1:-unit}"
EXPECTED_MIGRATION_HEAD=26
EXPECTED_MIGRATION_HEAD=27
MIGRATION_ROLLBACK_TARGET=11
MODEL_TEXT_CANONICAL_BASE=15
WORKSPACE_AI_PROFILE_BASE=16
Expand Down
2 changes: 1 addition & 1 deletion server/internal/api/api.go
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,7 @@ type MemoryService interface {
Restore(context.Context, memory.LifecycleCommand) (*memory.MutationResult, error)
Forget(context.Context, memory.ForgetCommand) (*memory.ForgetResult, error)
CreateRelation(context.Context, memory.CreateRelationCommand) (*memory.CreateRelationResult, error)
ListRelations(context.Context, memory.ListRelationsQuery) ([]memory.Relation, error)
ListRelations(context.Context, memory.ListRelationsQuery) (*memory.ListRelationsResult, error)
}

// DurableContextService is the scoped durable-context port (mem#70). Handlers
Expand Down
15 changes: 11 additions & 4 deletions server/internal/api/handlers_memory.go
Original file line number Diff line number Diff line change
Expand Up @@ -763,18 +763,21 @@ func (s *Server) handleListMemoryRelations(w http.ResponseWriter, r *http.Reques
}

tok := r.Context().Value(ctxToken).(*auth.Token)
relations, err := s.Memory.ListRelations(r.Context(), memory.ListRelationsQuery{
result, err := s.Memory.ListRelations(r.Context(), memory.ListRelationsQuery{
WorkspaceID: currentWorkspace(r).ID,
MemoryID: id,
Direction: direction,
RelationType: relationType,
AllowedPaths: tok.Paths,
Limit: limit,
Cursor: r.URL.Query().Get("cursor"),
})
if err != nil {
switch {
case errors.Is(err, memory.ErrInvalidCommand):
writeError(w, http.StatusBadRequest, "invalid_relation_query", err.Error())
case errors.Is(err, memory.ErrInvalidCursor):
writeError(w, http.StatusBadRequest, "invalid_cursor", err.Error())
case errors.Is(err, memory.ErrNotFound):
writeError(w, http.StatusNotFound, "not_found", "memory not found")
case errors.Is(err, memory.ErrForgotten):
Expand All @@ -789,8 +792,12 @@ func (s *Server) handleListMemoryRelations(w http.ResponseWriter, r *http.Reques
}
return
}
if relations == nil {
relations = []memory.Relation{}
resp := map[string]any{"relations": result.Relations}
if result.Relations == nil {
resp["relations"] = []memory.Relation{}
}
writeJSON(w, http.StatusOK, map[string]any{"relations": relations})
if result.NextCursor != "" {
resp["next_cursor"] = result.NextCursor
}
writeJSON(w, http.StatusOK, resp)
}
4 changes: 2 additions & 2 deletions server/internal/api/handlers_memory_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -107,10 +107,10 @@ func (s *memoryServiceStub) CreateRelation(
func (s *memoryServiceStub) ListRelations(
_ context.Context,
q memory.ListRelationsQuery,
) ([]memory.Relation, error) {
) (*memory.ListRelationsResult, error) {
s.calls++
s.listRelationsQuery = q
return s.relations, s.controlErr
return &memory.ListRelationsResult{Relations: s.relations}, s.controlErr
}

func memoryHandlerContext(req *http.Request, paths []string) (*http.Request, uuid.UUID, uuid.UUID, uuid.UUID) {
Expand Down
17 changes: 15 additions & 2 deletions server/internal/db/migration_sequence_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ package db
import (
"context"
"database/sql"
"errors"
"os"
"strconv"
"strings"
Expand All @@ -11,6 +12,7 @@ import (

"github.com/google/uuid"
"github.com/jackc/pgx/v5"
"github.com/jackc/pgx/v5/pgconn"
"github.com/pressly/goose/v3"
)

Expand Down Expand Up @@ -107,8 +109,12 @@ func TestMigrationUpgradeSequence(t *testing.T) {
}
}
var chunks int
if err := sqldb.QueryRowContext(ctx, "SELECT count(*) FROM embeddings_text WHERE file_id=$1", fileID).Scan(&chunks); err != nil || chunks != 2 {
t.Fatalf("preserved chunks=%d, want=2, err=%v", chunks, err)
wantChunks := 2
if version >= 27 {
wantChunks = 1
}
if err := sqldb.QueryRowContext(ctx, "SELECT count(*) FROM embeddings_text WHERE file_id=$1", fileID).Scan(&chunks); err != nil || chunks != wantChunks {
t.Fatalf("preserved chunks=%d, want=%d, err=%v", chunks, wantChunks, err)
}
if version >= 26 {
var producerIdx int
Expand All @@ -118,6 +124,13 @@ func TestMigrationUpgradeSequence(t *testing.T) {
}
t.Logf("PASS: strict Goose upgrade to %d; complete history and populated data preserved", version)
}
if head >= 27 {
_, err := sqldb.ExecContext(ctx, "INSERT INTO embeddings_text(file_id,chunk_index,chunk_text) VALUES($1,0,'duplicate')", fileID)
var pgErr *pgconn.PgError
if !errors.As(err, &pgErr) || pgErr.Code != "23505" {
t.Fatalf("expected duplicate rejection 23505, got %v", err)
}
}
// The real startup path must accept the upgraded history unchanged.
if err := (&DB{url: dsn}).Migrate(ctx); err != nil {
t.Fatalf("production startup after sequential upgrade: %v", err)
Expand Down
85 changes: 85 additions & 0 deletions server/internal/db/migrations/0027_data_plane_hygiene.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
-- +goose Up
-- Data-plane hygiene from the index audit (#178), sequenced after main's 0026.
-- Three independent fixes bundled into one migration because each is a single
-- DDL statement and none warrants its own schema version.

-- Item 1: embeddings_text uniqueness on (file_id, chunk_index).
-- The write path (indexer.go) DELETEs all chunks for a file before re-inserting,
-- so duplicates should never exist in practice. Deduplicate defensively before
-- adding the constraint: if any duplicates survived, keep the row with the
-- smallest UUID for deterministic selection (UUID order is not insert order).
-- +goose StatementBegin
DELETE FROM embeddings_text
WHERE id NOT IN (
SELECT DISTINCT ON (file_id, chunk_index) id
FROM embeddings_text
ORDER BY file_id, chunk_index, id
);
-- +goose StatementEnd

-- +goose StatementBegin
ALTER TABLE embeddings_text
ADD CONSTRAINT uq_embeddings_text_file_chunk UNIQUE (file_id, chunk_index);
-- +goose StatementEnd

-- Item 2: partial indexes for ON DELETE SET NULL lookups on memories.
-- Without these, every file or user delete takes a RowExclusiveLock on memories
-- and performs a sequential scan to find the rows to null out.
-- +goose StatementBegin
CREATE INDEX IF NOT EXISTS idx_memories_source_file_id
ON memories (source_file_id) WHERE source_file_id IS NOT NULL;
-- +goose StatementEnd

-- +goose StatementBegin
CREATE INDEX IF NOT EXISTS idx_memories_created_by_user_id
ON memories (created_by_user_id) WHERE created_by_user_id IS NOT NULL;
-- +goose StatementEnd

-- Item 3: memory_relations FKs must cascade with memories.
-- memories.workspace_id is ON DELETE CASCADE, so a workspace delete removes
-- memories rows. Without matching cascade on memory_relations, the delete then
-- fails on any edge touching those memories. Align the referential actions.
-- +goose StatementBegin
ALTER TABLE memory_relations
DROP CONSTRAINT IF EXISTS memory_relations_workspace_id_fkey,
DROP CONSTRAINT IF EXISTS memory_relations_source_id_fkey,
DROP CONSTRAINT IF EXISTS memory_relations_target_id_fkey;
-- +goose StatementEnd

-- +goose StatementBegin
ALTER TABLE memory_relations
ADD CONSTRAINT memory_relations_workspace_id_fkey
FOREIGN KEY (workspace_id) REFERENCES workspaces(id) ON DELETE CASCADE,
ADD CONSTRAINT memory_relations_source_id_fkey
FOREIGN KEY (source_id) REFERENCES memories(id) ON DELETE CASCADE,
ADD CONSTRAINT memory_relations_target_id_fkey
FOREIGN KEY (target_id) REFERENCES memories(id) ON DELETE CASCADE;
-- +goose StatementEnd

-- +goose Down
-- +goose StatementBegin
ALTER TABLE memory_relations
DROP CONSTRAINT IF EXISTS memory_relations_workspace_id_fkey,
DROP CONSTRAINT IF EXISTS memory_relations_source_id_fkey,
DROP CONSTRAINT IF EXISTS memory_relations_target_id_fkey;
-- +goose StatementEnd

-- +goose StatementBegin
ALTER TABLE memory_relations
ADD CONSTRAINT memory_relations_workspace_id_fkey
FOREIGN KEY (workspace_id) REFERENCES workspaces(id),
ADD CONSTRAINT memory_relations_source_id_fkey
FOREIGN KEY (source_id) REFERENCES memories(id),
ADD CONSTRAINT memory_relations_target_id_fkey
FOREIGN KEY (target_id) REFERENCES memories(id);
-- +goose StatementEnd

-- +goose StatementBegin
DROP INDEX IF EXISTS idx_memories_created_by_user_id;
DROP INDEX IF EXISTS idx_memories_source_file_id;
-- +goose StatementEnd

-- +goose StatementBegin
ALTER TABLE embeddings_text
DROP CONSTRAINT IF EXISTS uq_embeddings_text_file_chunk;
-- +goose StatementEnd
109 changes: 109 additions & 0 deletions server/internal/db/migrations_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
package db

import (
"context"
"errors"
"os"
"strings"
"testing"
"time"

"github.com/google/uuid"
"github.com/jackc/pgx/v5/pgconn"
"github.com/jackc/pgx/v5/pgxpool"
)

func TestEmbeddingsTextUniqueConstraint(t *testing.T) {
dsn := os.Getenv("MEM_TEST_DB")
if dsn == "" {
t.Skip("MEM_TEST_DB not set; skipping DB integration test")
}
config, err := pgxpool.ParseConfig(dsn)
if err != nil {
t.Fatalf("parse MEM_TEST_DB: %v", err)
}
if !strings.HasSuffix(config.ConnConfig.Database, "_test") {
t.Fatalf("refusing to modify non-test database %q", config.ConnConfig.Database)
}

ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()

pool, err := pgxpool.NewWithConfig(ctx, config)
if err != nil {
t.Fatalf("connect: %v", err)
}
defer pool.Close()

db := &DB{Pool: pool, url: dsn}
if err := db.Migrate(ctx); err != nil {
t.Fatalf("migrate: %v", err)
}
tx, err := pool.Begin(ctx)
if err != nil {
t.Fatal(err)
}
defer tx.Rollback(context.Background())

var userID, workspaceID, fileID uuid.UUID
if err := tx.QueryRow(ctx, `
INSERT INTO users (email, password_hash) VALUES ($1, 'x')
RETURNING id
`, uuid.NewString()+"@example.test").Scan(&userID); err != nil {
t.Fatalf("seed user: %v", err)
}
if err := tx.QueryRow(ctx, `
INSERT INTO workspaces (name, resource_owner_user_id)
VALUES ('unique-ctest-ws', $1)
ON CONFLICT (resource_owner_user_id) DO UPDATE SET name = EXCLUDED.name
RETURNING id
`, userID).Scan(&workspaceID); err != nil {
t.Fatalf("seed workspace: %v", err)
}
if err := tx.QueryRow(ctx, `
INSERT INTO files (user_id, name, path, size, sha256, mime, storage_key)
VALUES ($1, 'unique-ctest.txt', '/unique-ctest.txt', 0, '', 'text/plain', 'test://unique')
RETURNING id
`, userID).Scan(&fileID); err != nil {
t.Fatalf("seed file: %v", err)
}
if _, err := tx.Exec(ctx, `
INSERT INTO embeddings_text (file_id, chunk_index, chunk_text, provider)
VALUES ($1, 0, 'chunk zero', 'test')
`, fileID); err != nil {
t.Fatalf("first insert: %v", err)
}
// Replay the actual migration over populated, pre-constraint data inside
// this rollback-only transaction. A fresh-schema migrate alone misses this.
if _, err := tx.Exec(ctx, `ALTER TABLE embeddings_text DROP CONSTRAINT uq_embeddings_text_file_chunk`); err != nil {
t.Fatal(err)
}
if _, err := tx.Exec(ctx, `INSERT INTO embeddings_text (file_id, chunk_index, chunk_text)
VALUES ($1, 0, 'legacy duplicate')`, fileID); err != nil {
t.Fatal(err)
}
migration, err := migrationsFS.ReadFile("migrations/0027_data_plane_hygiene.sql")
if err != nil {
t.Fatal(err)
}
up := strings.SplitN(string(migration), "-- +goose Down", 2)[0]
if _, err := tx.Exec(ctx, up); err != nil {
t.Fatalf("migrate populated table: %v", err)
}
var survivors int
if err := tx.QueryRow(ctx, `SELECT count(*) FROM embeddings_text WHERE file_id=$1`, fileID).Scan(&survivors); err != nil || survivors != 1 {
t.Fatalf("deduplicated rows = %d, err=%v", survivors, err)
}

_, err = tx.Exec(ctx, `
INSERT INTO embeddings_text (file_id, chunk_index, chunk_text, provider)
VALUES ($1, 0, 'duplicate chunk', 'test')
`, fileID)
if err == nil {
t.Fatal("expected duplicate (file_id, chunk_index) to be rejected")
}
var pgErr *pgconn.PgError
if !errors.As(err, &pgErr) || pgErr.Code != "23505" {
t.Fatalf("expected unique violation 23505, got %v", err)
}
}
Loading
Loading