You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(search): scale the retirement scan window with the row limit and time out resumed rechecks like completion
A capped page re-reads its scan from its last mutated row, so a fixed 25,000-ID window re-read
most of the same IDs on every page once the row limit shrank. Each page now reads at most four IDs
per row of its limit, capped at 25,000, and any fast page doubles the limit so sparse stretches
widen the window again.
A retry that finds retirement already complete revalidates every captured KB, as completion does,
so it now runs under the same 30-minute timeout instead of the two-minute page timeout.
Copy file name to clipboardExpand all lines: packages/db/script-migrations/search-embedding-retirement.md
+8-6Lines changed: 8 additions & 6 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -30,18 +30,20 @@ resume the saved scope, phase, cursor and maintenance checkpoints.
30
30
The existing maintenance implementation rebuilds HNSW indexes and vacuums affected tables before
31
31
deployment continues.
32
32
33
-
Each page reads at most 25,000 IDs and mutates at most a row limit of them. Pages execute
33
+
Each page mutates at most a row limit of target rows and reads at most four IDs per row of that
34
+
limit, never more than 25,000 IDs. Pages execute
34
35
sequentially, and each is followed by a pause as long as the page took, up to five seconds, to
35
36
leave the primary headroom. Retiring a document is a non-HOT update that writes every index on
36
37
`document`, and deleting a chunk cascades into its projections, so a page's cost follows the target
37
38
rows it mutates, not the IDs it reads. A page that reaches the row limit advances the cursor only to
38
-
its last mutated row; the rest of its scan is read again by the next page. Documents that are
39
-
already retired never count against the limit.
39
+
its last mutated row; the rest of its scan is read again by the next page. Tying the scan window to
40
+
the limit keeps that re-reading proportional to the work, even after the limit shrinks. Documents
41
+
that are already retired never count against the limit.
40
42
41
43
The row limit starts at 2,000 rows. A page is timed from the start of its transaction through its
42
44
commit, including the synchronous-replication wait and any lock-timeout retries. A page slower than
43
-
30 seconds halves the limit. A fast page, one under 7.5 seconds, that reached the limit doubles it,
44
-
up to 8,000. The limit never drops below 25 rows. Phase changes do not adjust it.
45
+
30 seconds halves the limit. A fast page, one under 7.5 seconds, doubles it up to 8,000, which
46
+
also widens the scan window, so sparse stretches are not crawled in small windows. The limit never drops below 25 rows. Phase changes do not adjust it.
45
47
46
48
Materialized SQL pages keep the IDs inside PostgreSQL; the migration process receives only a cursor
47
49
and a validation result. Each page uses a two-minute statement timeout and a one-second lock
@@ -88,7 +90,7 @@ unrelated rows. Before completion, the cleanup checks for unretired documents an
88
90
behind either cursor and restarts the affected phase if needed. A final bounded pass validates all
89
91
captured KB markers, including empty KBs and KBs whose rows were already scanned, holding shared
90
92
marker locks until the completion checkpoint commits. Resuming a completed cleanup before maintenance
91
-
also revalidates the captured set. Keep target writers stopped and
93
+
also revalidates the captured set, with the same 30-minute timeout as the completion recheck. Keep target writers stopped and
92
94
do not change their Search markers during the pass.
0 commit comments