Skip to content

v0.9.7: embeddings cleanup, search expansion, navigation and execution speedup - #8439

Merged
waleedlatif1 merged 21 commits into
mainfrom
staging
Sep 30, 2026
Merged

waleedlatif1 merged 21 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

waleedlatif1 and others added 10 commits September 29, 2026 12:07
…8421)

* improvement(desktop): surface update actions in a persistent toast

* fix(desktop): distinguish update toast dismissal from eviction
…oute reload/back to the browser tab (#8429)

* fix(desktop-browser): keep the chat caret while the agent works and route reload/back to the browser tab

* fix(desktop-browser): return popup focus to the user's page and wait for a shown view
* feat(cli): add sim cli search and default output to JSON

- sim cli search ranks every command in the tree against a plain-language
  query, locally via minisearch, and returns the top 5
- root and group --help open with a discovery note when a coding agent runs
  the CLI, pointing it at sim cli search
- default output format is now json; table stays available via --output,
  SIM_OUTPUT, or sim configure --set-output
- reword command descriptions that lacked the words people search with

* fix(cli): describe billing status by what it reports, scope search privacy claim to the query

* chore(cli): release 2.2.0
* fix(search): retire legacy embeddings through scoped backfill

* fix(search): run retirement without cleanup flags and close ingestion gaps

* chore(tests): scope migration recovery journal assertion

* chore(tests): align document dispatch billing and quota fixtures
…ts first block (#8434)

* improvement(execution): cut the time between a workflow trigger and its first block

- Trigger.dev getJob only retrieves real run ids; caller-chosen ids (schedule_…,
  workflow-execution:…) go straight to the tag lookup instead of a ~10 s 404
- Execution-log start resolves the snapshot id (cached once referenced) and inserts
  with ON CONFLICT instead of select + full-state upsert RETURNING * + insert
- Preprocessing starts payer attribution and the ban/usage/subscription gate reads as
  soon as their inputs are known; subscription is only read when needed
- Execution core prefetches env + PII policy alongside custom blocks and reuses the
  webhook job's already-loaded environment
- Webhook lookups answer trigger-block deployment from the join they already do;
  versioned deployment loads serve the materialized cache first
- getHighestPrioritySubscription, custom-block rows, and env suspension lookups drop
  sequential round trips

* fix(execution): scope the snapshot FK retry, keep a suspended identity's access read out of the run, tidy tests

* improvement(execution): encapsulate early admission reads, plain snapshot identity, explicit suspension paths

* improvement(environment): reuse the actor's resolved access when withholding a suspended identity
…switches (#8435)

* improvement(settings): remove the 300ms floor under settings section switches

Every section switch mounted a fresh Suspense boundary (the empty section loading.tsx and a page-level <Suspense fallback={null}> around the code-split body). React 19 holds content that resolves into a just-committed fallback for at least 300ms, so every switch paid that before the section even rendered or started its queries.

- drop the section loading boundaries and page-level Suspense on all settings planes; navigations are transitions, so the outgoing section stays until the incoming one is ready
- move the sidebar selection on click so the click still reads as acknowledged
- warm each hot section's first-content queries on navigation intent
- load the fork sync editor and custom tool editor on open; they pulled the block and trigger registries into the list chunks

* fix(settings): tie the pending sidebar selection to its navigation

The pending row cleared only on a pathname change, so a navigation the server redirected back to the current section, or one that failed, left the clicked row selected and its click guard swallowing retries. Set the selection optimistically inside the navigation's own transition so React drops it when that transition settles.
…ion (#8431)

* feat(logs): allow omitting workflow snapshots from diagnostic reads

* feat(stripe): add optional subscription pagination cursors

* fix: clarify Stripe guidance and log snapshot defaults
* fix(search): finish retirement and index maintenance on deploy

* docs(search): clarify retirement migration session requirement
* improvement(chat): show live search sources by default

* test(chat): cover live search disclosure transitions

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
@waleedlatif1
waleedlatif1 requested a review from a team as a code owner September 30, 2026 00:18
@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 30, 2026 7:51am UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Desktop browser navigation, shortcuts, and update notifications.

This PR is not safe to merge until Search retirement handles multiple owner-scoped indexes and refuses to remove data from deployments still using indexed Search.

Findings

  1. P1 Multiple Search indexes block deployment ▶
  2. P1 Retirement deletes active Search data ▶

Summary

This PR expands the CLI and integration tools, changes desktop focus and update notifications, speeds execution and settings navigation, adds an optional log projection, and retires legacy Search embeddings. The retirement needs correction before deployment because it assumes a single global Search index and does not protect deployments still using indexed Search.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Deploy[Deployment migration runner] --> Select[0027 selects Search KB]
  Select -->|More than one| Fail[Migration fails; deployment stops]
  Select -->|One| Retire[Disable documents and delete embeddings]
  Retire --> Maintain[0028 rebuilds indexes and vacuums]
  Retire -->|Indexed Search still enabled| Lost[Active indexed results need reindexing]
Loading

Reviews (1) · Last reviewed commit: "improvement(chat): show live search sour..."

Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts Outdated
Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 149 files

Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts Outdated
Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts Outdated
Comment thread apps/desktop/src/main/browser-agent/session.ts Outdated
Comment thread apps/sim/background/webhook-execution.ts
Comment thread apps/sim/lib/knowledge/documents/service.ts
Comment thread apps/sim/app/workspace/[workspaceId]/settings/[section]/layout.tsx Outdated
…er and route hard reload through recovery (#8442)

* fix(desktop-browser): keep error-page focus recovery inside the browser and route hard reload through recovery

* fix(desktop-browser): recover issue-page focus only when nothing holds focus, and keep cache bypass except for hung pages

* fix(desktop-browser): retry the failed URL when hard reloading a load error
* fix(search): retire indexed content across all Search KBs

* fix(search): validate retirement scope through completion
* feat(planetscale): add management actions

* fix(planetscale): support default-parent backups and input limits

* fix(planetscale): retain active backup branch context

* fix(planetscale): address integration validation findings

* chore(planetscale): clarify input and restore guidance

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
…warm every section (#8440)

* improvement(settings): paint settings section clicks immediately and cut a round trip

- Share the in-flight section navigation between the sidebar and the content area: the clicked section's heading paints over an empty body while its route resolves, with the outgoing section kept laid out but invisible and inert. It is optimistic state inside the navigation's own transition, not a Suspense fallback, so nothing is held back and a redirect back restores the section untouched.
- A press warms data and code but no longer prefetches the route: a prefetch started at mousedown made the navigation wait on a second, two-step request instead of its own.
- Warm every section's chunk on intent from the persistent settings layout instead of the sidebar, so the chunk loads alongside the route payload instead of after it, without growing the workspace chrome's module graph.
- Seed fork availability for admins and read the server-seeded viewer permission, so Workspace Forks renders with the rest of the sidebar instead of after a session fetch and an availability fetch.

* fix(settings): keep press prefetch, guard fork direction switch, tighten docs

- Restore the sidebar link's press-time route prefetch. Measured with a realistic ~90ms press, the prefetch mousedown starts completes before the click commits, so removing it slowed every navigation (commit ~105ms -> ~190ms after mousedown).
- Route the fork sync direction switch through the unsaved-changes guard: switching Push/Pull drops the in-session mapping choices, so unsaved edits confirm first.
- Give SettingsNavigationProvider a props interface, and scope the section layout's 404/307 note to section authorization, which is what an unknown or legacy segment skips.

* fix(settings): confirm every discarded fork choice, keep settings UI out of the workspace chrome

- A fork sync direction switch drops every in-session choice, not just unsaved mapping edits. Expose hasSessionChoices (mapping edits, a copy selection that differs from the default, accepted dropped references, trigger URL choices) and confirm the switch whenever any is set.
- Move SettingsPendingSection into its own module so the provider in the workspace chrome no longer imports the settings header UI.

* fix(fork-sync): clear committed session choices after a sync

A successful sync applies the accepted drops and trigger URL choices, so clear exactly the snapshots it submitted, as it already does for mapping edits; choices made while the request was in flight stay. Compare the copy selection over the visible candidates only, the ones a sync sends, so keys left behind by a completed copy never count as a change.

* chore(settings): rebase module-graph baselines onto staging

Record the settings layout's section import() edges and the workspace chrome's smaller graph against the current staging baseline.
… dormant (#8445)

* fix(knowledge): withdraw a queued generation when its Search KB turns dormant

* fix(knowledge): refund a dormant queued generation only once

* fix(knowledge): withdraw only the exact dormant queue stamp
…failing parent instead of aborting the sync (#8443)

* fix(confluence): retry transient attachment metadata 500s and skip a failing parent instead of aborting the sync

* docs(confluence): list the knowledge base connector's service-account scopes, including attachments

* fix(confluence): log 5xx trace ids from headers without reading the error body

* fix(confluence): keep the 500 status when cancelling an errored response body
…8446)

* fix(execution): publish a pause only after its run log is finalized

* fix(execution): fail a paused run whose log was never finalized instead of publishing it

* fix(execution): restore the publish spy explicitly instead of in a hook
… S3_ENDPOINT in CSP (#8449)

* fix(uploads): sign S3-compatible upload metadata as headers and allow S3_ENDPOINT in CSP

* test(csp): pin S3_FORCE_PATH_STYLE unset in the CSP test env
…read's entitlement (#8451)

* improvement(whitelabeling): gate the settings page on the whitelabel read's entitlement

The page blocked on the full organization billing read (member page plus each member's usage ledger for the period) only to check the plan name. The whitelabel read now also returns the entitlement the update enforces, as the session-policy and data-retention reads do, as an additive field beside `data` so clients from before it parse the same response. The page renders after that one read, which the workspace branding provider usually has cached, and its gate now matches what saving allows.

* fix(whitelabeling): parse the whitelabel read from a server that predates the entitlement field
… stream events (#8450)

* fix(mothership): bound replay frames and end refused turns instead of handing them off

A leased Chat stream persists every event to its Redis replay buffer before
delivering or dispatching it. When the buffer refused a write (a frame over
the 1 MiB single-write ceiling, or a stream past its 32 MiB budget), the
writer threw a generic error, the controller classified it as a takeover and
released its lock without finalizing, and every reconnect started a recovery
controller that re-received and re-refused the same event until the run
deadline. The refused tool call was never dispatched, so the worker waited.

Refusals:
- The writer throws StreamReplayBudgetExhaustedError and rejects later
  publishes the same way, but still delivers the turn's terminal error and
  complete events unpersisted; flush no longer rethrows the refusal.
- The controller aborts with that error rather than a supersession, the
  lifecycle classifies it as an error (not a user cancel), the run is
  finalized as an error with code replay_budget_exhausted, and the worker is
  told to stop. Reconnects then see a terminal run and start no controller.
- Only a genuine StreamControllerSupersededError still hands the run off.

Compaction, applied only to the copy the writer delivers and persists:
- Payloads over 256 KiB have long string leaves cut to an 8 KiB head in
  place, then medium strings to 500 characters, and only then the largest
  unprotected subtrees replaced by { omitted, bytes }, targeting 128 KiB.
  A payload-level streamTruncation marker lists every changed field.
- Identity fields, UI-read keys, arguments of client-executable calls, file
  preview events, and generate_api_key results are never compacted; such a
  frame that stays too large is refused and ends the turn cleanly.
- Oversized assistant text splits into contiguous exact pieces, and each tool
  call forwards only the head of its argument deltas.

* fix(mothership): render stream-omitted tool values honestly

A call frame whose whole argument object was omitted by stream compaction no longer replaces a tool node's arguments, and the permission card and generic tool output render omitted values as "Too large to show (N KB)" instead of raw marker JSON.

* refactor(mothership): replace replay compaction with one string-truncation pass

Compaction is now a single linear walk: when a payload's strings could
serialize past 256 KiB, string leaves longer than 8K units are cut to their
head with an inline "…[truncated, N total]" note, copying only what changes.
Identity and UI-read keys, file previews, generate_api_key results, and the
arguments of calls the browser executes are never cut; anything still over
the write ceiling is refused and ends the turn. The subtree omission, second
cut pass, truncation marker, text splitting, argument-delta cap, and the UI
stub handling are removed.

One predicate, isClientExecutedToolCall, now names the tool calls the browser
starts from the call frame; both the stream compaction and the client
dispatch use it, and the workflow tool names move beside it.

Also fixes three defects in the refusal path:
- A superseded controller that is refused an oversized frame no longer
  expires its successor's stream or clears its abort marker; teardown keys on
  whether this controller handed off, not on the abort reason.
- A refusal after a user Stop stays a cancellation: the abort reason decides.
- The per-user hourly ceiling gets its own message instead of telling the
  user to continue immediately.

* fix(mothership): stop stale approval cards, raw backend bodies, and errored-turn spinners

- A replayed call frame is history and never gates anything, so Sim now clears
  the worker's awaiting_approval stamp on it as it already does for live
  frames. Replays no longer render Allow/Skip cards when approvals are off.
- Stop settles every unfinished tool row (pending, executing, or awaiting
  approval) through one shared predicate, isUnsettledToolState, on the
  persisted, live, and pending-message paths.
- An errored turn settles its unfinished tool rows as errored when it is
  persisted, so they no longer reload as spinners.
- A failed backend response no longer puts the upstream body in the error
  shown to users: 5xx and non-JSON bodies say the service is temporarily
  unavailable, 4xx uses the worker's displayMessage or a generic message.
  The status and body stay on the error for logs.

* fix(mothership): never forward an approval stamp on a frame that cannot be gated

The worker stamps awaiting_approval on resolved integration calls and relies
on Sim to keep or clear it. Sim cleared it on live complete frames that were
not gated, but a partial frame returned before that decision and kept the
stamp. Only a live, complete call frame can be held behind a prompt, so every
replayed or partial frame now drops the stamp at the one point each frame
passes through before it is persisted and delivered. With approvals off, no
frame, replay, snapshot, or persisted row can carry the stamp to the client.

* fix(mothership): keep retrying an unreachable worker through a task replacement

A worker task replacement leaves the load balancer answering 502/504 for
about a minute. Stream retries stopped after three attempts, about four
seconds, and ended the turn with an error. While no worker answers (a 5xx or
no connection), retries now continue with the existing backoff for up to two
minutes from the first failure, still bounded by the run deadline and by
Stop. Each attempt re-sends the same message identity, which the worker
treats as a reattach rather than a second run. Streams that end without a
terminal keep the three-retry limit.

* fix(mothership): settle runs whose terminal events fail, and tighten compaction and cleanup

- A controller that lost its lock while handling an ordinary failure no
  longer expires its successor's stream or clears its abort marker: teardown
  skips cleanup when this controller handed off or was superseded.
- The terminal run status is now written even when publishing the terminal
  events fails. The write stays scoped to this controller's token and never
  replaces a terminal status, so a successor's run is left alone. Before, such
  a failure left the run non-terminal with nothing to settle it.
- A failed backend response is described by its status alone: 5xx says the
  service is temporarily unavailable, anything else that the request could
  not be processed.
- Compaction no longer exempts whole subtrees by key name. Every string the
  UI reads for identity, status, titles, or targets is far shorter than the
  cut, and long text keeps its head, so only client-executed call arguments,
  file previews, and generate_api_key results are left whole.

* fix(mothership): settle tool rows left unfinished by any finished turn

- A completed turn now persists its unfinished tool rows settled as the live
  view settled them at complete, so they no longer reload as spinners.
- A stored message's unfinished tool rows render as interrupted, which settles
  rows persisted before these fixes; the live streaming message is unchanged.
- The client's error finalize settles pending and awaiting-approval rows too:
  finalizeResidualToolCalls reports whether it settled anything, and that
  decides whether the snapshot is rewritten.

* fix(mothership): harden replay compaction, stream retries, teardown, and backend errors

- Assistant text is never compacted: its length is part of the text receipt a
  replacement controller and the worker check. When cutting strings alone
  cannot bound an event, long arrays keep their first 100 items and a count.
- Only a gateway error page (502/503/504 without a JSON body) or a failed
  connection counts as an unreachable worker for the two-minute window; a
  JSON 5xx from a reachable worker keeps the three-retry limit. A retry that
  streams again resets the budget for the next outage.
- A controller cleans up its stream's buffer and abort marker only while it
  still holds the chat lock, checked just before cleanup and before releasing
  the lock, instead of inferring ownership from flags.
- A controller superseded while publishing its terminal events leaves the run
  to its successor; other publish failures still settle the run.
- A 4xx shows the worker's own short, plain reason (rate limit, BYOK, plan
  mode, model selection, version skew) and hides internal identifiers.
- File preview snapshots are spaced by size, so a large file's patch preview
  can no longer fill the stream's replay budget; small files keep their pace.

* fix(mothership): separate retry budgets, cap preview totals, and clean up only ended turns

- An unreachable worker and a reachable one now have independent retry
  budgets. Any answer from the worker restarts only the two-minute
  unreachable window; the three-retry, 30 s budget for failures of a worker
  that answered keeps its original semantics, so a failure that repeats
  after every reattach still stops, and a long outage no longer spends the
  replacement's retries.
- Intermediate file preview content is capped at 8M characters per edit;
  past it the preview holds until the final snapshot, which is always sent.
- A controller cleans up its stream only when it finalized the turn and
  still holds the lock, so a run it leaves recoverable keeps its buffer and
  any pending Stop.
- Remove the unreachable generate_api_key compaction exemption, avoid a
  second copy of trimmed arrays, and correct the size-estimate comment.

* fix(mothership): clean up turns whose terminal publish failed and cap previews per turn in bytes

- A controller that ended its turn but could not publish the terminal events
  still cleans up its stream: the run is settled either way, so the buffer no
  longer lingers for its full TTL. Only a superseded controller skips cleanup,
  and the lease check still protects a successor.
- The intermediate file preview cap is now 8 MiB of UTF-8 per stream loop,
  across all of its edits, instead of characters per edit. Each edit's final
  snapshot is still always sent.

* fix(mothership): bound preview frames per write and per turn, and narrow unreachable retries

- A file preview frame whose content would not fit one replay write, or that
  would exceed the turn's preview budget, is skipped instead of sent: the
  preview holds its last content and the client loads the stored file on
  completion. A completion too large to send whole drops its tool output.
  The per-frame limit derives from the replay buffer's single-write ceiling,
  and the 8 MiB budget now covers every preview frame in the turn, finals
  included, across all legs. A large file preview no longer ends the turn.
- Only a request that fails before any response headers counts as an
  unreachable worker. A TypeError raised after the response began, including
  one from handling delivered events, keeps the three-retry budget, so a
  deterministic failure can no longer retry until the run deadline.
- A worker that cannot be reached is reported with the generic "temporarily
  unavailable" message; the network error stays on the error's cause.
- Retry logs and spans count retries from both budgets.

* fix(mothership): omit the largest field of an event no cut can bound, and persist closed lanes

- An event whose many short values survive every string and array cut, such
  as an object with thousands of keys, now has its largest bulk field
  replaced by a size note when it would still exceed one replay write,
  instead of ending the turn. Identity fields and client-executed arguments
  are never replaced; such an event is still refused cleanly.
- finalizeResidualToolCalls reports closing an open subagent lane as a
  change, so the error finalizer persists it instead of leaving a stale lane.

* fix(mothership): keep sibling fields when omitting bulk, and never cache a skipped preview as file content

- The last-resort omission for an event no cut can bound now walks down from
  the largest field while one child holds most of its parent's bytes and
  replaces only that node, so the ids, resources, and status fields beside
  the bulk keep driving the UI's resource updates.
- On completion the client seeds the file content cache from preview text
  only when it received the completed version. When the server skipped the
  final content frame, the text is an earlier draft, so the stored file loads
  instead.

* fix(mothership): keep omitting bulk until an event fits one replay write

A result with more than one large object of short fields left the second
one over the write limit after the first was omitted, so the event was
refused and the turn ended. Omission now repeats until the event fits or
nothing large enough is left; each pass removes at least 64 KiB, so it
always terminates.

* fix(mothership): shed oversized replay events in one pass and retry mid-body stream cuts

- Replace the replay-compaction omit loop with one bounded pass over memoized
  exact serialized sizes: recurse into the smallest child that covers the
  overage, otherwise drop the largest children whole, keeping small siblings
- Map a non-abort body-read failure from the worker stream to a retryable
  WorkerStreamInterruptedError on the reachable budget, with the generic
  message and the original error on cause
- Log the cause message on retry warnings and orchestration failures
- Count replayed file preview content toward a restored turn's preview budget

* refactor(mothership): tidy replay compaction and stream retry after review

- Share the replay write ceiling between compaction and file previews
- Scan for the smallest sufficient child only once one can cover the need,
  and reuse one size memo across every compaction step
- Stop retrying a raw TypeError; fetch and body-read failures are already
  wrapped as WorkerUnreachableError and WorkerStreamInterruptedError
- Use absolute imports, a private retry counter, a typed unsettled-state set,
  and hasUnsettledTool naming
- Cover a timed-out body read and the preview completion edge cases

* fix(mothership): bound preview metadata frames and settle stopped rows on every save

- Compact every preview phase except content and completion, which the preview
  adapter already bounds; a model-written patch search string could otherwise
  exceed one replay write in an edit_meta frame
- Settle unfinished tool rows as stopped in buildPersistedAssistantMessage for a
  cancelled turn, so background and API callers that persist the result
  directly never save pending, executing, or awaiting-approval rows
@waleedlatif1
waleedlatif1 merged commit 7cf57f3 into main Sep 30, 2026
66 of 67 checks passed

This branch was previously deployed

1 inactive deployment
Preview — 85402d08 Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants