Skip to content

v0.9.15: mship improvements, memory improvements, nextjs bump, plane integration - #8748

Open
waleedlatif1 wants to merge 118 commits into
mainfrom
staging
Open

waleedlatif1 wants to merge 118 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

icecrasher321 and others added 30 commits October 5, 2026 00:52
…he-top-options-compare (#8612)

* docs(library): update agentic-ai-coding-tools-what-they-are-and-how-the-top-options-compare

* Pi Babysit: address PR #8612 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…8613)

* feat(library): How do you build an AI bot for Discord without code?

* Pi Babysit: address PR #8613 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…ead of extra E2B round trips (#8621)

* perf(sandbox): grant the run_code session lease in the reconnect instead of extra E2B round trips

Reused Mothership workbench calls made five E2B control-plane requests:
list, connect, getInfo + setTimeout on acquire (the set always fired), and
getInfo on release. Connect now asks for max(5 min, remaining, lease), the
handle records the deadline it requested, and acquisition/release skip the
provider while that lower bound covers the request. Final deadlines are
unchanged; unrequested deadlines are still read back.

* test(sandbox): model connect as setting the deadline so a dropped preserve is caught
The simple Chat effort picker labeled medium as Low, high as Medium, and
xhigh as High. Derive its options from MOTHERSHIP_EFFORT_OPTIONS so each
label names the effort it sends: Medium, High, Extra High. Values, the
default (high), and the stored preference are unchanged.
…un --stop-after` (#8622)

* feat(workflows): stop a manual v2 run after a block, and `workflows run --stop-after`

A manual v2 run can now name `run.stopAfterBlockId`; the run stops once that
block completes and downstream blocks do not execute. Combined with a block
entry on the same block, it re-runs exactly one block against a prior run's
persisted upstream outputs, server-side:

  sim workflows run W --from-block X --source-run R --stop-after X --select-output X.result

Agents verifying an edit no longer re-run every upstream block (often a slow
LLM or API call) or toggle blocks off to skip them.

- Contract: optional `stopAfterBlockId` on the manual run selection.
- Application: both manual operations refuse a block missing from the saved
  workflow or nested in a loop/parallel (the engine would otherwise run to the
  end or stop after one iteration), before anything runs.
- Execute service: threads the trusted value to the sync and stream paths.
- CLI: `--stop-after <blockId>` implies --manual and rejects --async.
- E2E: test-workflow-stop-after-e2e.ts against a running app; the http-e2e job
  gains a Redis service because hosted billing admits runs through a Redis
  usage reservation.

* fix(workflows): refuse stop targets a run cannot reach, and run the E2E self-hosted

- The manual operations refuse a stop block the run cannot reach from its
  entry (an upstream block would let the run finish everything after the
  entry), and look blocks up as own properties.
- The executor fails a run whose stop block is absent from the workflow it
  executes, instead of running everything; this closes the window between
  validation and the executor's own draft load, for every caller.
- The CLI refuses an empty --stop-after rather than dropping it.
- CI: the stop-after E2E gets its own self-hosted app step; the SCIM suite
  asserts PostgreSQL rate-limit storage, so Redis is not added to that app.
  Fixture cleanup waits for run logs to finalize before deleting.

* fix(workflows): refuse a disabled stop block, or one reached only through one

The executor omits disabled blocks from its graph, so a disabled stop target,
or one whose only path runs through a disabled block, is never reached and
the run would finish everything after the entry.

* fix(workflows): the executor refuses a disabled stop block too

The serialized workflow keeps disabled blocks, but the DAG skips them, so a
disabled stop target would never be reached.

---------

Co-authored-by: Waleed Latif <waleed@sim.ai>
…pletion (#8620)

* fix(executor): stop retaining duplicate copies of loop block outputs

* test(executor): guard output sharing in block logs and loop aggregates

* fix(logs): size execution data the way JSON.stringify writes it

* fix(logs): count JSON string bytes without copying and unbox primitive wrappers

* fix(logs): measure execution data iteratively and apply toJSON on functions

* fix(executor): drop the full JSON clone of execution state at run completion

* fix(executor): normalize only live state for PII masking and walk serializability lazily

* fix(executor): snapshot array lengths and unbox wrappers in JSON walks

* fix(logs): read boxed boolean and bigint values the way JSON.stringify does
* feat(projects): add project identity and lifecycle foundation

* feat(projects): create projects with their initial environment

* docs(projects): record project files follow-up

* docs(projects): explain project and workspace creation flows

* fix(projects): stage activation after compatible writers deploy

* refactor(projects): prepare compatible writers for the SQL backfill

* fix(projects): clean up automatically created fixture Projects

* fix(workflows): guard restore against concurrent workspace archive

* fix(projects): close lifecycle races and surface rollout conflicts

* fix(workflows): return not found when import loses archive race
* fix(mcp): restrict MCP server destination changes to admins

* fix(mcp): compare exact MCP paths and guard concurrent URL changes

* fix(mcp): treat setting a URL on a URL-less server as a destination change

* fix(mcp): guard re-registration against concurrent URL changes

* fix(mcp): narrow re-registration URL before the guarded update

* chore(mcp): use absolute import in utils test
…ck (#8635)

* fix(executor): fail a stop-after run whose routing skips the stop block

A run with stopAfterBlockId only stopped when the stop block completed. When a
router, condition, or untaken error path routed the run away from it, the stop
never triggered and the run finished every other branch, reporting success as
if it had stopped there. A static check before the run cannot see this.

- The engine ends the run as soon as every path into the stop block has been
  deactivated, before any further block starts, and fails it with
  `Stop block "<name>" (<id>) was not reached: no path this run took leads to it`.
- Any run that ends without completing its stop block fails the same way: a
  stop block missing from the executed graph, or a Response block that ended
  the run first.
- A loop or parallel stop with nothing to run completes at its start sentinel,
  whose end sentinel never runs, so that exit now counts as reaching it.
- The v2 contract and the CLI `--stop-after` help describe the failure.
- E2E: a condition fixture checks the stop on the taken branch still stops
  there, a stop on the skipped branch fails the run before the other branch's
  slow block finishes, and the CLI exits non-zero.

* fix(executor): a skipped stop block fails a run another branch paused, and names a Response ending

- A run whose stop block was proven unreachable fails even when another branch
  paused, instead of returning a paused run that would resume past it.
- When a Response block ended the run first, the error says so rather than
  claiming no path leads to the stop block.
…le (#8631)

* fix(projects): restore workspace deletion and tighten Project lifecycle

- Archive a Project with its last active environment instead of refusing the
  workspace delete; account deletion follows the same rule, and the implicit
  archive is audited
- Gate Project APIs on a `projects` AppConfig flag (PROJECT_API_ENABLED fallback)
- Run Project reads in a read-only snapshot without locks; list Projects from the
  caller's grants with batched authorization
- Batch workflow archival, Project transfer and owner reassignment; move Project
  ownership on organization ownership transfer
- Reuse shared advisory-lock and text-array helpers; narrow admin-move conflict
  mapping to ProjectConflictError

* improvement(projects): unify environment archive and align with shared patterns

- Archive a workspace's workflows atomically with it through one
  archiveEnvironmentInTransaction shared by workspace delete and Project archive;
  the workspace row is locked before the sweep so concurrent creates are covered
- Scope the Project lock timeout to lock acquisition and map lock timeouts and
  deadlocks to a retryable conflict; backfill and multi-Project locks use the
  shared advisory-lock helpers in code-unit order
- Shared orchestrationFailureResponse for raw routes; contracts use the ID
  primitives and export only what is consumed; audit enums and mock in sync
- Project restrictions section matches its sibling settings rows
- Batch account-deletion Project loads/locks; skip inconsistent Projects in lists
- Harden the foundation integration suite (user-keyed cleanup, poll helper,
  pid-scoped waits, precise assertions)

* fix(projects): trim the requested organization id before validating it

* fix(projects): address review on Project locking, archive notifications and list policy cost

* fix(projects): keep archive retries from re-stamping MCP servers and isolate post-commit notifications
…m, restore Low (#8634)

* feat(mothership): keep each chat's reasoning effort, default to medium, restore Low

The simple picker offers Low / Medium / High / Extra High again, each sending
exactly that effort. New chats and chats never changed run at medium instead
of high. An effort the user picks is stored on the chat (copilot_chats.config)
through a new PUT /api/mothership/chats/[chatId]/effort and at turn admission,
and later turns of that chat keep it. The global last-used effort is no longer
persisted. The Sim Chat block defaults to medium.

* fix(mothership): keep the latest effort pick through refetches, failed saves and abandoned new chats

* fix(mothership): leave a deduplicated send's chat on the pick its first attempt stored

* fix(mothership): show a recovered chat's pick while its details load

* fix(mothership): hand a withdrawn first send's effort pick back to the new-chat composer

* fix(mothership): hand a withdrawn send's pick back only while its new-chat surface is open
… Tools Compared (#8638)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
… Integrations (#8639)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…gine-optimization-actually-mean (#8647)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
#8643)

* fix(mothership): refuse desktop claims for unapproved or stopped calls

The desktop authorize route claimed any pending call, so a gated terminal
run that was still awaiting approval, or a call on a run the user had
stopped, could be claimed and executed. The claim now locks the run row and
refuses once tool admission has closed (Stop, a newer turn, or the run's
end), and refuses a call held for the user's decision until they allow it.
Whether a call is gated depends on the turn, so pre-persist records it on
the row (permission_requested_at).

Stop now settles the stopped runs' open desktop calls in the transaction
that closes admission: unclaimed calls as never started, claimed calls as
outcome unknown. A result for an already-settled call (a retry, or one that
lost to Stop) is acknowledged with the stored outcome instead of 404/500.

* fix(mothership): settle every open desktop call on Stop and answer 410 after it

Stop also settles delivered desktop calls and reads of granted local
folders. Authorize checks admission before the call's status, so a call
Stop already settled answers 410 rather than 404.

* test(mothership): assert the acknowledged outcome, not the publish mock

* test(mothership): give the hand-built tool call table the new permission column

* refactor(mothership): one claim primitive, one desktop-tool classifier, sealed Stop results

The desktop claim is now an option of the run-locked tool execution claim
(claimSimToolExecution becomes claimToolExecution) instead of a second copy
of the admission check. Stop picks the open desktop calls with the shared TS
classifier, which moves to lib/mothership/tools/desktop-tools.ts, instead of
a SQL restatement of it, and seals each result the way the confirm route
does, so a waiter restores what Stop did rather than failing to unseal it.

* fix(mothership): a declined call stays unclaimable without its gate marker

Calls gated before permission_requested_at existed carry no marker, so a
recorded decision that does not allow the call now disqualifies it too.

* test(mothership): assert authorize outcomes, not claim mock calls
…#8655)

Picks queued behind an in-flight save run onMutate at once, so after low, high, low a failed first save dropped the newest low and the picker fell back to the stale server effort. Each pick now carries a token and a failed save drops only its own pick.
… JSON serializability (#8658)

* fix(executor): unwrap boxed primitives by internal slot when checking JSON serializability

* fix(executor): check the boxed slot first and convert wrappers with ToNumber/ToString
… moved to (#8741)

* fix(mothership): send a late queue write to the chat a new-chat queue moved to

On the new-chat surface, a Send-now whose Stop saw the first message admitted
moved the queue to that chat, but the follow-up's busy refusal was re-queued
under the dead new-chat key it was dispatched from: gone from the chat, never
retried there, and liable to be adopted into a new chat later. migrate now
records where a key moved, and every write that captured a key before an
await (the dispatch's removal and restore, the direct send's re-queue, the
history check's defer and drop) resolves it at write time (liveQueueKey).
Every re-queue on a chatless surface also carries the surface, so one that
lands after the surface unmounted can still be adopted.

* fix(mothership): keep a late write behind what the chat's queue already held

When the new-chat queue moved into a chat queue that already had messages,
migrate put those first, but a late write still used its index in the
new-chat queue and could land ahead of them. The move now records how many
messages it went behind, and late writes resolve their position, not just
their key (liveQueuePosition).

* fix(mothership): anchor a late re-queue on the messages ahead of it, not an index

A restored message went back at an index captured at dispatch, offset by
how many messages a chat's queue held when the new-chat queue moved into it.
Removing any of those while the POST was out shifted it behind a newer
message. It now goes right after the last message still queued that was
ahead of it (at dispatch, or in the chat's queue before the move), else at
the head.
…send (#8743)

* refactor(mothership): one hold field and one retry field on a queued send

A queued message's wait is now hold ('user' or 'online') instead of
retryRequired plus heldUntilOnline, and its automatic retry is
retry: { attempt, notBefore } instead of sendRetries plus notBefore, which
folds ScheduledRetry away. Queues saved in the older shape are mapped when
the session restores them.

* refactor(mothership): tidy the forwarding of late queue writes

- one lookup instead of a hop loop: only a new-chat key moves, and only to
  its chat's key, which never does; liveQueueKey derives from
  liveQueuePosition
- migrate keeps the first record, so a repeat cannot rewrite what was ahead
- the history check drops its no-op forwarding: it never runs on a new-chat key
- the dead-key DOM test waits on the queue instead of a fixed sleep

* test(mothership): cover migrate keeping its first record
…the server already has (#8744)

* refactor(mothership): a pure resend verdict, one discard helper, and Send-now honours it

- resendVerdict(entry, history | null) returns send, wait or drop, and the
  queue drain acts on it, instead of a boolean check with side effects
- discardQueuedSend replaces the three copies of clear handoff, clear claim,
  remove
- Send-now checks the history too and drops a message the server already
  accepted instead of resending it (G5); a history it cannot read does not
  hold back a send the user asked for
- the own-id conflict test reaches that branch through a restored Send-now,
  since a message the history shows accepted is now dropped first

* test(mothership): lock the queue upgrade path and tighten held waits

- rehydrate rows for the production shape (retryRequired: false, no hold)
  and for a queue already in the new shape
- the offline-hold waits check hold === 'online' instead of any hold
- the one-lookup comment names where the one-hop invariant is enforced

* fix(mothership): don't stop the running turn for a Send-now gone during its history read

Send-now reads the chat's history before stopping the running turn. A
message removed, edited or dispatched in that time, or a view that moved on,
still stopped the turn. It now re-reads the queue and checks the chat and
mount are current before going on.

* test(mothership): cover Send-now's re-checks after its history read

Send-now no longer stops a turn in a chat the user moved to during the
read, or from a surface that unmounted during it, and leaves a message
the drain dispatched meanwhile to that dispatch. applyHeldResend is
renamed applyResendVerdict.
)

resumeUserMessageId is the only id a queue entry goes out under; the Stop
handoff seed no longer carries a copy and reusedRequestId is gone. The
stored handoff record keeps its userMessageId and is converted where it
enters and leaves the queue. Queues saved with the id on the seed move it
to resumeUserMessageId on rehydrate.
…ase its session renews (#8742)

* fix(desktop): keep a chat-view import alive while it works, by the lease its session renews

An import the chat view runs was failed as lost once it ran past the default tool budget (60 s plus
the 30 s resume grace), though it was still uploading. Its claim now takes the execution lease
under the claiming session, the chat view renews it through the existing lease route while the
import runs, and the turn's wait budget runs to the end of that lease. Without renewals the lease
lapses with the default budget, so a closed or crashed window still settles within about a lease.

* fix(desktop): keep renewing an import's lease through transient failures, and retry a failed lease lookup

- The chat view stops renewing only when the server refuses the call (410)
- The resume watchdog retries a failed lease lookup for up to one lease instead of treating it as a lapse
- The lifecycle tests assert what the agent is resumed with, and when

* fix(desktop): cap a renewed import's wait at the client tool limit, renew at once, and report the extended wait

- A chat-view import's lease extends its wait only up to the cap every client tool has
  (CLIENT_TOOL_RESULT_TIMEOUT_MS), so an import that hangs with its page alive still settles
- The page renews the lease as soon as the import starts, then every heartbeat
- The force-fail log names an extended wait and how long it lasted; the wait span's budget
  includes the extension
- Tests for the cap, the bound on failed lease lookups, and the client heartbeat

* fix(desktop): bound lease lookups, never fail a replaced call, and renew from the start of an import

- Each lease lookup gets 5 s (and Stop) before it counts as failed, so a stalled read cannot hold
  the wait past its deadlines
- A call replaced while its lease was read is left to its new watchdog
- The page renews from the moment it asks for the manifest; a refusal counts only once the claim
  is confirmed, and renewing stops on every exit
- Heartbeat tests check the lease a fake server holds, not request counts

* fix(desktop): judge a lease refusal by whether the claim was confirmed when the renewal was sent

* fix(desktop): renew an import's lease only after its claim, and give a capped call up without reading its lease

The first heartbeat now comes one beat in, after the desktop's bounded claim, so every renewal
follows the claim and a refusal always means the call was stopped, settled, or lapsed. A call at
its ceiling is given up before its lease is read, and the force-fail log names whether the
budget, the cap, a lapsed lease, or failed lease lookups ended the wait.

* test(desktop): give the renewal test's lease room for a slow round trip
When an entry saved with an id on its Stop handoff also has its own
resumeUserMessageId, the handoff's id is the one that build sent, so it
is the one kept. resumeUserMessageId's doc names every path that sets it.
@waleedlatif1
waleedlatif1 requested a review from a team as a code owner October 7, 2026 16:41
@vercel

vercel Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 7, 2026 7:26pm UTC

Request Review

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

We detected this is a high-risk PR and are running a free ultrareview. An ultrareview is a deeper, multi-pass review that catches hard-to-find bugs a standard review can miss. We'll post the findings when it completes.

This PR appears to change concurrency-sensitive code such as locks, queues, or retries, where a missed race may only surface under production load, so a deeper multi-pass review is worth running.

Want an ultrareview on every high-risk PR? Set up automated ultrareviews.

Comment thread packages/desktop-bridge/src/local-filesystem-tools.ts Fixed
@greptile-apps

greptile-apps Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

Fix desktop cancellation before merging so Stop reaches commands whose results are waiting for delivery.

Findings

  1. P1 Stop leaves a command running ▶
  2. P2 Row-change log stays empty ▶

Summary

This release adds background desktop tool execution, improves chat send recovery, adds Plane tools and triggers, and updates workflow execution, settings, background jobs, and Next.js.

  • Stop can leave a terminal command running while its result waits for delivery.
  • The new table row-change log is not connected to production writes.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  Offer[Sim offers a desktop call] --> Claim[Desktop claims and records it]
  Claim --> Run[Run the tool]
  Run --> Result[Tool returns a result]
  Result --> Report[Retry delivery to Sim]
  Report --> Accepted[Sim accepts the result]
  Accepted --> Release[Release the call]
  Run --> StillRunning[Terminal command may still be running]
  StillRunning --> Report
  Stop[User presses Stop] --> Cancel[Sim cancels the call]
  Cancel --> Phase{Desktop call phase}
  Phase -->|running| NativeStop[Cancel native work]
  Phase -->|reporting| Gap[Native cancellation is skipped]
Loading

Reviews (1) · Last reviewed commit: "refactor(mothership): a saved Stop hando..." · Reviewed by Greptile

Comment thread apps/desktop/src/main/desktop-executor/executor.ts
Comment thread packages/db/migrations/0396_user_table_row_changes.sql

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We detected this is a high-risk PR and ran a free ultrareview. An ultrareview is a deeper, multi-pass review that catches hard-to-find bugs a standard review can miss.

This PR appears to change concurrency-sensitive code such as locks, queues, or retries, where a missed race may only surface under production load, so a deeper multi-pass review is worth running.

Want an ultrareview on every high-risk PR? Set up automated ultrareviews.

24 issues found across 831 files

Confidence score: 2/5

  • shell.ts misses arithmetic contexts in heredoc bodies and [[ ... -eq ... ]], so placeholders can be compiled where they should be rejected. Preserve the enclosing context and detect these arithmetic operands.
  • e2b.ts can promise a lease that extends beyond the sandbox’s absolute lifetime, leaving the workbench unavailable before the lease ends. Cap the lease to the provider limit and recover or create a sandbox when it won’t fit.
  • outbox-events.ts can leave Stripe’s cancel_at_period_end out of sync with the final database value when updates for one subscription run concurrently. Make processing converge on the latest value.
  • mothership.tsx can discard the generated key when the user switches tabs or leaves settings without warning, even though it is labeled as shown only once. Include generatedKey in the dirty check.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apps/docs/lib/openapi.ts">

<violation number="1" location="apps/docs/lib/openapi.ts:6">
P2: These top-level imports eagerly load all seven API documents for non-API docs pages as well, adding roughly 2.5 MB of spec data to normal server startup. Keep the API specs behind a lazy server-only module boundary while still bundling them for deployment.</violation>

<violation number="2" location="apps/docs/lib/openapi.ts:74">
P3: This repeats the same seven imports and filename map already maintained in `apps/docs/lib/openapi-download.ts`; future spec additions or renames now require synchronized edits in two modules. Move the shared map into one module and consume it from both callers.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/home/components/user-input/components/model-selector.tsx">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/home/components/user-input/components/model-selector.tsx:58">
P2: This cleanup erases the selected new-chat effort when the optimistic first-send view swaps composers, so a failed or withdrawn send returns with the default instead of the user’s pick. Clear the pick only when the chatless surface is actually abandoned, not when its composer swaps during an optimistic send.</violation>
</file>

<file name="apps/sim/app/api/desktop/inbox/stream/route.ts">

<violation number="1" location="apps/sim/app/api/desktop/inbox/stream/route.ts:32">
P2: An abort during authentication or device setup can hand an already-aborted signal to `createSSEStream`; its listener misses that abort and leaves the Redis doorbell subscription alive. Make `createSSEStream` handle a pre-aborted request before installing the listener (or return 499 here before opening the stream).</violation>
</file>

<file name="apps/sim/app/api/v1/admin/workflows/import/route.ts">

<violation number="1" location="apps/sim/app/api/v1/admin/workflows/import/route.ts:127">
P2: `createWorkflowWithState` can throw a classified `conflict` after its name-race retries, but this route catches only `not_found`; concurrent imports therefore return a misleading 500. Map `conflict` to `conflictResponse` so callers receive a retryable 409.</violation>

<violation number="2" location="apps/sim/app/api/v1/admin/workflows/import/route.ts:164">
P2: The broad `not_found` branch mislabels a concurrent folder deletion as a missing workspace. Handle the helper's folder-not-found result separately before mapping the workspace archive race, preserving the folder-specific response.</violation>
</file>

<file name="apps/sim/ee/access-control/components/group-detail.tsx">

<violation number="1" location="apps/sim/ee/access-control/components/group-detail.tsx:875">
P3: `hasConfigChanges` now performs an O(config-array-size) comparison and allocates new sets on every render. Restore the `useMemo` around this comparison so unrelated GroupDetail renders do not repeatedly rebuild the same collections.</violation>
</file>

<file name="apps/sim/ee/workspace-forking/components/fork-sync-detail-view/fork-sync-detail-view.tsx">

<violation number="1" location="apps/sim/ee/workspace-forking/components/fork-sync-detail-view/fork-sync-detail-view.tsx:70">
P2: `Open workspace` bypasses this guard. When `hasSessionChoices` is true but `dirty` is false, or while `submitting` is true, the header still exposes that imperative navigation and users can lose session choices or leave an in-flight sync; route this leaving action through `guard.guardBack` or otherwise guard/disable it.</violation>
</file>

<file name="apps/sim/ee/access-control/components/project-issue-restrictions.tsx">

<violation number="1" location="apps/sim/ee/access-control/components/project-issue-restrictions.tsx:28">
P3: This guard hides configured restrictions after their project is archived, leaving admins no way to see or remove the stale ID. Preserve selected project metadata in the UI or clean the ID when archiving the project.</violation>

<violation number="2" location="apps/sim/ee/access-control/components/project-issue-restrictions.tsx:50">
P3: An initial project-list failure leaves this editor without an in-place retry after React Query exhausts its retry. Render `SettingsQueryErrorState` with `projects.refetch` so admins can recover without reloading the settings page.</violation>
</file>

<file name="apps/sim/lib/execution/code-placeholders/shell.ts">

<violation number="1" location="apps/sim/lib/execution/code-placeholders/shell.ts:495">
P1: Each heredoc-body scan resets `arithmeticDepth`, so it loses an enclosing arithmetic substitution. Carry the enclosing arithmetic context into heredoc scans or reject body placeholders whose output can be consumed by outer arithmetic.</violation>

<violation number="2" location="apps/sim/lib/execution/code-placeholders/shell.ts:541">
P1: `[[ ... -eq ... ]]` is an arithmetic context, but this detector only recognizes delimiter-based forms, so placeholders there are compiled instead of rejected. Detect arithmetic operands in `[[ ]]` and the other shell arithmetic contexts before resolving the placeholder; otherwise values containing `$(...)` are re-evaluated by Bash.</violation>
</file>

<file name="apps/sim/lib/execution/remote-sandbox/e2b.ts">

<violation number="1" location="apps/sim/lib/execution/remote-sandbox/e2b.ts:1133">
P1: This reconnect can promise a lease past E2B’s absolute sandbox lifetime. Cap the usable deadline against the sandbox’s provider-limit deadline and create or recover a new workbench when the requested lease cannot fit.</violation>
</file>

<file name="apps/sim/app/api/copilot/tool-permission/route.ts">

<violation number="1" location="apps/sim/app/api/copilot/tool-permission/route.ts:141">
P3: `desktopDeviceId` identifies the whole turn, not the approved call, so this also rings the executor for non-desktop approvals such as `run_workflow`. The inbox filters those rows out, but every answer still causes an unnecessary authenticated inbox read/reconcile; gate the ring on `isDesktopToolCall(claimed.toolName, existing.args)` after normalizing the stored args.</violation>
</file>

<file name="apps/sim/app/_shell/consent/google-analytics-page-view-tracker.tsx">

<violation number="1" location="apps/sim/app/_shell/consent/google-analytics-page-view-tracker.tsx:17">
P2: This context update does not run for same-path query navigations, so Google can retain the previous sanitized URL and referrer after the URL changes. Track the search portion as well, such as by adding `useSearchParams()` (or an equivalent URL key) to the effect dependencies, and cover a query-only navigation.</violation>
</file>

<file name="apps/sim/ee/sso/components/sso-settings.tsx">

<violation number="1" location="apps/sim/ee/sso/components/sso-settings.tsx:240">
P2: `VerifiedDomainsSection` now mounts on every SSO tab, but `useOrganizationDomains` remains enabled whenever `organizationId` exists, so opening Sign-in or Provisioning also fetches the domains endpoint unnecessarily. Gate the domains query with `active` while keeping the component mounted to preserve drafts.</violation>
</file>

<file name="apps/sim/app/api/cron/fold-table-row-changes/route.ts">

<violation number="1" location="apps/sim/app/api/cron/fold-table-row-changes/route.ts:31">
P2: `foldPendingTableRowChanges` can start another page query after its 45-second deadline. A slow query can consume the route's remaining 15 seconds and hit `maxDuration`, so the cron exits without reporting or folding the remaining tables; check the deadline before fetching each page or pass an absolute deadline into the helper.

(Based on your team's feedback about per-page sweep budgets.)</violation>
</file>

<file name="apps/sim/lib/billing/webhooks/outbox-events.ts">

<violation number="1" location="apps/sim/lib/billing/webhooks/outbox-events.ts:5">
P1: This event does not guarantee convergence when rows for one subscription run concurrently. An older DB read can finish its Stripe request after a newer update, leaving `cancel_at_period_end` opposite the final DB value; serialize or coalesce per-subscription processing, or recheck and repair before completing the event.</violation>
</file>

<file name="apps/sim/app/api/superuser/import-workflow/route.ts">

<violation number="1" location="apps/sim/app/api/superuser/import-workflow/route.ts:215">
P2: `createWorkflowWithState` can throw a classified `conflict` after exhausting its name-race retries, but this catch handles only `not_found` and returns a generic 500. Map `conflict` to 409 (or use the shared orchestration status mapping) so concurrent imports remain retryable.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/settings/components/mothership/mothership.tsx">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/settings/components/mothership/mothership.tsx:370">
P1: `generatedKey` is not included in the dirty check, although it is the only copy of a key labeled “only shown once.” Include it so switching tabs or leaving settings requires confirmation before local state is unmounted and the key is lost.</violation>
</file>

<file name="apps/docs/app/llms-full.txt/route.ts">

<violation number="1" location="apps/docs/app/llms-full.txt/route.ts:7">
P2: `force-dynamic` removes the previous indefinite cache from this public, build-time docs endpoint. Each request now walks all bundled pages again, so crawler traffic repeatedly pays the full generation cost; keep the streamed response cacheable with an explicit public revalidation/CDN policy.</violation>

<violation number="2" location="apps/docs/app/llms-full.txt/route.ts:33">
P2: A `getLLMText` failure after the first chunk now terminates a response that has already been sent with status 200. Clients that do not surface body-stream errors can accept an incomplete documentation file; preserve an atomic error/status contract or expose an explicit terminal failure.</violation>
</file>

<file name="apps/sim/connectors/plane/utils.ts">

<violation number="1" location="apps/sim/connectors/plane/utils.ts:24">
P2: `parsePlanePage` discards the current page and fails the sync when Plane reports more results without an advancing cursor. Return valid results as non-authoritative partial progress with a notice to narrow the source instead of failing the account.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/settings/components/admin/admin.tsx">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/settings/components/admin/admin.tsx:247">
P1: Disable every impersonate control while one impersonation is transitioning. `pendingUserIds` contains only the current target, so another row can start a second session-changing request before the first cleanup and redirect finish; include the existing `impersonatingUserId` state in this condition.</violation>
</file>

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.

Turn on auto-fix | Re-trigger cubic

Comment thread .github/scripts/desktop-live-changes.sh Outdated
@@ -235,7 +244,7 @@ export function Admin() {
<Chip
aria-label={`Impersonate ${u.email}`}
onClick={() => handleImpersonate(u.id, u.email)}
disabled={pendingUserIds.has(u.id)}
disabled={importWorkflow.isPending || pendingUserIds.has(u.id)}

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: Disable every impersonate control while one impersonation is transitioning. pendingUserIds contains only the current target, so another row can start a second session-changing request before the first cleanup and redirect finish; include the existing impersonatingUserId state in this condition.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/app/workspace/[workspaceId]/settings/components/admin/admin.tsx, line 247:

<comment>Disable every impersonate control while one impersonation is transitioning. `pendingUserIds` contains only the current target, so another row can start a second session-changing request before the first cleanup and redirect finish; include the existing `impersonatingUserId` state in this condition.</comment>

<file context>
@@ -235,7 +244,7 @@ export function Admin() {
               aria-label={`Impersonate ${u.email}`}
               onClick={() => handleImpersonate(u.id, u.email)}
-              disabled={pendingUserIds.has(u.id)}
+              disabled={importWorkflow.isPending || pendingUserIds.has(u.id)}
             >
               {impersonatingUserId === u.id ? 'Switching...' : 'Impersonate'}
</file context>
Suggested change
disabled={importWorkflow.isPending || pendingUserIds.has(u.id)}
disabled={importWorkflow.isPending || impersonatingUserId !== null || pendingUserIds.has(u.id)}
Fix with cubic

@@ -488,6 +492,7 @@ function collectShellOccurrenceContexts(
{ kind: 'root', quote: 'none', parenthesisDepth: 0, literalRoot },
]
let skippedRangeIndex = 0
let arithmeticDepth = 0

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: Each heredoc-body scan resets arithmeticDepth, so it loses an enclosing arithmetic substitution. Carry the enclosing arithmetic context into heredoc scans or reject body placeholders whose output can be consumed by outer arithmetic.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/lib/execution/code-placeholders/shell.ts, line 495:

<comment>Each heredoc-body scan resets `arithmeticDepth`, so it loses an enclosing arithmetic substitution. Carry the enclosing arithmetic context into heredoc scans or reject body placeholders whose output can be consumed by outer arithmetic.</comment>

<file context>
@@ -488,6 +492,7 @@ function collectShellOccurrenceContexts(
     { kind: 'root', quote: 'none', parenthesisDepth: 0, literalRoot },
   ]
   let skippedRangeIndex = 0
+  let arithmeticDepth = 0
 
   for (let index = start; index < end; ) {
</file context>
Fix with cubic

@@ -533,6 +538,24 @@ function collectShellOccurrenceContexts(
}
continue
}
const arithmeticExpansion =

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: [[ ... -eq ... ]] is an arithmetic context, but this detector only recognizes delimiter-based forms, so placeholders there are compiled instead of rejected. Detect arithmetic operands in [[ ]] and the other shell arithmetic contexts before resolving the placeholder; otherwise values containing $(...) are re-evaluated by Bash.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/lib/execution/code-placeholders/shell.ts, line 541:

<comment>`[[ ... -eq ... ]]` is an arithmetic context, but this detector only recognizes delimiter-based forms, so placeholders there are compiled instead of rejected. Detect arithmetic operands in `[[ ]]` and the other shell arithmetic contexts before resolving the placeholder; otherwise values containing `$(...)` are re-evaluated by Bash.</comment>

<file context>
@@ -533,6 +538,24 @@ function collectShellOccurrenceContexts(
       }
       continue
     }
+    const arithmeticExpansion =
+      character === '$' &&
+      ((code[index + 1] === '(' && code[index + 2] === '(') || code[index + 1] === '[')
</file context>
Fix with cubic

) : undefined
}
>
{projects.error && (

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: An initial project-list failure leaves this editor without an in-place retry after React Query exhausts its retry. Render SettingsQueryErrorState with projects.refetch so admins can recover without reloading the settings page.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/ee/access-control/components/project-issue-restrictions.tsx, line 50:

<comment>An initial project-list failure leaves this editor without an in-place retry after React Query exhausts its retry. Render `SettingsQueryErrorState` with `projects.refetch` so admins can recover without reloading the settings page.</comment>

<file context>
@@ -0,0 +1,81 @@
+        ) : undefined
+      }
+    >
+      {projects.error && (
+        <p className='pl-2 text-[var(--text-error)] text-caption'>{projects.error.message}</p>
+      )}
</file context>
Fix with cubic

return null
}
const choices = projects.data?.pages.flatMap((page) => page.projects) ?? []
if (!projects.error && choices.length === 0) return null

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This guard hides configured restrictions after their project is archived, leaving admins no way to see or remove the stale ID. Preserve selected project metadata in the UI or clean the ID when archiving the project.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/ee/access-control/components/project-issue-restrictions.tsx, line 28:

<comment>This guard hides configured restrictions after their project is archived, leaving admins no way to see or remove the stale ID. Preserve selected project metadata in the UI or clean the ID when archiving the project.</comment>

<file context>
@@ -0,0 +1,81 @@
+    return null
+  }
+  const choices = projects.data?.pages.flatMap((page) => page.projects) ?? []
+  if (!projects.error && choices.length === 0) return null
+  const selected = new Set(value)
+  return (
</file context>
Fix with cubic

decidedAt: claimed.permissionDecidedAt?.toISOString(),
})
// A bound device lists the call for approval; the answer turns it into a call or drops it.
if (run.desktopDeviceId) ringDesktopInbox(run.desktopDeviceId, 'approval')

@cubic-dev-ai cubic-dev-ai Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: desktopDeviceId identifies the whole turn, not the approved call, so this also rings the executor for non-desktop approvals such as run_workflow. The inbox filters those rows out, but every answer still causes an unnecessary authenticated inbox read/reconcile; gate the ring on isDesktopToolCall(claimed.toolName, existing.args) after normalizing the stored args.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At apps/sim/app/api/copilot/tool-permission/route.ts, line 141:

<comment>`desktopDeviceId` identifies the whole turn, not the approved call, so this also rings the executor for non-desktop approvals such as `run_workflow`. The inbox filters those rows out, but every answer still causes an unnecessary authenticated inbox read/reconcile; gate the ring on `isDesktopToolCall(claimed.toolName, existing.args)` after normalizing the stored args.</comment>

<file context>
@@ -136,6 +137,8 @@ async function applyDecision(
     decidedAt: claimed.permissionDecidedAt?.toISOString(),
   })
+  // A bound device lists the call for approval; the answer turns it into a call or drops it.
+  if (run.desktopDeviceId) ringDesktopInbox(run.desktopDeviceId, 'approval')
 
   return { toolCallId, decision, applied: true }
</file context>
Fix with cubic

Comment thread apps/desktop/src/main/terminal/session.ts
Comment thread apps/sim/lib/core/outbox/processor.ts
* fix(desktop): stop a run the model never got the result of

* fix(desktop): never stop a run the model may hold

* fix(desktop): only a recovered result leaves a duplicate open

* fix(desktop): keep a recovered result's mark while it is parked
…fort across a failed first send (#8754)

* fix(mothership): close a pre-aborted SSE stream, keep the new-chat effort across a failed first send

- createSSEStream closes at once when the request aborted before start(), instead of subscribing until rotation.
- The new-chat effort pick is dropped by useChat when its chatless surface is left, not by each composer's unmount, so a failed first send keeps it and the pending chat view shows it.
- The chat response's effort is optional, so a new client loads chats from a server that predates it.

* fix(mothership): drop the new-chat effort when the surface adopts a chat

A first send stopped before admission adopted its chat without moving the pick, so the next new chat on the same Home mount showed and sent it. adoptResolvedChatId now drops the pick when the surface leaves the new chat. The rollback restore is gone: nothing clears the pick while a send is pending, and it overwrote a pick made during the send.
…the none effort (#8752)

* improvement(mothership): relabel the Sol pick GPT-6.1 Sol and retire the none effort

* improvement(mothership): check the effort table against the outgoing request

This branch was previously deployed

1 inactive deployment
Preview — 14e6c0ad Deployed Oct 7, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants