Skip to content

fix(mothership): run same-workflow workflow tools in order instead of failing as busy - #8709

Merged
waleedlatif1 merged 6 commits into
stagingfrom
fix/serialize-workflow-tool-runs
Oct 7, 2026
Merged

waleedlatif1 merged 6 commits into
stagingfrom
fix/serialize-workflow-tool-runs

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

When the agent issues several workflow run tools (run_block, run_workflow, run_from_block, run_workflow_until_block) against the same workflow in one step, the browser now runs them one after another, in the order they arrived. Before, every call after the first failed with WORKFLOW_EXECUTION_BUSY, and the agent had to spend a round re-issuing it.

Why

WORKFLOW_EXECUTION_BUSY is raised only in the browser (run-tool-execution.ts). The editor keeps one visible execution per workflow: the running flag, the current execution id, active blocks and the terminal pointer. Two runs of one workflow in a tab would overwrite each other, so the lock is right. Failing the queued calls is not: they are the agent's own fan-out, and nearly all of the busy results were this case, waiting behind a run tool from the same chat that finishes within seconds.

Design

  • One run slot per workflow. Each workflow has a single record, { owner, waiters }. The owner is either a run tool or an interrupted execution. acquireRunToolSlot, admitNextRunTool, releaseRunToolSlot and releaseInterruptedHold are the only functions that change it, and every run tool releases through releaseRunToolSlot.
  • FIFO hand-off. A release passes the slot to the oldest waiter synchronously, so a call issued later can't take it first.
  • Stop. stopRunToolExecutions removes stopped waiters and reports them cancelled without starting them. A call the server already took over is cancelled by the same Stop: the server's run uses the turn's abort signal.
  • Chat recovery. bindRunToolToExecution treats a waiting call as owned by this tab, so recovery never starts it twice. It also re-sends a stored, undelivered completion even while another run owns the workflow, using the report's own execution id.
  • Dropped streams. If a run's stream is interrupted while the server keeps executing it, that execution keeps the slot (owner interrupted), so a waiting call can't overwrite its terminal pointer. This is not run-tool ownership, so the editor's reconnect can still re-attach.
    • The hold is released once two things are true: the server reports the execution settled (or 403/404), and no reconnect stream is still following it.
    • Status is checked only while a call is actually waiting, never while a reconnect stream is open or the tab is hidden, and with backoff (max 15s).
    • Once the server reports the execution settled, the browser stops asking and only waits for the reconnect to close. After repeated status errors it gives up and releases the hold.
    • On release it clears any visible state that an abandoned reconnect left behind.
  • Paused runs count as settled. A pause ends the stream and can last for days.
  • What still returns BUSY: a run no run tool drives, such as the user's own run. That behaviour is unchanged.
  • No queue timeout of its own: a call no browser claims within the server's pickup grace is run by the server. A late launch from the queue then gets the execute route's existing benign 409, so a workflow never runs twice and a long wait can't stall the agent.

Tests

run-tool-execution.test.ts covers:

  • in-order runs;
  • no skipping ahead during a hand-off;
  • a stopped waiter is cancelled and never started;
  • recovery doesn't double-start a waiter;
  • a different workflow isn't held up;
  • a dropped run holds the workflow until it settles;
  • status is checked only while a call waits and no reconnect is open;
  • an abandoned reconnect doesn't block the queue;
  • a stored completion is re-sent while another run owns the workflow.

Each regression test fails with its guard removed. apps/sim type-check, biome, check:audits, and the run-tool, editor-hook and home-hook suites pass.

Verify

In a browser chat, ask the agent to test several blocks of one workflow in parallel. The runs should appear one after another in the editor, every result should succeed with no busy errors, and Stop during the queue should cancel the remaining calls.

… failing as busy

When the agent fans out run_block/run_workflow/run_from_block/run_workflow_until_block
calls on one workflow in a single step, the browser executor rejected every call after
the first with WORKFLOW_EXECUTION_BUSY, because the editor keeps one visible execution
per workflow. The agent then spent a round re-issuing them one by one. That was about
46% of run_block failures.

Calls behind a run tool already running that workflow in this tab now wait in arrival
order and launch when it releases. Stop cancels waiting calls without launching them,
a chat recovery treats a waiting call as owned by this tab, and a run the agent did not
start (the user's manual run) is still reported as busy. The wait needs no timeout of
its own: a call nobody claims within the server's pickup grace is run by the server,
and the late launch here gets the execute route's benign 409.
@vercel

vercel Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 7, 2026 1:59am UTC

Request Review

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 2 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts
@greptile-apps

greptile-apps Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Changes how workflow tool execution queuing works.

The PR appears safe to merge; no blocking issues remain.

What we checked:

  • Saved pointer cleanup finishes: Both pointer helpers use synchronous sessionStorage calls and catch storage errors. The new waits do not depend on a network request or another long-running operation.

Summary

Runs same-workflow tools in arrival order instead of rejecting later calls as busy.

  • Interrupted runs keep the workflow available for reconnect until they settle.
  • The latest change clears the abandoned run’s saved pointer before admitting the next call.
  • The regression test now checks running state and pointer cleanup order.
  • All three earlier findings are addressed. No new actionable issues were found.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Run tool arrives] --> B{Workflow slot held?}
  B -->|No| C[Start run]
  B -->|Yes| D[Wait in arrival order]
  C --> E{Stream interrupted?}
  E -->|No| F[Release slot]
  E -->|Yes| G[Keep interrupted execution in slot]
  G --> H[Wait until settled and reconnect closed]
  H --> I[Clear old visible state and saved pointer]
  I --> F
  F --> J[Admit oldest waiting call]
  J --> C
Loading

Reviews (6) · Last reviewed commit: "fix(mothership): clear a dropped run's s..." · Reviewed by Greptile

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts Outdated
… it settles

A run tool whose execute stream is interrupted leaves the run executing server-side and
keeps its terminal pointer so the editor can re-attach. Admitting the next waiting call
right away overwrote that pointer and claimed the workflow, so reconnect skipped it and
the still-running execution lost its live output and the editor's Stop.

The interrupted execution now holds the workflow until the server reports it settled
and any reconnect has drained it. The hold is not run-tool ownership, so reconnect can
still claim the pointer. Waiting calls start after it; one still waiting past the
server's pickup grace is run by the server, and its late start here gets the benign 409.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 2 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts
Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts Outdated
…fter an abandoned reconnect

The hold waited for the execution store to stop showing the interrupted execution as
current. Navigating away cancels the editor's reconnect without clearing that, so a
settled run kept later calls waiting until the server's pickup fallback ran them.

The hold now waits for the server to settle the execution and for no reconnect stream to
be open for it, then clears the stale visible state it left behind before admitting the
next call.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts Outdated
Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts Outdated
Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.test.ts Outdated
…e another run owns the workflow

A run whose completion report failed keeps it in session storage for recovery to re-send.
Recovery refused to touch the call whenever another run tool owned the workflow, which the
queue now makes common, so the stored report was never delivered. Re-sending a stored
report needs nothing from the workflow; recovery now does it first, and only skips the
pointer lookup and cleanup that belong to the run that owns the workflow.

Also makes the abandoned-reconnect test track the running flag, so it fails without the
stale-state clearing it covers.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 3 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Turn on auto-fix | Re-trigger cubic

…cord and poll a dropped run only when needed

Replaces the three parallel maps (owning run tool, waiters, interrupted execution) with one
per-workflow record whose owner is a run tool or an interrupted execution. Acquire, admit
and release are its only mutators, and every run tool releases through one path.

A dropped run is now checked only while a call is actually waiting behind it, never while
the editor's reconnect stream is following it or the tab is hidden, stops fetching once the
server reports it settled, and gives up the hold after repeated status failures. Also
extracts the visible-execution reset, makes waiter removal a single pass and moves the
design notes to the PR description.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts
… released

A reconnect cancelled by navigating away leaves the settled run's execution pointer saved.
Releasing the hold now clears it when it still names that execution, before the next run
is admitted, so recovery cannot later report the finished run as a live one.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/run-tool-execution.ts
@waleedlatif1
waleedlatif1 merged commit b5a6d53 into staging Oct 7, 2026
33 checks passed
@waleedlatif1
waleedlatif1 deleted the fix/serialize-workflow-tool-runs branch October 7, 2026 02:17

This branch was previously deployed

1 inactive deployment
Preview — e65f6c75 Deployed Oct 7, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant