Repository navigation
fix(mothership): run same-workflow workflow tools in order instead of failing as busy - #8709
Conversation
… failing as busy When the agent fans out run_block/run_workflow/run_from_block/run_workflow_until_block calls on one workflow in a single step, the browser executor rejected every call after the first with WORKFLOW_EXECUTION_BUSY, because the editor keeps one visible execution per workflow. The agent then spent a round re-issuing them one by one. That was about 46% of run_block failures. Calls behind a run tool already running that workflow in this tab now wait in arrival order and launch when it releases. Stop cancels waiting calls without launching them, a chat recovery treats a waiting call as owned by this tab, and a run the agent did not start (the user's manual run) is still reported as busy. The wait needs no timeout of its own: a call nobody claims within the server's pickup grace is run by the server, and the late launch here gets the execute route's benign 409.
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
There was a problem hiding this comment.
All reported issues were addressed across 2 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Turn on auto-fix | Re-trigger cubic
|
… it settles A run tool whose execute stream is interrupted leaves the run executing server-side and keeps its terminal pointer so the editor can re-attach. Admitting the next waiting call right away overwrote that pointer and claimed the workflow, so reconnect skipped it and the still-running execution lost its live output and the editor's Stop. The interrupted execution now holds the workflow until the server reports it settled and any reconnect has drained it. The hold is not run-tool ownership, so reconnect can still claim the pointer. Waiting calls start after it; one still waiting past the server's pickup grace is run by the server, and its late start here gets the benign 409.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 2 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Turn on auto-fix | Re-trigger cubic
…fter an abandoned reconnect The hold waited for the execution store to stop showing the interrupted execution as current. Navigating away cancels the editor's reconnect without clearing that, so a settled run kept later calls waiting until the server's pickup fallback ran them. The hold now waits for the server to settle the execution and for no reconnect stream to be open for it, then clears the stale visible state it left behind before admitting the next call.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Turn on auto-fix | Re-trigger cubic
…e another run owns the workflow A run whose completion report failed keeps it in session storage for recovery to re-send. Recovery refused to touch the call whenever another run tool owned the workflow, which the queue now makes common, so the stored report was never delivered. Re-sending a stored report needs nothing from the workflow; recovery now does it first, and only skips the pointer lookup and cleanup that belong to the run that owns the workflow. Also makes the abandoned-reconnect test track the running flag, so it fails without the stale-state clearing it covers.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
…cord and poll a dropped run only when needed Replaces the three parallel maps (owning run tool, waiters, interrupted execution) with one per-workflow record whose owner is a run tool or an interrupted execution. Acquire, admit and release are its only mutators, and every run tool releases through one path. A dropped run is now checked only while a call is actually waiting behind it, never while the editor's reconnect stream is following it or the tab is hidden, stops fetching once the server reports it settled, and gives up the hold after repeated status failures. Also extracts the visible-execution reset, makes waiter removal a single pass and moves the design notes to the PR description.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Turn on auto-fix | Re-trigger cubic
… released A reconnect cancelled by navigating away leaves the settled run's execution pointer saved. Releasing the hold now clears it when it still names that execution, before the next run is admitted, so recovery cannot later report the finished run as a live one.
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Turn on auto-fix | Re-trigger cubic
Summary
When the agent issues several workflow run tools (
run_block,run_workflow,run_from_block,run_workflow_until_block) against the same workflow in one step, the browser now runs them one after another, in the order they arrived. Before, every call after the first failed withWORKFLOW_EXECUTION_BUSY, and the agent had to spend a round re-issuing it.Why
WORKFLOW_EXECUTION_BUSYis raised only in the browser (run-tool-execution.ts). The editor keeps one visible execution per workflow: the running flag, the current execution id, active blocks and the terminal pointer. Two runs of one workflow in a tab would overwrite each other, so the lock is right. Failing the queued calls is not: they are the agent's own fan-out, and nearly all of the busy results were this case, waiting behind a run tool from the same chat that finishes within seconds.Design
{ owner, waiters }. The owner is either a run tool or an interrupted execution.acquireRunToolSlot,admitNextRunTool,releaseRunToolSlotandreleaseInterruptedHoldare the only functions that change it, and every run tool releases throughreleaseRunToolSlot.stopRunToolExecutionsremoves stopped waiters and reports themcancelledwithout starting them. A call the server already took over is cancelled by the same Stop: the server's run uses the turn's abort signal.bindRunToolToExecutiontreats a waiting call as owned by this tab, so recovery never starts it twice. It also re-sends a stored, undelivered completion even while another run owns the workflow, using the report's own execution id.interrupted), so a waiting call can't overwrite its terminal pointer. This is not run-tool ownership, so the editor's reconnect can still re-attach.Tests
run-tool-execution.test.tscovers:Each regression test fails with its guard removed.
apps/simtype-check, biome,check:audits, and the run-tool, editor-hook and home-hook suites pass.Verify
In a browser chat, ask the agent to test several blocks of one workflow in parallel. The runs should appear one after another in the editor, every result should succeed with no busy errors, and Stop during the queue should cancel the remaining calls.