Repository navigation
[Fix] Controller shutdown hangs when worker launches stall - #3445
Merged
Merged
Conversation
Contributor
|
No code issues found. See task Reviewed the shared shutdown deadline, timer cleanup, late-result dispatch guards, spawn callers, teardown and persisted orphan recovery. Current-head test, lint, type-check and Knip checks passed; Docker build and JavaScript/TypeScript CodeQL were still running when checked. Reviewed a6944e4 |
This was referenced Oct 9, 2026
roomote-roomote
Bot
requested review from
brunobergher,
daniel-lxs and
mrubens
as code owners
October 9, 2026 06:38
daniel-lxs
approved these changes
Oct 9, 2026
daniel-lxs
left a comment
Member
There was a problem hiding this comment.
Verified stalled-launch shutdown through real controller processes, queue and database state, plus concurrent durable orphan recovery with a controlled provider. Focused lifecycle tests, controller checks/build and current CI pass at this head.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Controller shutdown now spends one existing wait budget across the current dequeue iteration and in-flight worker launches. A stalled provider launch no longer prevents watcher cleanup and teardown from starting. Deadline warnings identify unfinished run IDs, normal completion clears its timers, and the progress interval cannot keep the process alive by itself.
Late queue, database-claim and provider-resolution results do not admit new provider work after shutdown begins. Their durable Pending/Dequeued rows remain eligible for the existing recovery path. Skipped dispatch is not logged as a successful worker launch.
Why this change was made
The current iteration wait was bounded, but the following
Promise.allSettledover provider launches was not. Since SIGTERM/SIGINT handling awaitsstop(), a never-settling provider promise could hang graceful shutdown indefinitely. Bounding the second wait also requires guarding late results so cleanup cannot be followed by a fresh dispatch.Impact
The existing 60-second budget limits the waits before controller teardown; no new timeout setting is added. Task runs are not failed, canceled or completed because shutdown times out. Provider operations already in progress are not canceled or remotely fenced, and may settle while the process is alive. Teardown itself retains its existing behavior and is not given a new deadline. The controller entrypoint exits after stop/flush; persisted recovery leases remain authoritative.
How it was tested
35201fbf: actual controller start/dequeue with real Postgres and Redis, controlled queue/provider/auth/monitoring. With a 50ms test budget, stop remained pending at 125ms, teardown had not run, and the run/event were durably Dequeued/nonterminal; releasing the provider allowed stop to finish.a6944e453e1ff32e58d678c617e592345af268ec: repeated independent-process proof on Node 24.13.1. A hung-provider child naturally exited without force after stop (66ms including test client teardown with a 50ms wait budget). Parent reads confirmed Dequeued state, an actual dequeue event and no terminal/cancellation writes. Two separate recovery processes concurrently reclaimed exactly once and renewed the persisted lease.Checklist
[Fix],[Feat],[Improve],[Refactor],[Docs], or[Chore]followed by a user-facing descriptionpnpm lintandpnpm check-typespass locally — normal fast repository gates passed; full formatting-inclusive commands were not runpnpm changesetRelated PRs
These are independent reliability fixes from the same shift; none depends on another.