fix(execution): complete background jobs on workflow failures, fault only on platform errors - #8368
waleedlatif1 wants to merge 3 commits into
Conversation
…only on platform errors Async API and resume jobs re-threw every execution error, so Trigger.dev marked the run failed and alerted on user workflow failures (a block's own error, a missing required field). The workflow-execution task also checked core's finalized signal before core had set it, so its existing guard never applied. - Add one classifier (lib/workflows/executor/job-failure) built on classifyExecutionError: a failure core recorded and attributed to a block is the workflow's outcome and the job completes with success: false; anything else (engine, setup, unrecorded, or a programming error in Sim code) faults. - Await core's post-execution work before classifying. - Carry programming errors thrown inside tools across the executeTool flattening as ToolResponse.isSystemError so they still fault the job. - Throw WorkflowValidationError with block attribution for missing required fields so pre-execution validation is attributed to its block. - Apply the rule to workflow-execution, resume-execution and schedule-execution; schedule re-throws platform faults only after its own failure bookkeeping. - Project a completed failure result back to failed in /api/jobs and the queue-job status fallback so pollers still see a failed run.
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
There was a problem hiding this comment.
All reported issues were addressed across 19 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
…t of system faults
- Add SystemError and isSystemError to @sim/utils/errors; throw SystemError
where an internal tool operation has no registered handler so a broken
registry still faults the job after executeTool flattens the error.
- isProgrammingError excludes errors carrying a cause, so fetch's
TypeError('fetch failed', { cause }) for an unreachable host is not treated
as a defect in Sim code.
- Track the schedule job fault in a box so a thrown undefined still faults.
- Drop the partial executor-errors mock; the real hasExecutionResult suffices.
|
@greptileai review — second commit (e3d6098) addresses the prior findings. |
|
@cubic-dev review |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
|
Closing in favor of a stricter design: page by default and only complete quietly for failures positively tagged as user-caused at their source. The block-attribution rule here could silence platform faults inside blocks (e.g. database errors, handler invariants). Replacement PR incoming. |
Summary
workflow-executionalso checked core's finalized signal before core set it, so its guard never appliedlib/workflows/executor/job-failure.ts, built on the existingclassifyExecutionError: a failure core recorded and attributed to a block is the workflow's outcome, and the job completes withsuccess: false. Everything else (engine, setup, unrecorded failures, programming errors in Sim code) still faults and alertsTypeError/ReferenceError/RangeErrorwith nocause) and a newSystemError(@sim/utils/errors) thrown at invariant sites (a declared internal operation with no registered handler) are carried across theexecuteToolflattening asToolResponse.isSystemError. Afetchnetwork failure (TypeError('fetch failed', { cause })) is not treated as a defectWorkflowValidationErrorwith block attribution, so it counts as the block's outcome (the editor console also names the block now)workflow-execution,resume-executionandschedule-execution. Schedule previously swallowed every error; it now re-throws platform faults only after its own failure bookkeeping (failure count, next run, claim)/api/jobsand the queue-job status fallback project a completed failure result back tofailed, so pollers still see a failed runwebhookIdempotency(retryFailures: true), so re-throwing there would let a provider redelivery re-run a partly executed workflow. Faulting webhook engine errors needs to happen outside the idempotency wrapper, so it's a follow-upType of Change
Testing
workflow-executioncompletes on a recorded block failure,resume-executioncompletes on a recorded block failure,schedule-executionfaults on an unrecorded failure after bookkeeping, and the serializer attributes missing required fields to their blockexecuteToolsetsisSystemErroronly for programming errorsChecklist