Skip to content

feat(tests): workflow tests as a workspace resource - #8773

Merged
TheodoreSpeaks merged 18 commits into
stagingfrom
feat/tests
Oct 9, 2026
Merged

TheodoreSpeaks merged 18 commits into
stagingfrom
feat/tests

Conversation

@TheodoreSpeaks

@TheodoreSpeaks TheodoreSpeaks commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Workflow tests as a workspace resource. A test is a plain vitest file (tests/<name>.test.js, a workspace file with context='test'). Mothership writes and runs these files through a new tests tool. They run in the sandbox against draft or deployed workflows, and the result is only pass or fail.

  • Harness (sim:test): runWorkflow(name, input, { trigger }), mockBlock, mockTool (answers an Agent's tool calls while the model still runs), .mockSampleOutput(overrides) (builds mock data shaped like the real output), spyOnBlock, and toMatchRubric (an LLM judge).
  • Executor test hooks: mocks and spies are matched by block name inside each workflow, and they carry into child workflows.
  • Page: the latest run (or one picked from the run dropdown), with Passing / Failing / Duration, live progress while running, every case, and what the run executed. A row there opens the deployed version or the snapshot of the draft as it ran. A run is marked out of date when the test file or a workflow it ran has changed since. The page has Edit, Split and Preview modes over the file.
  • Owned test files never open as file tabs in chat. The check uses the file's context, not its folder.
  • Rollout: behind the workflow-tests feature flag, off by default.
  • Migration: 0400_workflow_tests is additive only (new tables).
  • Also includes fix(chat): the public v2 chat API now sends entitlements and mode: 'agent'.

Type of Change

  • New feature

Testing

  • Ran Mothership end to end locally against workflows imported from staging (Infra Investigation Store, Infra Incident Core). It wrote and ran suites of 25/25 and 15/15 cases, with tool mocks and child workflow mocks.
  • Checked the page in the browser: the run picker, live progress, the out-of-date marking, and the draft snapshot opened from a ran-against row.
  • Ran bun run lint:check.
  • There are also unit tests for the executor test hooks.

Checklist

  • Code follows project style guidelines
  • Self-reviewed my changes
  • Tests added/updated and passing
  • No new warnings introduced
  • I confirm that I have read and agree to the terms outlined in the Contributor License Agreement (CLA)

Companion PRs

  • simstudioai/mothership#644

🤖 Generated with Claude Code

https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

TheodoreSpeaks and others added 3 commits October 7, 2026 17:04
`/api/v2/chat` (what `sim chat` uses) stopped sending `entitlements` and `mode`
in #8208, so CLI and API chats got no entitlement-gated tools and every CLI
service refused them with "CLI services require agent mode". The route now
computes entitlements per turn like the workspace chat and sends
`mode: 'agent'`. Its test still mocked the old entitlements function and
listed `mode` as a forbidden legacy field; both are updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
Workflow tests are a workspace resource whose source is a plain vitest file,
`tests/<name>.test.js`, owned by the test (workspace_files.context = 'test').
Sim creates a test's metadata with the tests tool and writes its cases with
the file tools; every write is collected in the sandbox and refused if the
file does not load.

- Runner: test files run in the isolated-vm sandbox against draft or
  deployed workflows. `runWorkflow` executes real runs; `mockBlock`,
  `mockTool` (Agent tool calls) and `spyOnBlock` reach blocks in the tested
  workflow and in every child workflow it runs, matched by name as each
  workflow starts. `.mockSampleOutput()` builds outputs shaped like the real
  block or tool. `toMatchRubric` asks a model judge for pass or fail.
- Runs record live per-case progress, the source hash, and the deployment of
  every workflow they ran, so results show as out of date once the test or
  a workflow changes.
- UI: Tests page and test page (Edit / Split / Preview over the file, the
  preview a dashboard of the selected run), a test resource type in chat,
  and a Tests sidebar entry behind the `workflow-tests` flag.
- Owned files never open as file tabs in chat: only workspace files and
  chat uploads do.
- Migration 0400 adds workflow_test and workflow_test_run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
…ainst

The test page shows one run at a time, chosen from a run picker with status
dots and Draft / Out of date chips. Case statuses use the Badge status chip.
Each ran-against entry records one execution, so a draft row opens the
workflow snapshot from that run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 8, 2026 11:12pm UTC

Request Review

@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

# Conflicts:
#	apps/sim/app/workspace/[workspaceId]/layout.tsx

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/workflow-tests/judge.ts
Comment thread apps/sim/providers/runtime-context.ts Outdated
Comment thread apps/sim/sandbox-tasks/workflow-test.ts
Comment thread apps/sim/lib/workflow-tests/application/tests.ts Outdated
Comment thread apps/sim/lib/workspace-files/application/workspace-file-context.ts Outdated
Comment thread apps/sim/lib/workflows/application/run-workflow-for-test.ts
Comment thread apps/sim/lib/mothership/agent-cli/resource-effects.ts
Comment thread apps/sim/lib/api/contracts/mothership-tests.ts Outdated
Comment thread apps/sim/providers/runtime-context.ts
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium impact] The latest change appears safe to merge and prevents unknown selected cases from producing a passing run.

Summary

Adds workflow tests as workspace resources, with sandbox runs, mocks, spies, saved results, and a feature-gated page.

  • The latest change rejects unknown only names before any cases run. The server saves the run as an error instead of a passing result.
  • No new actionable issues were found.
  • TheodoreSpeaks accepted the remaining long-queue risk and deferred it: queued files can appear stopped after more than 30 minutes of combined waiting. A follow-up will add a queued state or bounded parallel runs.
  • The earlier schema-example finding was withdrawn because the examples are required fixtures for existing coverage.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Collect cases from saved source] --> B[Read selected case names]
  B --> C{Every selected name exists?}
  C -->|Yes| D[Run selected cases]
  D --> E[Save case results]
  C -->|No| F[Throw with unknown names]
  F --> G[Save run as error]
Loading

Reviews (7) · Last reviewed commit: "fix(tests): fail a run whose selected ca..." · Reviewed by Greptile

Comment thread apps/sim/providers/runtime-context.ts Outdated
Comment thread apps/sim/executor/handlers/workflow/workflow-handler.ts Outdated
Comment thread apps/sim/lib/workspace-files/application/update-workspace-file-content.ts Outdated
Comment thread apps/sim/sandbox-tasks/workflow-test.ts Outdated
Comment thread apps/sim/sandbox-tasks/workflow-test.ts
Comment thread apps/sim/lib/workflow-tests/application/run-tests.ts Outdated
Comment thread apps/sim/lib/workflow-tests/application/tests.ts Outdated
Comment thread apps/sim/lib/execution/sandbox/bundles/build.ts
Comment thread apps/sim/app/workspace/[workspaceId]/tests/tests.tsx
Comment thread apps/sim/stores/workflow-tests/store.ts
- Narrow the test principal to the kinds workflow_tests.run admits before
  handing it to executeWorkflow.
- Select progress with the latest-run rows, guard file upsert ids in the
  tab filter, and set the sandbox Event polyfills through Reflect.
- Cover the tests tool in the management tool contract, expect content
  writes to reach test files, and stub test availability in the payload test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
- Redact each run's resolved secrets from what returns to the sandbox
  (output, errors, mocked tool inputs); mocked tools get only declared params.
- Custom blocks no longer receive the consumer's test hooks.
- Draft runs go stale when the draft changes; children a run calls are
  recorded in ran-against.
- Test cases commit in the same transaction as the source file write.
- Harness: runWorkflow is rejected in suite hooks, a timed-out case stops
  the file, and expect.assertions/hasAssertions are supported.
- Insert run rows in one statement and start each run's clock with its file;
  check bans before each workflow run; restrict owned-file access to Copilot
  delegation; validate names in the tool contract.
- Delete soft-deletes the test file and removes the chat tab; a finished run
  shows its own cases; polling at 3s on a separate read bucket; list error
  state; store reset; tests stay in the org Add Resource picker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

Comment thread apps/sim/providers/runtime-context.ts Outdated
Comment thread apps/sim/sandbox-tasks/workflow-test.ts
Comment thread apps/sim/lib/workflow-tests/session.ts
Comment thread apps/sim/lib/workflows/application/run-workflow-for-test.ts
Comment thread apps/sim/lib/api/contracts/mothership-management-tools.test.ts
Comment thread apps/sim/hooks/queries/workflow-tests.ts Outdated
Comment thread apps/sim/app/workspace/[workspaceId]/tests/components/test-dashboard.tsx Outdated
# Conflicts:
#	apps/sim/app/workspace/[workspaceId]/layout.tsx
#	apps/sim/lib/workflows/executor/execution-core.ts
#	packages/db/migrations/meta/0400_snapshot.json
#	packages/db/migrations/meta/_journal.json
… tool args

- Each test workflow run reserves and releases an execution slot.
- A case waits for assertions it did not await and fails if one fails.
- Mocked MCP and custom tools keep the arguments their schema declares.
- A closed session refuses starts still awaiting their lookups.
- Stable refresh callback; scroll fade on the results pane.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

Comment thread apps/sim/lib/mothership/request/tools/resources.ts
Comment thread apps/sim/lib/mothership/tools/server/tests.ts
Comment thread apps/sim/lib/mothership/generated/resources.ts
…eopen tests

- File-edit tool results mark a non-tab file `fileTab: false`, and the browser
  skips promoting it.
- Idle test pages poll every 15s so runs started elsewhere appear; a Mothership
  run returns its tests as resource changes.
- open_resource accepts test resources through an authorized read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

Comment thread apps/sim/lib/mothership/resources/file-tabs.ts Outdated
Comment thread apps/sim/lib/mothership/tools/server/tests.ts
- A finished Mothership run refreshes its tests instead of upserting tabs, so a
  test deleted mid-run does not come back.
- A failed file-tab lookup after a saved edit opens no tab instead of
  reporting the edit as failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
The vitest-expect sandbox bundle builds from these packages; rebuilt with the
Reflect-based event polyfills.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
…itest/spy

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
# Conflicts:
#	apps/sim/app/o/[organizationId]/home/components/composer/composer.test.tsx
#	apps/sim/app/o/[organizationId]/layout.tsx
#	apps/sim/app/workspace/[workspaceId]/layout.tsx
#	apps/sim/app/workspace/[workspaceId]/providers/feature-flags-provider.tsx
MCP tool ids embed the server's database id, which changes when a server is
re-added or a workspace is forked, so a stale mock silently stopped matching
and the real server was called. Tests now name workspace tools the way the
workspace does: mockTool({ mcp: 'Server', tool: 'name' }) resolved per run
(failing on an unknown or ambiguous server), and mockTool({ customTool:
'Title' }) matched case- and space-insensitively. Raw mcp- and custom_ ids
are rejected; built-in catalog ids are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

# Conflicts:
#	packages/db/migrations/meta/0401_snapshot.json
#	packages/db/migrations/meta/_journal.json
Comment thread apps/sim/lib/api/contracts/mothership-tests.ts
Comment thread apps/sim/app/workspace/[workspaceId]/tests/components/test-detail.tsx Outdated
Comment thread apps/sim/lib/workflow-tests/session.ts
…judges on close

- A test named "run" collided with the static run endpoint, so its detail
  page got a 405; the name is now reserved.
- Run is disabled while the open editor holds edits the server has not
  saved (including a refused save), so a run never uses the previous source.
- toMatchRubric model calls are aborted when the sandbox run ends, so a
  stopped test no longer keeps calling or billing the judge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

Comment thread apps/sim/sandbox-tasks/workflow-test.ts
A renamed or misspelled `only` path skipped every case, and the run was then
saved as passing. The harness now rejects unknown names, so the run is
recorded as an error with the names it could not find.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
@TheodoreSpeaks

Copy link
Copy Markdown
Collaborator Author

@greptile review

@github-actions github-actions Bot added the requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep label Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

⚠️ Cross-repo companion check

One or more companion PRs aren't merged into staging yet. Merging this without them will leave copilot and sim out of sync — merge them in lockstep.

  • ❌ simstudioai/mothership#644 — OPEN, not merged (targets staging) — feat(worker): write and run workflow tests behind the tests entitlement

@TheodoreSpeaks
TheodoreSpeaks merged commit 900afc3 into staging Oct 9, 2026
50 checks passed
@TheodoreSpeaks
TheodoreSpeaks deleted the feat/tests branch October 9, 2026 02:20

This branch was previously deployed

1 inactive deployment
Preview — 21cc7452 Deployed Oct 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant