Repository navigation
feat(tests): workflow tests as a workspace resource - #8773
Conversation
`/api/v2/chat` (what `sim chat` uses) stopped sending `entitlements` and `mode` in #8208, so CLI and API chats got no entitlement-gated tools and every CLI service refused them with "CLI services require agent mode". The route now computes entitlements per turn like the workspace chat and sends `mode: 'agent'`. Its test still mocked the old entitlements function and listed `mode` as a forbidden legacy field; both are updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
Workflow tests are a workspace resource whose source is a plain vitest file, `tests/<name>.test.js`, owned by the test (workspace_files.context = 'test'). Sim creates a test's metadata with the tests tool and writes its cases with the file tools; every write is collected in the sandbox and refused if the file does not load. - Runner: test files run in the isolated-vm sandbox against draft or deployed workflows. `runWorkflow` executes real runs; `mockBlock`, `mockTool` (Agent tool calls) and `spyOnBlock` reach blocks in the tested workflow and in every child workflow it runs, matched by name as each workflow starts. `.mockSampleOutput()` builds outputs shaped like the real block or tool. `toMatchRubric` asks a model judge for pass or fail. - Runs record live per-case progress, the source hash, and the deployment of every workflow they ran, so results show as out of date once the test or a workflow changes. - UI: Tests page and test page (Edit / Split / Preview over the file, the preview a dashboard of the selected run), a test resource type in chat, and a Tests sidebar entry behind the `workflow-tests` flag. - Owned files never open as file tabs in chat: only workspace files and chat uploads do. - Migration 0400 adds workflow_test and workflow_test_run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
…ainst The test page shows one run at a time, chosen from a run picker with status dots and Draft / Out of date chips. Case statuses use the Badge status chip. Each ran-against entry records one execution, so a draft row opens the workflow snapshot from that run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
@greptile review |
# Conflicts: # apps/sim/app/workspace/[workspaceId]/layout.tsx
There was a problem hiding this comment.
All reported issues were addressed
Reply with feedback, questions, or to request a fix.
Turn on auto-fix | Re-trigger cubic
|
- Narrow the test principal to the kinds workflow_tests.run admits before handing it to executeWorkflow. - Select progress with the latest-run rows, guard file upsert ids in the tab filter, and set the sandbox Event polyfills through Reflect. - Cover the tests tool in the management tool contract, expect content writes to reach test files, and stub test availability in the payload test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
- Redact each run's resolved secrets from what returns to the sandbox (output, errors, mocked tool inputs); mocked tools get only declared params. - Custom blocks no longer receive the consumer's test hooks. - Draft runs go stale when the draft changes; children a run calls are recorded in ran-against. - Test cases commit in the same transaction as the source file write. - Harness: runWorkflow is rejected in suite hooks, a timed-out case stops the file, and expect.assertions/hasAssertions are supported. - Insert run rows in one statement and start each run's clock with its file; check bans before each workflow run; restrict owned-file access to Copilot delegation; validate names in the tool contract. - Delete soft-deletes the test file and removes the chat tab; a finished run shows its own cases; polling at 3s on a separate read bucket; list error state; store reset; tests stay in the org Add Resource picker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
# Conflicts: # apps/sim/app/workspace/[workspaceId]/layout.tsx # apps/sim/lib/workflows/executor/execution-core.ts # packages/db/migrations/meta/0400_snapshot.json # packages/db/migrations/meta/_journal.json
… tool args - Each test workflow run reserves and releases an execution slot. - A case waits for assertions it did not await and fails if one fails. - Mocked MCP and custom tools keep the arguments their schema declares. - A closed session refuses starts still awaiting their lookups. - Stable refresh callback; scroll fade on the results pane. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
…eopen tests - File-edit tool results mark a non-tab file `fileTab: false`, and the browser skips promoting it. - Idle test pages poll every 15s so runs started elsewhere appear; a Mothership run returns its tests as resource changes. - open_resource accepts test resources through an authorized read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
- A finished Mothership run refreshes its tests instead of upserting tabs, so a test deleted mid-run does not come back. - A failed file-tab lookup after a saved edit opens no tab instead of reporting the edit as failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
The vitest-expect sandbox bundle builds from these packages; rebuilt with the Reflect-based event polyfills. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
…itest/spy Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
# Conflicts: # apps/sim/app/o/[organizationId]/home/components/composer/composer.test.tsx # apps/sim/app/o/[organizationId]/layout.tsx # apps/sim/app/workspace/[workspaceId]/layout.tsx # apps/sim/app/workspace/[workspaceId]/providers/feature-flags-provider.tsx
MCP tool ids embed the server's database id, which changes when a server is
re-added or a workspace is forked, so a stale mock silently stopped matching
and the real server was called. Tests now name workspace tools the way the
workspace does: mockTool({ mcp: 'Server', tool: 'name' }) resolved per run
(failing on an unknown or ambiguous server), and mockTool({ customTool:
'Title' }) matched case- and space-insensitively. Raw mcp- and custom_ ids
are rejected; built-in catalog ids are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
# Conflicts: # packages/db/migrations/meta/0401_snapshot.json # packages/db/migrations/meta/_journal.json
…judges on close - A test named "run" collided with the static run endpoint, so its detail page got a 405; the name is now reserved. - Run is disabled while the open editor holds edits the server has not saved (including a refused save), so a run never uses the previous source. - toMatchRubric model calls are aborted when the sandbox run ends, so a stopped test no longer keeps calling or billing the judge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
A renamed or misspelled `only` path skipped every case, and the run was then saved as passing. The harness now rejects unknown names, so the run is recorded as an error with the names it could not find. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R
|
@greptile review |
|
Summary
Workflow tests as a workspace resource. A test is a plain vitest file (
tests/<name>.test.js, a workspace file withcontext='test'). Mothership writes and runs these files through a newteststool. They run in the sandbox against draft or deployed workflows, and the result is only pass or fail.sim:test):runWorkflow(name, input, { trigger }),mockBlock,mockTool(answers an Agent's tool calls while the model still runs),.mockSampleOutput(overrides)(builds mock data shaped like the real output),spyOnBlock, andtoMatchRubric(an LLM judge).context, not its folder.workflow-testsfeature flag, off by default.0400_workflow_testsis additive only (new tables).fix(chat): the public v2 chat API now sends entitlements andmode: 'agent'.Type of Change
Testing
bun run lint:check.Checklist
Companion PRs
🤖 Generated with Claude Code
https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R