Skip to content

Commit dba34e9

Browse files
committed
Create owned conversations for every benchmark execution
1 parent b6dc150 commit dba34e9

7 files changed

Lines changed: 136 additions & 31 deletions

File tree

‎apps/sim/lib/benchmarks/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@ Open **Run history**, choose **View** on a result, then **Compare** on another t
2121

2222
Grading opens the saved report for human review. On any detail, choose **Mark as correct** or **Mark as incorrect**, add an optional note, and save. Human decisions update the report score and comparisons, while the AI grade, explanation, answers, and spec remain intact. **Edit review** changes the note; **Use AI grade** removes the override. The history list and report show the original AI score when human overrides exist. Reviews retain the same private ownership and current source-access checks as the benchmark; concurrent edits cannot silently overwrite each other.
2323

24-
Reference specs can be prose or JSON text; **Import spec** accepts Markdown, text, and JSON files without reformatting their contents. **Generate blanks with AI** creates the redacted reference and expected answers together. **Edit JSON** also lets you paste or edit a mapping such as `{"queue":"Customer Escalations","handoff":"Engineering explicitly accepts the case"}` alongside the matching redacted reference. Keys are blank IDs and values are exact original passages as strings. Applying validates that the mapping restores the reference exactly and updates the draft; **Save changes** persists it.
24+
Generated reference specs are readable Markdown, with code blocks where exact schemas or mappings help. Imported reference specs can be prose or JSON text; **Import spec** accepts Markdown, text, and JSON files without reformatting their contents. **Generate blanks with AI** creates the redacted reference and expected answers together. **Edit JSON** also lets you paste or edit a mapping such as `{"queue":"Customer Escalations","handoff":"Engineering explicitly accepts the case"}` alongside the matching redacted reference. Keys are blank IDs and values are exact original passages as strings. Applying validates that the mapping restores the reference exactly and updates the draft; **Save changes** persists it.
2525

2626
Benchmark stages have no benchmark-specific duration cutoff. The active process renews its persistence lease every 30 seconds; a two-minute lease detects an abandoned process rather than limiting a healthy run. Concurrent starts or stale completions cannot overwrite newer results. Interrupted steps are retryable; a hard server failure becomes retryable when its lease expires. This version runs a stage within its HTTP request, so the deployment must support long-lived requests. Page reload may interrupt a pending request; saved completed stages remain available.
2727

‎apps/sim/lib/benchmarks/application/operations.ts‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -3,8 +3,8 @@ import { defineWorkspaceOperation } from '@/lib/core/application/workspace-opera
33

44
export const benchmarkOperations = {
55
/** permission-group-exempt: platform superusers administer benchmarks; source access is checked for the selected user. */
6-
prepareReference: defineOperation({
7-
id: 'benchmarks.reference.prepare',
6+
prepareExecution: defineOperation({
7+
id: 'benchmarks.execution.prepare',
88
principalKinds: ['session'],
99
capability: 'none',
1010
}),

apps/sim/lib/benchmarks/application/prepare-reference.ts renamed to apps/sim/lib/benchmarks/application/prepare-execution.ts

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -6,9 +6,9 @@ import { requireBenchmarkCaseAccess } from '@/lib/benchmarks/application/cases'
66
import { benchmarkOperations } from '@/lib/benchmarks/application/operations'
77
import { MOTHERSHIP_CHAT_DEFAULT_MODEL } from '@/lib/mothership/constants'
88

9-
/** Source reads use an owned workspace chat, separate from the planner's enterprise discovery context. */
10-
export const prepareBenchmarkReference = defineAuthorizedBenchmarkUseCase({
11-
operation: benchmarkOperations.prepareReference,
9+
/** Each JSON execution uses a fresh owned workspace chat, separate from the planner's enterprise discovery context. */
10+
export const prepareBenchmarkExecution = defineAuthorizedBenchmarkUseCase({
11+
operation: benchmarkOperations.prepareExecution,
1212
async execute({
1313
principal,
1414
input,
@@ -32,7 +32,7 @@ export const prepareBenchmarkReference = defineAuthorizedBenchmarkUseCase({
3232
lastSeenAt: new Date(),
3333
})
3434
.returning({ id: copilotChats.id })
35-
if (!chat) throw new Error('Failed to create benchmark reference conversation')
35+
if (!chat) throw new Error('Failed to create benchmark execution conversation')
3636
return { chatId: chat.id, userId }
3737
},
3838
})

‎apps/sim/lib/benchmarks/application/run-stage.ts‎

Lines changed: 12 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -5,7 +5,6 @@ import { z } from 'zod'
55
import { defineAuthorizedBenchmarkUseCase } from '@/lib/benchmarks/application/access'
66
import { requireBenchmarkCaseAccess } from '@/lib/benchmarks/application/cases'
77
import { benchmarkOperations } from '@/lib/benchmarks/application/operations'
8-
import { prepareBenchmarkReference } from '@/lib/benchmarks/application/prepare-reference'
98
import {
109
BENCHMARK_LEASE_MS,
1110
withBenchmarkStageLease,
@@ -44,7 +43,14 @@ import { OrchestrationError } from '@/lib/core/orchestration/types'
4443
const logger = createLogger('BenchmarkStage')
4544

4645
const distillationSchema = z
47-
.object({ taskBrief: benchmarkBriefSchema.min(1), referenceSpec: benchmarkSpecSchema.min(1) })
46+
.object({
47+
taskBrief: benchmarkBriefSchema.min(1),
48+
referenceSpec: benchmarkSpecSchema
49+
.min(1)
50+
.describe(
51+
'A complete, human-readable Markdown specification, with headings and prose. Use code blocks for exact mappings or code where needed.'
52+
),
53+
})
4854
.strict()
4955
const redactionSchema = z
5056
.object({
@@ -91,23 +97,19 @@ async function performStage(
9197
const artifacts = benchmark.artifacts
9298
switch (stage) {
9399
case 'distill': {
94-
const target = await prepareBenchmarkReference.execute({
95-
principal,
96-
input: { organizationId: benchmark.organizationId, benchmarkId: benchmark.id },
97-
})
98-
signal.throwIfAborted()
99100
const result = await executeBenchmarkJson({
101+
principal,
100102
benchmark,
101103
signal,
102104
schema: distillationSchema,
103105
messages: distillationMessages(artifacts.taskBrief),
104106
profile: { stage: 'distill' },
105-
chatId: target.chatId,
106107
})
107108
return { artifacts: applyBenchmarkPatch(artifacts, result), plannerChatId: null }
108109
}
109110
case 'redact': {
110111
const result = await executeBenchmarkJson({
112+
principal,
111113
benchmark,
112114
signal,
113115
schema: redactionSchema,
@@ -134,6 +136,7 @@ async function performStage(
134136
}
135137
case 'reconstruct': {
136138
const result = await executeBenchmarkJson({
139+
principal,
137140
benchmark,
138141
signal,
139142
schema: reconstructionSchema,
@@ -145,6 +148,7 @@ async function performStage(
145148
}
146149
case 'grade': {
147150
const result = await executeBenchmarkJson({
151+
principal,
148152
benchmark,
149153
signal,
150154
schema: gradingSchema,

‎apps/sim/lib/benchmarks/prompts.ts‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,7 @@ function messages(instruction: string, data: unknown): ExecuteMessage[] {
1313

1414
export function distillationMessages(taskBrief: string): ExecuteMessage[] {
1515
return messages(
16-
'Explore the selected workspace using the read-only workspace CLI and describe its implemented behavior as a detailed, self-contained reference specification. Begin by listing all active workflows, follow list pagination, inspect each workflow and its referenced workspace resources, and follow large-output continuations. Cover exact triggers and input shapes, conditions, ownership, field mappings, prompts and code behavior, actions, destinations, cross-workflow relationships, outputs and failure/recovery behavior. Preserve concrete names, values and business rules; do not compress them into a high-level overview. Distinguish implemented behavior from unresolved configuration or inferred intent. Do not invent missing values. Omit editor layout and credential values. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. The reference is for human review before evaluation. Use tools to inspect; return the final JSON only when inspection is complete.',
16+
'Explore the selected workspace using the read-only workspace CLI and describe its implemented behavior as a detailed, self-contained reference specification. Begin by listing all active workflows, follow list pagination, inspect each workflow and its referenced workspace resources, and follow large-output continuations. Cover exact triggers and input shapes, conditions, ownership, field mappings, prompts and code behavior, actions, destinations, cross-workflow relationships, outputs and failure/recovery behavior. Preserve concrete names, values and business rules; do not compress them into a high-level overview. Distinguish implemented behavior from unresolved configuration or inferred intent. Do not invent missing values. Omit editor layout and credential values. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. Write referenceSpec as a human-readable Markdown document with descriptive headings, paragraphs and lists. Use fenced JSON or code blocks only for exact schemas, mappings or code that help explain the behavior. Do not serialize the entire specification as a JSON document inside referenceSpec. The outer JSON object is only the response envelope containing taskBrief and the Markdown referenceSpec string. The reference is for human review before evaluation. Use tools to inspect; return the final JSON only when inspection is complete.',
1717
taskBrief.trim() ? { taskBrief } : {}
1818
)
1919
}

‎apps/sim/lib/benchmarks/repository.integration.ts‎

Lines changed: 106 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -8,21 +8,65 @@ import { drizzle, type PostgresJsDatabase } from 'drizzle-orm/postgres-js'
88
import postgres from 'postgres'
99
import { afterAll, beforeAll, beforeEach, describe, expect, it, vi } from 'vitest'
1010

11-
const database = vi.hoisted(() => {
11+
const database = await vi.hoisted(async () => {
1212
process.env.MOTHERSHIP_BENCHMARK_ENABLED = 'true'
13-
return { current: undefined as PostgresJsDatabase | undefined }
13+
const { createServer } = await import('node:http')
14+
const requests: Record<string, unknown>[] = []
15+
const worker = createServer(async (request, response) => {
16+
let body = ''
17+
for await (const chunk of request) body += chunk
18+
if (request.url !== '/api/mothership/execute') {
19+
response.writeHead(200, { 'content-type': 'application/json' }).end('{}')
20+
return
21+
}
22+
const payload = JSON.parse(body) as Record<string, unknown>
23+
requests.push(payload)
24+
const frames = [
25+
{ type: 'text', payload: { channel: 'assistant', text: '{"ok":true}' } },
26+
{ type: 'complete', payload: { status: 'complete' } },
27+
]
28+
response.writeHead(200, { 'content-type': 'text/event-stream' })
29+
response.end(
30+
frames
31+
.map(
32+
(frame, index) =>
33+
`data: ${JSON.stringify({
34+
v: 1,
35+
seq: index + 1,
36+
ts: new Date().toISOString(),
37+
stream: { streamId: payload.messageId },
38+
...frame,
39+
})}\n\n`
40+
)
41+
.join('')
42+
)
43+
})
44+
await new Promise<void>((resolve) => worker.listen(0, '127.0.0.1', resolve))
45+
const address = worker.address()
46+
if (!address || typeof address === 'string') throw new Error('Worker fixture did not bind')
47+
process.env.MOTHERSHIP_BENCHMARK_URL = `http://127.0.0.1:${address.port}`
48+
process.env.COPILOT_API_KEY = 'local-benchmark-fixture'
49+
return { current: undefined as PostgresJsDatabase | undefined, worker, requests }
1450
})
1551
vi.mock('server-only', () => ({}))
16-
vi.mock('@sim/db', () => ({
17-
get db() {
18-
if (!database.current) throw new Error('Benchmark test database is not initialized')
19-
return database.current
20-
},
21-
}))
52+
vi.mock('@sim/db', () => {
53+
const scopedDatabase = new Proxy(
54+
{},
55+
{
56+
get(_target, property) {
57+
if (!database.current) throw new Error('Benchmark test database is not initialized')
58+
const value = Reflect.get(database.current, property)
59+
return typeof value === 'function' ? value.bind(database.current) : value
60+
},
61+
}
62+
)
63+
return { db: scopedDatabase, dbFor: () => scopedDatabase }
64+
})
2265

66+
import { z } from 'zod'
2367
import { createBenchmark, getBenchmark, listBenchmarks } from '@/lib/benchmarks/application/cases'
68+
import { prepareBenchmarkExecution } from '@/lib/benchmarks/application/prepare-execution'
2469
import { prepareBenchmarkPlan } from '@/lib/benchmarks/application/prepare-plan'
25-
import { prepareBenchmarkReference } from '@/lib/benchmarks/application/prepare-reference'
2670
import {
2771
getBenchmarkRun,
2872
listBenchmarkRuns,
@@ -46,6 +90,7 @@ import {
4690
updateBenchmarkRecord,
4791
} from '@/lib/benchmarks/repository'
4892
import { type BenchmarkArtifacts, emptyBenchmarkArtifacts } from '@/lib/benchmarks/types'
93+
import { executeBenchmarkJson } from '@/lib/benchmarks/worker'
4994
import { listOrganizationChats } from '@/lib/mothership/chat/organization-chats'
5095

5196
describe('private benchmark persistence and attempt fencing', () => {
@@ -113,6 +158,9 @@ describe('private benchmark persistence and attempt fencing', () => {
113158
'user',
114159
'settings',
115160
'copilot_chats',
161+
'copilot_runs',
162+
'copilot_request_stops',
163+
'copilot_organization_request_stops',
116164
'mothership_memory_selections',
117165
'mothership_memory_spaces',
118166
'organization',
@@ -128,6 +176,7 @@ describe('private benchmark persistence and attempt fencing', () => {
128176
]) {
129177
await connection`CREATE TABLE ${connection(table)} (LIKE ${connection(`public.${table}`)} INCLUDING ALL)`
130178
}
179+
await connection`ALTER TABLE copilot_runs ADD CONSTRAINT benchmark_chat_fk FOREIGN KEY (chat_id) REFERENCES copilot_chats(id)`
131180
database.current = drizzle(connection)
132181
await connection`INSERT INTO "user" (id, name, email, email_verified, created_at, updated_at) VALUES ('owner', 'Owner', 'owner@benchmark.test', true, now(), now()), ('peer', 'Peer', 'peer@benchmark.test', true, now(), now())`
133182
await connection`INSERT INTO "user" (id, name, email, email_verified, created_at, updated_at) VALUES ('target', 'Target', 'target@benchmark.test', true, now(), now())`
@@ -138,7 +187,7 @@ describe('private benchmark persistence and attempt fencing', () => {
138187
})
139188

140189
beforeEach(async () => {
141-
await connection`TRUNCATE mothership_benchmark_runs, mothership_benchmarks, permissions, copilot_chats`
190+
await connection`TRUNCATE mothership_benchmark_runs, mothership_benchmarks, permissions, copilot_runs, copilot_chats`
142191
await connection`UPDATE "user" SET role = CASE WHEN id IN ('owner', 'peer') THEN 'admin' ELSE 'user' END, banned = false`
143192
await connection`UPDATE settings SET super_user_mode_enabled = true`
144193
await connection`INSERT INTO member (id, organization_id, user_id, role) VALUES ('owner-member', 'org', 'owner', 'member'), ('target-member', 'org', 'target', 'member') ON CONFLICT (id) DO UPDATE SET organization_id = 'org'`
@@ -157,7 +206,52 @@ describe('private benchmark persistence and attempt fencing', () => {
157206
} finally {
158207
database.current = undefined
159208
await connection.end()
209+
await new Promise<void>((resolve, reject) => {
210+
database.worker.close((error) => (error ? reject(error) : resolve()))
211+
database.worker.closeAllConnections()
212+
})
213+
}
214+
})
215+
216+
it('completes isolated JSON executions through the real lifecycle with fresh target-owned chats', async () => {
217+
const { benchmark } = await createBenchmark.execute({
218+
principal,
219+
input: {
220+
organizationId: 'org',
221+
sourceWorkspaceId: 'workspace',
222+
runAsUserId: 'target',
223+
name: 'JSON execution',
224+
},
225+
})
226+
database.requests.length = 0
227+
for (let attempt = 0; attempt < 2; attempt++) {
228+
const input = {
229+
principal,
230+
benchmark,
231+
messages: [{ role: 'user' as const, content: 'Return the fixture result.' }],
232+
schema: z.object({ ok: z.literal(true) }),
233+
signal: new AbortController().signal,
234+
}
235+
expect(await executeBenchmarkJson(input)).toEqual({ ok: true })
160236
}
237+
const runs =
238+
await connection`SELECT r.status, r.user_id, c.user_id AS chat_user_id, c.config FROM copilot_runs r JOIN copilot_chats c ON c.id = r.chat_id`
239+
expect(runs).toHaveLength(2)
240+
for (const run of runs)
241+
expect(run).toMatchObject({
242+
status: 'complete',
243+
user_id: 'target',
244+
chat_user_id: 'target',
245+
config: { benchmark: { id: benchmark.id, operatorUserId: 'owner' } },
246+
})
247+
expect(database.requests).toHaveLength(2)
248+
expect(new Set(database.requests.map((request) => request.chatId)).size).toBe(2)
249+
for (const request of database.requests)
250+
expect(request).toMatchObject({
251+
userId: 'target',
252+
useConversationHistory: false,
253+
messages: [{ role: 'user', content: 'Return the fixture result.' }],
254+
})
161255
})
162256

163257
it('requires a current superuser role and enabled toggle even when the deployment flag is on', async () => {
@@ -233,7 +327,7 @@ describe('private benchmark persistence and attempt fencing', () => {
233327
},
234328
})
235329
expect(benchmark).toMatchObject({ userId: 'owner', runAsUserId: 'target' })
236-
const reference = await prepareBenchmarkReference.execute({
330+
const reference = await prepareBenchmarkExecution.execute({
237331
principal,
238332
input: { organizationId: 'org', benchmarkId: benchmark.id },
239333
})
@@ -287,7 +381,7 @@ describe('private benchmark persistence and attempt fencing', () => {
287381
})
288382
).rejects.toMatchObject({ code: 'forbidden' })
289383
await expect(
290-
prepareBenchmarkReference.execute({
384+
prepareBenchmarkExecution.execute({
291385
principal,
292386
input: { organizationId: 'org', benchmarkId: benchmark.id },
293387
})

‎apps/sim/lib/benchmarks/worker.ts‎

Lines changed: 10 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,7 @@ import type { Principal } from '@sim/auth/principal'
22
import { createLogger } from '@sim/logger'
33
import { generateId } from '@sim/utils/id'
44
import { z } from 'zod'
5+
import { prepareBenchmarkExecution } from '@/lib/benchmarks/application/prepare-execution'
56
import { prepareBenchmarkPlan } from '@/lib/benchmarks/application/prepare-plan'
67
import { getBenchmarkMothershipUrl } from '@/lib/benchmarks/config'
78
import type { BenchmarkCase } from '@/lib/benchmarks/types'
@@ -58,20 +59,26 @@ async function stopIncompleteRun(
5859

5960
/** Each invocation uses a fresh conversation; optional profiles expose only the stage's permitted reads. */
6061
export async function executeBenchmarkJson<S extends z.ZodType>(input: {
62+
principal: Principal
6163
benchmark: BenchmarkCase
6264
messages: ExecuteMessage[]
6365
schema: S
6466
signal: AbortSignal
6567
profile?: BenchmarkExecution
66-
chatId?: string
6768
}): Promise<z.output<S>> {
6869
getBenchmarkMothershipUrl()
69-
const executionUserId = input.benchmark.runAsUserId ?? input.benchmark.userId
70+
input.signal.throwIfAborted()
71+
const target = await prepareBenchmarkExecution.execute({
72+
principal: input.principal,
73+
input: { organizationId: input.benchmark.organizationId, benchmarkId: input.benchmark.id },
74+
})
75+
input.signal.throwIfAborted()
76+
const executionUserId = target.userId
7077
const messageId = generateId()
7178
const payload: ExecuteRequest = {
7279
protocolVersion: PROTOCOL_VERSION,
7380
messageId,
74-
chatId: input.chatId ?? generateId(),
81+
chatId: target.chatId,
7582
benchmark: input.profile,
7683
userId: executionUserId,
7784
workspaceId: input.benchmark.sourceWorkspaceId,

0 commit comments

Comments
 (0)