Skip to content

Add super-user Plan and knowledge graph settings with benchmark tooling - #8637

Open
Sg312 wants to merge 51 commits into
stagingfrom
feat/plan-discovery-hillclimb
Open

Sg312 wants to merge 51 commits into
stagingfrom
feat/plan-discovery-hillclimb

Conversation

@Sg312

@Sg312 Sg312 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add private named knowledge graphs, detailed Plan-mode specifications and an organization benchmark console. Plan and knowledge use the same server-side eligibility as Benchmark UI: the existing benchmark deployment opt-in plus platform admin status and enabled Super User Mode. The UI, chat creation/admission, graph settings and worker memory scope enforce that policy; ordinary organization admin status is insufficient.

The branch also adds separately gated desktop computer use, read-only Search integration tooling, catalog credential/scope guidance and client-tool result recovery fixes. Computer use and Search tools retain their separate rollout gates; no production flag is enabled by this PR. A benchmark's selected user must independently qualify for Plan.

Product update

Eligible super users can investigate an organization, author a detailed automation specification as a native Sim Page, and reuse private knowledge across chats and workspaces. Plan and graphs require the existing benchmark deployment opt-in and Super User Mode; computer use and Search integration tools retain independent rollout gates.

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation
  • Other

Testing

Latest review fixes recheck the workspace organization while publishing a fork, reject repeated benchmark answers left outside redaction markers, and preserve workspace-specific authorization errors. Their regressions were shown failing before repair. The Stripe concurrency test now waits until its parking transaction actually holds the locks before starting competitors and releases them in a finally block. Fork route fixtures use table-specific results, and the ledger aggregate unit test pins its clock; assertions and timeouts are unchanged. Existing real-PostgreSQL tests still prove elapsed-time deduction and deadline exhaustion.

  • Actual production Next build passes. The CI prerender failure was reproduced before extracting the unchanged trigger builder into a registry-free module; trigger definitions now import that owner directly, and the cycle audit covers the registry loop.

  • The existing packaged desktop local-files E2E passes with its real feature-flag provider. Page-exit regressions cover retained capacity and late-ack identity races. Additional review regressions cover unavailable Plan policy, scoped cancellation, benchmark retries/redaction/versioning and null native status; each relevant guard was demonstrated red before its fix.

  • Full root bun run test: 410 script tests and all 20 workspace tasks pass; Sim app 37,115 pass, 21 opt-in skips.

  • All 26 type/lint tasks, 58 architecture audits, workflow lint, block-registry checks, generated docs, migration safety and schema drift pass.

  • Complete migrations apply to fresh local PostgreSQL; 37 real graph/benchmark/access/concurrency cases passed in the preceding round. The latest graph ownership and Stripe regressions pass 49 cases; seven PostgreSQL ledger-budget tests also pass. Coverage includes workspace transfer during chat creation or fork and benchmark target deletion.

  • Streaming regressions cover proxy idle periods, rate limits, access denial, cancellation, incomplete streams and split UTF-8. Private query/mutation caches remain isolated across account changes.

  • The final combined-tree gate passes all 20 workspace tasks, including 1,127 desktop and 478 CLI tests. Focused handler/unload/format checks pass 142 tests. An existing cancellation test now synchronizes on actual tool entry and completion instead of assuming completion within 10 milliseconds; its assertions are unchanged.

  • Client unload delivery and native 16 MiB reply-bound regressions fail against the old behavior and pass with the fixes. The existing benchmark UI fixture now supplies its authenticated session; no assertions or timeouts were weakened.

Staging is integrated with a normal merge. Feature eligibility and Slack scope metadata use narrow imports; all 82 guarded module graphs pass their unchanged baseline. The desktop build retains staging's filesystem modules and compiles the dormant helper without executing it or requesting permissions.

Review fixes preserve acknowledged tool outcomes during page exit, bound native replies without losing dispatch metadata, serialize graph binding with workspace transfers, cascade target-user deletion and keep long benchmark requests alive with heartbeats. Computer-use and Search rollout gates remain separate. No live configuration, deployment, PR merge, paid trial or customer-data read was performed in this validation pass.

Additive migrations 0404–0410 add graph/benchmark storage and bounded-lock NOT VALID constraints. Dev databases retaining an earlier feature migration journal still need the documented reconciliation before redeployment; no live journal was changed.

Preview limitation: account deletion does not yet erase Graphiti records. Hosted offboarding is documented as unimplemented; owner-to-namespace tracking, durable erasure and indexer fencing are required before enabling Graphiti for production users. This PR keeps the existing dev infrastructure and super-user/deployment gates.

Checklist

  • Code follows project style guidelines
  • Self-reviewed my changes
  • Tests added/updated and passing
  • No new warnings introduced (existing upload-test namespace-import and script-template warnings remain)
  • I confirm that I have read and agree to the terms outlined in the Contributor License Agreement (CLA)

Screenshots/Videos

No new screenshots captured for the eligibility change. Plan and knowledge settings use the existing composer and settings layouts.

Companion PR

Companion: https://github.com/simstudioai/mothership/pull/598

@Sg312
Sg312 requested a review from a team as a code owner October 5, 2026 22:36
@vercel

vercel Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 9, 2026 8:08pm UTC

Request Review

@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High impact] The PR appears safe to merge based on this review; no new actionable finding or outstanding previous finding remains.

Summary

The PR adds gated Plan chats, private knowledge graphs, an organization benchmark console, desktop computer use, and read-only Search tooling. The latest changes recheck workspace ownership before publishing a fork, tighten benchmark redaction validation, preserve workspace-specific refusal messages, and stabilize test synchronization.

Reviews (19) · Last reviewed commit: "fix(plan): preserve fork scope and bench..." · Reviewed by Greptile

Comment thread apps/sim/lib/benchmarks/application/prepare-plan.ts Outdated
Comment thread apps/sim/lib/mothership/memory/application/read-scope.ts
Comment thread apps/sim/lib/benchmarks/application/stage-lease.ts
@github-actions github-actions Bot added the requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep label Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

⚠️ Cross-repo companion check

One or more companion PRs aren't merged into staging yet. Merging this without them will leave copilot and sim out of sync — merge them in lockstep.

  • ❌ simstudioai/mothership#598 — OPEN, not merged (targets staging) — Add super-user Plan, private knowledge graphs, and benchmark orchestration

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 245 files

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.

Re-trigger cubic

Comment thread apps/sim/lib/mothership/tools/client/computer-tool-execution.ts
Comment thread apps/sim/app/o/[organizationId]/layout.tsx Outdated
Comment thread apps/sim/components/settings/navigation.ts
Comment thread packages/desktop-bridge/src/index.ts
Comment thread apps/sim/lib/mothership/chat/organization-chats.ts
Comment thread apps/sim/lib/benchmarks/application/stage-lease.ts
Comment thread apps/desktop/native/computer-use/Info.plist
Comment thread apps/sim/lib/benchmarks/worker.ts
Comment thread apps/sim/hooks/queries/computer-use.ts Outdated
Comment thread apps/sim/app/o/[organizationId]/benchmark/components/benchmark-results.tsx Outdated
@Sg312
Sg312 force-pushed the feat/plan-discovery-hillclimb branch from ff42b48 to 7931397 Compare October 6, 2026 00:36
@Sg312

Sg312 commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@greptile

@Sg312

Sg312 commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@Sg312 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 252 files

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.

Re-trigger cubic

Comment thread apps/sim/hooks/queries/benchmarks.ts
Comment thread apps/sim/lib/benchmarks/application/prepare-plan.ts
Comment thread apps/sim/lib/mothership/tools/client/computer-tool-execution.ts Outdated
Comment thread apps/sim/lib/benchmarks/application/stage-lease.ts
Comment thread apps/sim/lib/workspaces/admin-move.ts
Comment thread apps/sim/lib/mothership/chat/organization-chats.ts
Comment thread apps/sim/lib/benchmarks/application/run-stage.ts
Comment thread apps/sim/hooks/queries/benchmarks.ts
Comment thread apps/sim/app/workspace/[workspaceId]/settings/components/desktop/computer-use.tsx Outdated

Copy link
Copy Markdown
Collaborator

@greptile

Copy link
Copy Markdown
Collaborator

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@icecrasher321 I have started the AI code review. It will take a few minutes to complete.

Copy link
Copy Markdown
Collaborator

@greptile

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 261 files

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.

Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/mothership/agent-cli/index.ts
Comment thread apps/sim/lib/mothership/tools/desktop-tools.ts
Comment thread apps/sim/lib/computer-use/application/authorize.ts
Comment thread apps/sim/app/o/[organizationId]/benchmark/benchmark.tsx Outdated
Comment thread apps/desktop/src/main/computer-use/native-client.ts Outdated
Comment thread apps/sim/lib/mothership/request/lifecycle/run.ts
Comment thread apps/desktop/src/main/ipc.ts
Comment thread apps/sim/lib/desktop/index.ts
Comment thread apps/sim/lib/benchmarks/evidence.ts
Comment thread apps/sim/lib/mothership/tools/client/computer-tool-execution.ts
Comment thread apps/sim/app/api/organizations/[id]/benchmarks/[benchmarkId]/compare/route.ts Outdated

Copy link
Copy Markdown
Collaborator

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@icecrasher321 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@icecrasher321 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 285 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.

Turn on auto-fix | Re-trigger cubic

@icecrasher321

Copy link
Copy Markdown
Collaborator

Scope correction for the latest summary: the failed-load Back / browser-dialog implementation entered staging in 51cdbe6. All nine files from that commit are byte-identical between this head and current staging, and none appears in the complete 285-file PR diff (GitHub and local file lists match). This feature therefore adds no browser-dialog navigation implementation delta. This only corrects attribution; it does not disprove the separate stale-Back callback concern, and no unrelated browser implementation was changed or dismissed here.

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@greptile

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@Sg312 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 285 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.

Turn on auto-fix | Re-trigger cubic

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@greptile

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@Sg312 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 286 files

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.

Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/app/workspace/[workspaceId]/layout.tsx Outdated
Comment thread apps/sim/app/o/[organizationId]/benchmark/benchmark.tsx
Comment thread apps/sim/lib/mothership/assistant/tool-policy.ts
Comment thread apps/sim/triggers/slack/oauth.ts
Comment thread packages/db/migrations/0406_mothership_benchmark_runs.sql
Comment thread apps/sim/triggers/slack/shared.ts
Comment thread apps/sim/app/o/[organizationId]/settings/[section]/settings.tsx
Comment thread apps/sim/lib/benchmarks/types.ts Outdated
Comment thread apps/sim/lib/computer-use/transport.ts Outdated
Comment thread apps/desktop/native/computer-use/ComputerUse.swift Outdated
@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@greptile

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@Sg312 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 581 files

Confidence score: 3/5

  • In apps/sim/lib/benchmarks/artifacts.ts, redacting one occurrence can leave another gold answer visible while still passing the reference check, so reconstruction receives the answer. Reject selected answers if any occurrence remains visible.
  • In apps/sim/app/api/mothership/chats/route.ts, workspace authorization failures hit the generic forbidden branch first, so workspace chats return Organization access denied instead of the workspace-specific response. Handle workspace failures before the generic branch.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apps/sim/app/api/mothership/chats/route.ts">

<violation number="1" location="apps/sim/app/api/mothership/chats/route.ts:99">
P3: Workspace authorization failures from this use case are classified as `forbidden`, so the earlier generic branch returns `Organization access denied` for a workspace chat before the workspace-specific check. Handle workspace-use-case authorization failures before the organization error mapping so callers receive `Workspace access denied`.</violation>
</file>

<file name="apps/sim/lib/benchmarks/artifacts.ts">

<violation number="1" location="apps/sim/lib/benchmarks/artifacts.ts:56">
P2: This accepts redactions that leave repeated occurrences of a gold answer visible because restoring the marker still matches the reference. The reconstruction stage receives that text; reject selected answers that remain outside markers.</violation>
</file>

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.

Turn on auto-fix | Re-trigger cubic

Comment thread apps/sim/lib/benchmarks/application/prepare-execution.ts
Comment thread apps/sim/blocks/blocks/gmail.ts
Comment thread apps/sim/lib/mothership/tool-executor/executor.ts
Comment thread apps/sim/lib/mothership/chat/application/fork.ts Outdated
Comment thread apps/sim/lib/benchmarks/artifacts.ts
Comment thread apps/sim/blocks/blocks/github.ts
Comment thread apps/sim/app/api/mothership/chats/route.ts
@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@greptile

@Sg312

Sg312 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@Sg312 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 586 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.

Turn on auto-fix | Re-trigger cubic

This branch was previously deployed

1 inactive deployment
Preview — 2a701a7f Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants