Repository navigation
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
|
There was a problem hiding this comment.
All reported issues were addressed across 245 files
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Re-trigger cubic
ff42b48 to
7931397
Compare
|
@cubic-dev-ai review this PR |
@Sg312 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 252 files
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Re-trigger cubic
7931397 to
c0bf037
Compare
0f35b12 to
c88d1a8
Compare
|
@cubic-dev-ai review this PR |
@icecrasher321 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 261 files
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Turn on auto-fix | Re-trigger cubic
c88d1a8 to
468fc3f
Compare
|
@cubic-dev-ai review this PR |
@icecrasher321 I have started the AI code review. It will take a few minutes to complete. |
@icecrasher321 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
No issues found across 285 files
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.
Turn on auto-fix | Re-trigger cubic
…nto feat/plan-discovery-hillclimb
|
Scope correction for the latest summary: the failed-load Back / browser-dialog implementation entered staging in 51cdbe6. All nine files from that commit are byte-identical between this head and current staging, and none appears in the complete 285-file PR diff (GitHub and local file lists match). This feature therefore adds no browser-dialog navigation implementation delta. This only corrects attribution; it does not disprove the separate stale-Back callback concern, and no unrelated browser implementation was changed or dismissed here. |
|
@cubic-dev-ai review this PR |
@Sg312 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
No issues found across 285 files
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.
Turn on auto-fix | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@Sg312 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 286 files
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.
Turn on auto-fix | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@Sg312 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
2 issues found across 581 files
Confidence score: 3/5
- In
apps/sim/lib/benchmarks/artifacts.ts, redacting one occurrence can leave another gold answer visible while still passing the reference check, so reconstruction receives the answer. Reject selected answers if any occurrence remains visible. - In
apps/sim/app/api/mothership/chats/route.ts, workspace authorization failures hit the genericforbiddenbranch first, so workspace chats returnOrganization access deniedinstead of the workspace-specific response. Handle workspace failures before the generic branch.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="apps/sim/app/api/mothership/chats/route.ts">
<violation number="1" location="apps/sim/app/api/mothership/chats/route.ts:99">
P3: Workspace authorization failures from this use case are classified as `forbidden`, so the earlier generic branch returns `Organization access denied` for a workspace chat before the workspace-specific check. Handle workspace-use-case authorization failures before the organization error mapping so callers receive `Workspace access denied`.</violation>
</file>
<file name="apps/sim/lib/benchmarks/artifacts.ts">
<violation number="1" location="apps/sim/lib/benchmarks/artifacts.ts:56">
P2: This accepts redactions that leave repeated occurrences of a gold answer visible because restoring the marker still matches the reference. The reconstruction stage receives that text; reject selected answers that remain outside markers.</violation>
</file>
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.
Turn on auto-fix | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@Sg312 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
No issues found across 586 files
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
You've manually re-run cubic several times on this PR. Each manual re-review checks the full PR again and counts toward your usage quota. To preserve your usage limits, we recommend letting cubic automatically review new commits.
Turn on auto-fix | Re-trigger cubic
Summary
Add private named knowledge graphs, detailed Plan-mode specifications and an organization benchmark console. Plan and knowledge use the same server-side eligibility as Benchmark UI: the existing benchmark deployment opt-in plus platform admin status and enabled Super User Mode. The UI, chat creation/admission, graph settings and worker memory scope enforce that policy; ordinary organization admin status is insufficient.
The branch also adds separately gated desktop computer use, read-only Search integration tooling, catalog credential/scope guidance and client-tool result recovery fixes. Computer use and Search tools retain their separate rollout gates; no production flag is enabled by this PR. A benchmark's selected user must independently qualify for Plan.
Product update
Eligible super users can investigate an organization, author a detailed automation specification as a native Sim Page, and reuse private knowledge across chats and workspaces. Plan and graphs require the existing benchmark deployment opt-in and Super User Mode; computer use and Search integration tools retain independent rollout gates.
Type of Change
Testing
Latest review fixes recheck the workspace organization while publishing a fork, reject repeated benchmark answers left outside redaction markers, and preserve workspace-specific authorization errors. Their regressions were shown failing before repair. The Stripe concurrency test now waits until its parking transaction actually holds the locks before starting competitors and releases them in a finally block. Fork route fixtures use table-specific results, and the ledger aggregate unit test pins its clock; assertions and timeouts are unchanged. Existing real-PostgreSQL tests still prove elapsed-time deduction and deadline exhaustion.
Actual production Next build passes. The CI prerender failure was reproduced before extracting the unchanged trigger builder into a registry-free module; trigger definitions now import that owner directly, and the cycle audit covers the registry loop.
The existing packaged desktop local-files E2E passes with its real feature-flag provider. Page-exit regressions cover retained capacity and late-ack identity races. Additional review regressions cover unavailable Plan policy, scoped cancellation, benchmark retries/redaction/versioning and null native status; each relevant guard was demonstrated red before its fix.
Full root
bun run test: 410 script tests and all 20 workspace tasks pass; Sim app 37,115 pass, 21 opt-in skips.All 26 type/lint tasks, 58 architecture audits, workflow lint, block-registry checks, generated docs, migration safety and schema drift pass.
Complete migrations apply to fresh local PostgreSQL; 37 real graph/benchmark/access/concurrency cases passed in the preceding round. The latest graph ownership and Stripe regressions pass 49 cases; seven PostgreSQL ledger-budget tests also pass. Coverage includes workspace transfer during chat creation or fork and benchmark target deletion.
Streaming regressions cover proxy idle periods, rate limits, access denial, cancellation, incomplete streams and split UTF-8. Private query/mutation caches remain isolated across account changes.
The final combined-tree gate passes all 20 workspace tasks, including 1,127 desktop and 478 CLI tests. Focused handler/unload/format checks pass 142 tests. An existing cancellation test now synchronizes on actual tool entry and completion instead of assuming completion within 10 milliseconds; its assertions are unchanged.
Client unload delivery and native 16 MiB reply-bound regressions fail against the old behavior and pass with the fixes. The existing benchmark UI fixture now supplies its authenticated session; no assertions or timeouts were weakened.
Staging is integrated with a normal merge. Feature eligibility and Slack scope metadata use narrow imports; all 82 guarded module graphs pass their unchanged baseline. The desktop build retains staging's filesystem modules and compiles the dormant helper without executing it or requesting permissions.
Review fixes preserve acknowledged tool outcomes during page exit, bound native replies without losing dispatch metadata, serialize graph binding with workspace transfers, cascade target-user deletion and keep long benchmark requests alive with heartbeats. Computer-use and Search rollout gates remain separate. No live configuration, deployment, PR merge, paid trial or customer-data read was performed in this validation pass.
Additive migrations
0404–0410add graph/benchmark storage and bounded-lockNOT VALIDconstraints. Dev databases retaining an earlier feature migration journal still need the documented reconciliation before redeployment; no live journal was changed.Preview limitation: account deletion does not yet erase Graphiti records. Hosted offboarding is documented as unimplemented; owner-to-namespace tracking, durable erasure and indexer fencing are required before enabling Graphiti for production users. This PR keeps the existing dev infrastructure and super-user/deployment gates.
Checklist
Screenshots/Videos
No new screenshots captured for the eligibility change. Plan and knowledge settings use the existing composer and settings layouts.
Companion PR
Companion: https://github.com/simstudioai/mothership/pull/598