Skip to content

fix(sandbox): withhold workbench certification after an unprovenanced file mount - #8456

Merged
waleedlatif1 merged 4 commits into
stagingfrom
fix/sim-cli-result-provenance
Sep 30, 2026
Merged

waleedlatif1 merged 4 commits into
stagingfrom
fix/sim-cli-result-provenance

Conversation

@waleedlatif1

Copy link
Copy Markdown
Collaborator

Summary

A chat's persistent workbench is certified "clean" only while every input it has received was classified secret-free. The scratch-file read relies on that certification to hand workbench bytes to the model. Function file mounts from platform file objects counted toward that history only when the file had a provenance source. A mounted file with no source (no principal to bind one, or a key without a canonical metadata record) left the machine certified clean, although its bytes were never classified. Model-supplied Function parameters could reach this resolver, and a model-supplied _sandboxFiles entry could skip input resolution entirely.

The fix:

  • resolveUserFileMounts now reports how many mounts had no provenance source. Storage contexts other than workspace and execution always count, which is conservative by design.
  • When a workbench session receives any such mount, the session request carries unprovenancedInputs, and the code boundary records the machine as unknown. Both the code and shell paths do this.
  • The Copilot Function handler strips model-supplied _sandboxFiles. Only resolved inputs, which carry their provenance, may populate it.

This is the right layer because the workbench's input history is the single record that the scratch read and any later certification consult. Marking the shared execution registry incomplete inside the resolver would instead latch every workflow Function block that mounts a legacy file. Workflow runs have no persistent session, so they keep their existing policy for files with no provenance.

Why not certifying clean-machine uploads

The investigation began with successful sim_cli reads being withheld from the model, mostly because an uploaded file's provenance was unknown. We considered classifying receipt-less uploads from a clean workbench as secret-free and rejected it. Every workbench command runs with the sandbox's own session credential in its environment, which the clean-machine history does not track. A byte scan cannot prove a file free of that credential, so a secret-free label on such an upload cannot be proven. That track stays open, and the withheld reads are not addressed here.

Testing

  • sandbox-mounts.test.ts: the unprovenanced count is 1 for a key without a metadata record, 0 for one with a record, and 1 when no principal can bind a source.
  • session-input-certification.integration.ts (real Redis): the real code and shell sandbox boundaries run against the real history script. A session with unprovenanced mounts leaves the machine uncertified, and one without keeps it clean. It fails with the boundary gate reverted.
  • execute-request.test.ts: the route marks the workbench session when the resolver reports an unprovenanced mount, and leaves it unmarked otherwise.
  • function-execute-provenance.test.ts: a model-supplied _sandboxFiles URL mount never reaches the Function request. It fails with the omit reverted.
  • Gates: bun run lint, bun run type-check, bun run check:audits, and the affected unit suites (function-execution, remote-sandbox, Copilot tool handlers, agent CLI, chat application).

… file mount

A persistent chat workbench stays "clean" only while every input it received
was classified secret-free, and the scratch-file read hands its bytes to the
model on that basis. File mounts resolved from platform file objects were
counted only when a provenance source existed, so a mount whose key has no
canonical metadata record (or no principal to bind one) left the machine
certified. Both the `files` parameter and a mount marker in context variables
reach this resolver from model-supplied Function parameters.

The resolver now reports how many mounts had no provenance source. When a
workbench session receives any, the session request carries
`unprovenancedInputs` and the code boundary records the machine as unknown.
Workflow runs have no session and keep their existing absence policy.
… the mount bypass

- Replace the certification unit test, which restated the history script in a
  fake Redis, with an integration suite that runs the real code boundary against
  the real script in a disposable Redis.
- Have the route test's mocked mount resolver return a fixed count per test
  rather than restating the counting rule.
- Strip model-supplied `_sandboxFiles` from Copilot Function calls. Only resolved
  inputs may populate it, and a supplied URL mount would skip their provenance.
- Document that public storage contexts always count as unprovenanced mounts.
@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 30, 2026 9:25am UTC

Request Review

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 9 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread apps/sim/lib/execution/remote-sandbox/session-input-certification.integration.ts Outdated
Comment thread apps/sim/lib/execution/remote-sandbox/session-input-certification.integration.ts Outdated
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Critical risk] Adds provenance tracking to sandbox file mounts in persistent workbenches.

The PR appears safe to merge based on the changes since the previous review.

Summary

The PR withholds clean-workbench certification when a mounted file lacks provenance and prevents model-supplied sandbox mounts from bypassing input resolution. The change since the previous review adds Redis connection cleanup to the certification integration test.

  • No new actionable issue was found in that cleanup.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Function file mount] --> B[Resolve provenance]
  B -->|Source found| C[Record contributing input]
  B -->|No source| D[Mark session input unprovenanced]
  D --> E[Withhold clean-machine certification]
Loading

Reviews (3) · Last reviewed commit: "test(sandbox): close the shared Redis cl..."

The suite's cleanup assigned undefined to REDIS_URL when it had been
unset, which stores the string "undefined" for later suites in the
worker, and asked for a Redis client even when the suite was skipped.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 9 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 9 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Re-trigger cubic

@waleedlatif1
waleedlatif1 merged commit e09970a into staging Sep 30, 2026
32 checks passed
@waleedlatif1
waleedlatif1 deleted the fix/sim-cli-result-provenance branch September 30, 2026 18:28

This branch was previously deployed

1 inactive deployment
Preview — cfce8b3e Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant