Skip to content

New example: examples/servers/admission-gate — a real approve/deny ServerMiddleware demo #3272

Description

@fede-kamel

Summary

Propose adding a new example server, examples/servers/admission-gate, demonstrating a real approve/deny use of ServerMiddleware — the SDK has this real, documented pre-execution veto point (runner.py's own comment calls it a "middleware veto") but no existing example actually denies anything. The one middleware example (stories/middleware) is audit-logging only; the SDK's only human-in-the-loop pattern (stories/refund_desk) uses elicitation for mid-call parameter confirmation, a different mechanism from a pre-call approve/deny gate on the whole tool call.

What it is

A filesystem server (read_file/write_file/delete_file) whose write_file/delete_file calls go through a middleware that queries an admission model with a written policy and the proposed action, and raises MCPError (the same mechanism handler errors already use) instead of calling call_next when the model says the policy requires denial.

The model backend is real tulip-agents code (tulip.models.native.openai.OpenAIModel, whose base_url override is documented for vLLM endpoints) — a dependency of this one example only, declared in its own pyproject.toml, not the rest of the workspace. Demoed against Clusiana, a real (unreleased) checkpoint trained for this exact three-way decision; any tulip-compatible chat model works the same way.

Verified

Real MCP client, real stdio transport, the server run as its own installed console script, the gate backed by a live model over a real network call — 4/4 correct on a representative probe, with the denied writes/deletes independently confirmed to have genuinely not touched the filesystem (not the tool's own claimed result). Full methodology: gist.

Scope

Doesn't touch src/mcp at all — new example directory only, own pyproject.toml, matches the existing examples/servers/* pattern (each with independent dependencies, e.g. simple-auth adds pydantic-settings).

Have a working, tested implementation ready — opening this first per CONTRIBUTING.md before submitting the PR.

Activity

  1. Ecocitizenz commented on Aug 9, 2026

    @Ecocitizenz
  2. added a commit that references this issue on Aug 10, 2026
  3. fede-kamel commented on Aug 10, 2026

    @fede-kamel
    Author

    Thanks — this was a genuinely useful review, and it's a real gap the example had. Pushed a commit addressing each point:

    • Decisions now bind to an args_hash (tool + arguments), not just the tool name — a stored decision is checkable against a later claim of "the same call."
    • The policy is versioned (POLICY_VERSION).
    • Failure modes are a fixed set of ReasonCodes, not freeform strings — unconfigured gate, timeout, unreachable backend, off-schema response are each their own code. All four still resolve to require_human, never allow — uncertainty escalates rather than being read as permission, which was already true in the old code but is now testable per-failure-mode rather than asserted. Added a parametrized test covering all four explicitly.
    • A minimal, secret-redacted DecisionRecord is emitted per gated call: request_id, tool_id, capability_class, args_hash, policy_version, model_id, decision, reason_code, issued_at — logged structurally and kept in decision_log(). It deliberately never carries the raw arguments (write_file's content could be arbitrary file data).
    • Bounded timeout on the admission call itself (GATE_TIMEOUT_SECONDS, default 10s) — an unbounded call was its own fail-open risk if the backend hangs rather than errors.
    • Added tests/test_admission_gate.py — 11 real tests exercising the middleware's veto logic directly (allow/deny/escalate paths, the decision record each produces, every gate-failure code), independent of whether a live admission model is reachable. Matches your "independently testable" framing directly.

    Two things I did not build, called out explicitly in the README rather than left implicit:

    • Replay of a stale approval against different arguments — turns out not to be a live risk in this design's current shape, not just by policy: nothing here caches or reuses a decision across calls, so every tools/call gets a fresh classification against its own args_hash. There's no stored approval to replay against different arguments in the first place. Worth flagging if you think that's the wrong assumption for where this pattern would actually get used.
    • Tool-schema and model/checkpoint pinning — not attempted. model_id is recorded on every decision so a checkpoint swap is at least visible after the fact, but nothing enforces re-approval on a schema or checkpoint change. Agree this is real, non-trivial work; out of scope for an example this size.

    Re-verified live against the same real backing model after the change (4/4 correct, real MCP client/server round trip, real filesystem checks) — no regression.

    Separately: I still can't get GitHub to let me open the PR itself from this account (CreatePullRequest is rejected specifically for this org — permissions issue on my end, not yours, been stuck on it a few sessions now). Branch is pushed and current at fede-kamel:feat/admission-gate-example if anyone with the right access wants to open it in the meantime.

  4. Ecocitizenz commented on Aug 10, 2026

    @Ecocitizenz
  5. added
    documentationImprovements or additions to documentation
    P3Nice to haves, rare edge cases
    v2Affects the v2 line (2.x on main)
    on Aug 11, 2026
  6. HadiAbu commented on Sep 19, 2026

    @HadiAbu

    Is the issue still open?

  7. hippoley commented on Sep 22, 2026

    @hippoley

    +1 on adding a denial example. The useful part here is showing that middleware can be an execution boundary, not only an observability wrapper.

    Two things I would make very explicit in the example:

    1. Separate admission policy from the mechanism that predicts a decision.
    If the example uses an LLM-backed gate, it would help to keep the middleware contract generic and make the model one optional policy implementation. Otherwise readers may copy the pattern as "trust a second model to guard the first model", when many deployments will want deterministic policy, ACLs, or human approval at this boundary.

    2. Emit a small decision receipt.
    For a denied or modified call, I would log something like:

    tool_call_id
    tool_name
    decision = allow | deny
    policy_id / policy_version
    reason_category
    call_next_invoked = false
    

    No raw sensitive args required. The important invariant is that the trace can prove the proposed call existed, the admission decision happened before handler execution, and the denied filesystem mutation never occurred.

    One more subtlety: current ServerMiddleware runs before params validation, so I think the example should be very clear about whether policy sees raw params or a validated/sanitized representation. That choice affects what people can safely put in a real admission layer.

    A small test asserting that call_next was not reached and the filesystem stayed unchanged would make the example especially strong.

  8. maxisbey commented on Oct 2, 2026

    @maxisbey
    Contributor

    Closed with #3613, and thanks for pointing out the gap.

    The Middleware docs page now has a worked, tested example of a middleware that refuses tool calls: a small concurrency cap that raises MCPError instead of calling call_next.

    That's a good deal smaller than the admission-gate server you proposed; we'd like examples in this repo to run on the SDK alone, with no third-party package or live model call, so the model-backed gate is better kept in its own repository.

    It's merged to main, and if it doesn't cover what you were after you're welcome to open a new issue.

    AI Disclaimer

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P3Nice to haves, rare edge casesdocumentationImprovements or additions to documentationv2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions