Repository navigation
New example: examples/servers/admission-gate — a real approve/deny ServerMiddleware demo #3272
Description
Activity
Thanks — this was a genuinely useful review, and it's a real gap the example had. Pushed a commit addressing each point:
- Decisions now bind to an
args_hash(tool + arguments), not just the tool name — a stored decision is checkable against a later claim of "the same call." - The policy is versioned (
POLICY_VERSION). - Failure modes are a fixed set of
ReasonCodes, not freeform strings — unconfigured gate, timeout, unreachable backend, off-schema response are each their own code. All four still resolve torequire_human, neverallow— uncertainty escalates rather than being read as permission, which was already true in the old code but is now testable per-failure-mode rather than asserted. Added a parametrized test covering all four explicitly. - A minimal, secret-redacted
DecisionRecordis emitted per gated call:request_id, tool_id, capability_class, args_hash, policy_version, model_id, decision, reason_code, issued_at— logged structurally and kept indecision_log(). It deliberately never carries the raw arguments (write_file'scontentcould be arbitrary file data). - Bounded timeout on the admission call itself (
GATE_TIMEOUT_SECONDS, default 10s) — an unbounded call was its own fail-open risk if the backend hangs rather than errors. - Added
tests/test_admission_gate.py— 11 real tests exercising the middleware's veto logic directly (allow/deny/escalate paths, the decision record each produces, every gate-failure code), independent of whether a live admission model is reachable. Matches your "independently testable" framing directly.
Two things I did not build, called out explicitly in the README rather than left implicit:
- Replay of a stale approval against different arguments — turns out not to be a live risk in this design's current shape, not just by policy: nothing here caches or reuses a decision across calls, so every
tools/callgets a fresh classification against its ownargs_hash. There's no stored approval to replay against different arguments in the first place. Worth flagging if you think that's the wrong assumption for where this pattern would actually get used. - Tool-schema and model/checkpoint pinning — not attempted.
model_idis recorded on every decision so a checkpoint swap is at least visible after the fact, but nothing enforces re-approval on a schema or checkpoint change. Agree this is real, non-trivial work; out of scope for an example this size.
Re-verified live against the same real backing model after the change (4/4 correct, real MCP client/server round trip, real filesystem checks) — no regression.
Separately: I still can't get GitHub to let me open the PR itself from this account (
CreatePullRequestis rejected specifically for this org — permissions issue on my end, not yours, been stuck on it a few sessions now). Branch is pushed and current atfede-kamel:feat/admission-gate-exampleif anyone with the right access wants to open it in the meantime.- Decisions now bind to an
- addeddocumentationImprovements or additions to documentationImprovements or additions to documentationP3Nice to haves, rare edge casesNice to haves, rare edge casesv2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)
on Aug 11, 2026 Is the issue still open?
hippoley commented
on Sep 22, 2026 More actions+1 on adding a denial example. The useful part here is showing that middleware can be an execution boundary, not only an observability wrapper.
Two things I would make very explicit in the example:
1. Separate admission policy from the mechanism that predicts a decision.
If the example uses an LLM-backed gate, it would help to keep the middleware contract generic and make the model one optional policy implementation. Otherwise readers may copy the pattern as "trust a second model to guard the first model", when many deployments will want deterministic policy, ACLs, or human approval at this boundary.2. Emit a small decision receipt.
For a denied or modified call, I would log something like:tool_call_id tool_name decision = allow | deny policy_id / policy_version reason_category call_next_invoked = falseNo raw sensitive args required. The important invariant is that the trace can prove the proposed call existed, the admission decision happened before handler execution, and the denied filesystem mutation never occurred.
One more subtlety: current ServerMiddleware runs before params validation, so I think the example should be very clear about whether policy sees raw params or a validated/sanitized representation. That choice affects what people can safely put in a real admission layer.
A small test asserting that
call_nextwas not reached and the filesystem stayed unchanged would make the example especially strong.Closed with #3613, and thanks for pointing out the gap.
The Middleware docs page now has a worked, tested example of a middleware that refuses tool calls: a small concurrency cap that raises
MCPErrorinstead of callingcall_next.That's a good deal smaller than the
admission-gateserver you proposed; we'd like examples in this repo to run on the SDK alone, with no third-party package or live model call, so the model-backed gate is better kept in its own repository.It's merged to
main, and if it doesn't cover what you were after you're welcome to open a new issue.
Summary
Propose adding a new example server,
examples/servers/admission-gate, demonstrating a real approve/deny use ofServerMiddleware— the SDK has this real, documented pre-execution veto point (runner.py's own comment calls it a "middleware veto") but no existing example actually denies anything. The one middleware example (stories/middleware) is audit-logging only; the SDK's only human-in-the-loop pattern (stories/refund_desk) uses elicitation for mid-call parameter confirmation, a different mechanism from a pre-call approve/deny gate on the whole tool call.What it is
A filesystem server (
read_file/write_file/delete_file) whosewrite_file/delete_filecalls go through a middleware that queries an admission model with a written policy and the proposed action, and raisesMCPError(the same mechanism handler errors already use) instead of callingcall_nextwhen the model says the policy requires denial.The model backend is real tulip-agents code (
tulip.models.native.openai.OpenAIModel, whosebase_urloverride is documented for vLLM endpoints) — a dependency of this one example only, declared in its ownpyproject.toml, not the rest of the workspace. Demoed against Clusiana, a real (unreleased) checkpoint trained for this exact three-way decision; any tulip-compatible chat model works the same way.Verified
Real MCP client, real stdio transport, the server run as its own installed console script, the gate backed by a live model over a real network call — 4/4 correct on a representative probe, with the denied writes/deletes independently confirmed to have genuinely not touched the filesystem (not the tool's own claimed result). Full methodology: gist.
Scope
Doesn't touch
src/mcpat all — new example directory only, ownpyproject.toml, matches the existingexamples/servers/*pattern (each with independent dependencies, e.g.simple-authaddspydantic-settings).Have a working, tested implementation ready — opening this first per CONTRIBUTING.md before submitting the PR.