Repository navigation
Python: [Feature]: Supply-chain: MCP server allowlist & signature-verification primitive #5864
Description
Activity
- added.NETUsage: [Issues, PRs], Target: .NetUsage: [Issues, PRs], Target: .NetpythonUsage: [Issues, PRs], Target: PythonUsage: [Issues, PRs], Target: PythontriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on May 14, 2026 - changed the title
[-][Feature]: Supply-chain: MCP server allowlist & signature-verification primitive[/-][+].NET: [Feature]: Supply-chain: MCP server allowlist & signature-verification primitive[/+]on May 14, 2026 - changed the title
[-].NET: [Feature]: Supply-chain: MCP server allowlist & signature-verification primitive[/-][+]Python: [Feature]: Supply-chain: MCP server allowlist & signature-verification primitive[/+]on May 14, 2026 - removedtriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on May 21, 2026 Santoshkumarpuppala commented
on Sep 2, 2026 More actionsThe three controls sketched here —
allowed_origins,require_signature,pinned_version— are all evaluated once, atMCPClientconstruction. Four things follow from that, checkable against this repo.1. Connect-time checks do not cover this codebase's discovery surface.
python/packages/core/agent_framework/_mcp.py:1662-1665handlesnotifications/tools/list_changedandnotifications/prompts/list_changedby schedulingload_tools()/load_prompts(). A server that passes origin, signature and version at connect can push an entirely new set of tool and prompt definitions afterwards, and nothing in the sketch re-evaluates. Discovery here is not a connect-time event, so a connect-time primitive covers the first listing only.2. Two of the three controls are inert against the incident cited to justify them. In a Postmark-shaped rug pull the malicious update ships from the legitimate publisher: the origin allowlist passes because the URL is unchanged, and signature verification passes because the key is the real one. Only content pinning bears on that case — and only if the pin covers the tool definition, not a
pinned_version="1.4.2"string the server itself reports. A version the server asserts about itself is not a control against a server that has decided to lie.3. Three defects we shipped in exactly this design, each mapping onto something already in this file.
- Pin identity. Ours is trust-on-first-use keyed on (server, tool name). A rename produces a key with no prior, so it registers as new rather than changed and the integrity check never runs. That maps onto the existing reload path:
_mcp.py:~1905skips bylocal_nameand treats anything not already present as new, so a renamed tool takes the new branch. A name-keyed pin would inherit that. - Field coverage. Ours pins an enumerated list of fields, which makes it an allowlist — anything the spec adds later ships unpinned with no signal that coverage shrank. There is a live instance here:
#7824tracks MCP 2026-07-28 including the Tasks Extension, and this file already readstool.meta(672, 1877) andexecution.taskSupport(351, 1840). A fixed field list drafted today would not cover them. - Scope. Ours pins tools only; prompts, resources and
initialize.instructionsare scanned but never hashed. That one is load-bearing here in a specific way: MAF turns prompts intoFunctionTools in the sameself._functionslist, and_filtered_functions()already appliesallowed_toolsacross the merged list — so the allowlist got the scope right. A pinning primitive needs the same scope the allowlist already has, or it covers half of what this client actually exposes to the model.
4. A number on
on_violation="halt"as the default. We shipped that posture and measured its cost: it withheldcreate_directoryfrom the official MCP filesystem server, because the description contains the phrase "succeed silently" and a concealment rule matched the adverb bare. Upstream served 14 tools, the model saw 13, and the only signal was one line on stderr. Worth setting a false-positive budget alongside the acceptance criteria rather than discovering one.Related, for the audit criterion rather than the gate: we also shipped a defect where a policy refusal was persisted to the audit trail as an allow, because the record was built by a different code path than the one that enforced. Both paths were individually correct. If this grows an audit story, the enforcing path and the recording path should be the same object rather than two readings of the same event.
Disclosure: I maintain an MCP policy enforcement point, which is where all of the above was measured — offered as what the design cost us, not as a recommendation to adopt anything.
- Pin identity. Ours is trust-on-first-use keyed on (server, tool name). A rename produces a key with no prior, so it registers as new rather than changed and the integrity check never runs. That maps onto the existing reload path:
The failure modes you list — a previously-trusted server shipping a malicious update, typosquatted names riding the LLM-suggestion path — are exactly the cases where signature verification plus content-hash pinning earn their keep. On the signing piece specifically: Ed25519-signed server/tool descriptors with a digest pin work well in practice — the signature proves who published it, the SHA-256 pin fixes the exact revision, and pubkey rotation is handled by verifying against a key registry instead of one pinned key.
We built this shape in AOTrust (open provenance layer): each allow/deny decision at the client boundary gets sealed as a short hash-chained signed record (tool, origin, pinned digest, decision, timestamp), so deployments get the ISO 27001 / SOC 2 audit trail you map this control to without logging full payloads — a third party can verify the chain offline. Agree on fail-closed
on_violation="halt"as the default; happy to compare notes on the trust_policy surface.Field data on the "post-connection" gap Santosh Kumar Puppala (@Santoshkumarpuppala) identified. We run a weekly drift tracker over the most-installed MCP servers on npm (pin approved contracts, diff on update, grade by direction). The three controls sketched here -
allowed_origins,require_signature,pinned_version- all operate at connect time. What our tracker shows is that the contract your agent obeys changes between connections, silently, without the version string moving:- One package erased "Requires confirmation" from all 48 destructive SSH tools in a patch release (2.0.2 -> 2.0.3). Same version bump class as a bug fix.
pinned_versionat the minor range would not have caught it. - The This repo is missing a LICENSE file #1 server on npm (1.7M installs/week) re-serialized all 28 tool schemas in five weeks via an SDK migration - every hash-level pin fires, but zero parameters actually changed. Without grading, this is a false-positive storm that trains users to ignore the gate.
- A Lightning wallet gained five betting tools (spending from the agent's balance) in a minor update.
allowed_originswas unchanged - same server, same URL, new capabilities.
218 silent contract changes in our October audit across the npm top. The full distribution: 4 of 9 servers drifted (concentrated in fast movers), 5 were clean for comparable windows. Receipts: https://github.com/Paraphern/rugsnare/blob/main/audits/npm-top-mcp-drift-2026-10.md
What this suggests for the framework primitives:
-
pinned_versionneeds a content-hash component, not just a semver string. Version numbers don't capture contract changes within the same version (mid-session description rewrites are invisible to semver entirely). A sha-256 over each tool's{name, description, inputSchema}pinned at approval, compared on every reconnect ortools/list_changednotification, closes the gap. -
The diff needs grading. A naive hash-comparison flags SDK upgrades (key reordering, dialect changes,
additionalPropertiesform) identically to a rewritten safety instruction. We grade three ways - BREAKING (parameters changed), LOOSENED (constraints dropped), NOTATION (semantically void) - and all three still exit non-zero (it's drift), but the label tells the deployer whether to page someone or click through. Without grading, the false-positive rate on legitimate SDK upgrades will train users to disable the control. -
tools/list_changedshould trigger a diff, not just a re-fetch. If the framework already handles this notification, the re-listed contract can be compared against the pinned baseline at near-zero cost. We described this pattern for the cline CLI (CLI: MCP notifications/tools/list_changed is ignored — tools added mid-session never appear (extension fixed in #12619) cline/cline#14924, fix(core): refresh MCP tools onlist_changednotifications in the session runtime cline/cline#14946) - same architecture applies here.
Happy to contribute the hash-grading implementation or test material (we keep a public corpus of benign/weaponized contract pairs for exactly this kind of integration).
- One package erased "Requires confirmation" from all 48 destructive SSH tools in a patch release (2.0.2 -> 2.0.3). Same version bump class as a bug fix.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
Description
Context
The framework currently supports MCP server integration via
MCPClient(Python) andModelContextProtocol(NuGet). Library versions are pinned (mcp[ws]==1.27.0,ModelContextProtocol Version="1.1.0") — but the MCP servers themselves that a deployer connects to are not subject to any framework-level trust governance:SECURITY.mdorTRANSPARENCY_FAQ.mdfor MCP source trustThe closest existing item is #4927 (declarative MCP binding ergonomics), which is adjacent but distinct — that's about referencing servers, this is about trusting them.
Why this matters now (2026 supply-chain precedents)
The MCP ecosystem has had three publicly-discussed incidents in the last ~6 months:
In each case, deployer-side mitigation alone was insufficient because there was no framework primitive to enforce it.
Proposed primitive (sketch — open to other shapes)
A built-in allowlist + verification layer at the
MCPClientboundary:For .NET, an equivalent
MCPTrustPolicybuilder option onModelContextProtocolClientFactory.Acceptance criteria suggestion:
on_violation="halt"for new clients (fail-closed)Regulatory context
This control maps to:
Without a framework-level primitive, deployers building under any of these frameworks must implement the equivalent layer themselves — which is fragile and not easy to audit consistently.
Source
Found during an internal Aegis Release Gate v2.5 security assessment of MAF
python-1.3.0+dotnet-1.6.1. Happy to share more findings or PoC test cases if helpful.Cross-link: #4927.
Code Sample
Language/SDK
Both