Repository navigation
Python: [Bug]: MCP tools/call retry can duplicate non-idempotent side effects after connection loss #9204
Description
Activity
- addedpythonUsage: [Issues, PRs], Target: PythonUsage: [Issues, PRs], Target: PythontriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 8, 2026 - addedreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflowUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
on Oct 8, 2026 🤖 Automated triage reproduction notes (agent-authored — trust but verify)
Agent analysis
Repro:
python/packages/core/agent_framework/_mcp.py::MCPTool._call_tool_with_retriesaround lines 2875-2909 replays ordinarytools/callrequests afterClosedResourceErroror a session-terminatedMcpError. The trigger is a server completing a non-idempotent operation before the response connection closes. Minimal repro: mocksession.call_tool()to increment persistent state then raiseClosedResourceErroron its first call and succeed on its second; onetool.call_tool()invocation increments the state twice.- Failing test:
python/packages/core/tests/core/test_mcp.py::test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnect - Files examined: python/packages/core/agent_framework/_mcp.py, python/packages/core/tests/core/test_mcp.py, python/packages/core/agent_framework/exceptions.py, python/packages/core/pyproject.toml
- Tests run: test_mcp_tool_reconnects_after_session_terminated_error, test_mcp_tool_connection_properly_invalidated_after_closed_resource_error, test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnect
- Reported version:
1.21.0 - Current version:
1.21.0
- Failing test:
- addedagentsUsage: [Issues, PRs], Target: Single agentUsage: [Issues, PRs], Target: Single agentmcpUsage: [Issues, PRs], Target: MCPUsage: [Issues, PRs], Target: MCPand removedtriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 8, 2026 One retry-policy detail: I would not use a remote
idempotentHintalone to authorize replay after a lost response. MCP tool annotations are hints, not guarantees of runtime behavior; the failure here is precisely that the server may have committed before the response was lost. See the MCP guidance on tool annotations: https://blog.modelcontextprotocol.io/posts/2026-03-16-tool-annotations/No automatic replay on that ambiguous path seems like the right default. If opt-in retries are added later, I would gate them on client-owned trust policy or a provider idempotency/reconciliation contract, not the annotation alone.
A useful availability control alongside the existing lost-response test: after the first call commits and loses its response, assert one effect and one
tools/callplus the unknown-outcome exception; then issue a distinct second call and assert that it reconnects and succeeds without reissuing the first. That tests recovery separately from replay.suhasagg commented
on Oct 8, 2026 ContributorAuthorMore actionsThanks gomission.ai (@gomission) for the detailed feedback.
I agree with the key distinction: preventing replay of an operation with an unknown outcome and restoring connection availability for future operations are separate concerns.
1. Retry safety and execution semantics
The underlying issue is that a transport failure does not necessarily imply that the remote MCP operation failed.
If the server commits a non-idempotent operation but the response is lost, automatically replaying the same
tools/callrequest can produce duplicate side effects.I agree that
idempotentHintalone should not be treated as sufficient authorization for replay. Tool annotations are advisory and do not provide a guaranteed deduplication or exactly-once execution contract.My proposed default behavior is therefore:
- Do not automatically replay a
tools/callrequest when its remote execution outcome is unknown. - Propagate a
ToolExecutionExceptionthat clearly communicates the uncertain outcome and preserves the underlying transport exception. - Keep retry and reconciliation decisions under application control unless an explicit, trustworthy idempotency contract exists.
- Preserve existing behavior for successful calls and ordinary MCP tool-error responses.
This does not guarantee exactly-once execution across distributed systems; it specifically prevents the framework from introducing duplicate execution through an automatic replay.
2. Connection recovery must remain independent of replay
I also agree with your suggestion to validate availability after the ambiguous failure.
The local fix currently prevents replay and supports explicit reconnection. However, it does not yet establish automatic reconnection for a subsequent independent tool call.
There are two possible recovery policies:
Option A — Explicit recovery: The failed operation raises an unknown-outcome exception, and the application explicitly reconnects before initiating another operation.
Option B — Recovery on the next independent call: The failed operation raises an unknown-outcome exception without replay, while a later, distinct tool invocation can automatically restore the connection and execute normally.
I lean toward Option B for usability, provided reconnection can be implemented safely without interfering with concurrent calls, connection lifecycle ownership, or caller-managed MCP sessions.
Importantly, reconnecting must never implicitly replay the original operation.
3. Proposed regression coverage
Once the preferred recovery policy is confirmed, I can extend the tests to verify the complete failure-and-recovery sequence:
- The first
tools/callproduces a simulated server-side effect and then raisesClosedResourceError, representing a lost response. - The original call is attempted exactly once, with no automatic replay.
- The caller receives a
ToolExecutionExceptionindicating that the remote execution outcome is unknown. - The underlying connection exception remains available.
- A distinct second tool invocation is initiated.
- The second invocation succeeds after the appropriate reconnection process.
- The first operation's side-effect count remains exactly one.
I would also retain coverage for session-terminated
McpError, successful tool responses, and normal MCP tool errors.4. Current implementation and validation
I have prepared a local fix, but have not opened a PR.
Local commit:
c10b7ef28Current validation:
- Full MCP test suite: 395 passed, 2 skipped
- Focused duplicate-side-effect regression test: passed
- Ruff lint and formatting: passed
git diff --check: passed
The implementation intentionally avoids introducing a new automatic reconnection policy until the expected lifecycle behavior is agreed upon.
5. Maintainer guidance requested
Eduard van Valkenburg (@eavanvalkenburg) — could you please confirm whether Option A (explicit reconnection) or Option B (automatic recovery on a subsequent independent call) is the preferred behavior for Python MCP tools?
It would also be helpful to confirm whether retrying a potentially non-idempotent
tools/callafter an ambiguous transport failure should be disabled by default, with any future opt-in retry policy requiring stronger guarantees than MCP tool annotations alone.I am happy to adjust the implementation and regression tests to align with the framework's intended semantics.
I will wait for maintainer agreement on the recovery policy before opening the PR.
Thanks again for the review and suggestions.
- Do not automatically replay a
Thanks, that distinction helps. I recommend Option B when the framework owns the session, while keeping caller-managed sessions under Option A.
After an ambiguous call, return the unknown-outcome exception and never replay it. A later, independent tool call can establish a fresh session and proceed. The reconnect should be serialized with connection state so it cannot race active calls or reset a caller-owned session.
The regression should assert one side effect and one
tools/callfor the ambiguous operation, then a separate second call that reconnects and succeeds. That restores availability without implying the first operation failed or promising exactly-once execution.Verified on current main (
_call_tool_with_retries): the firstClosedResourceError/ session-terminatedMcpErrortriggersconnect(reset=True)and an unconditional replay of the sametools/call. Since the error can surface after the server executed the request but before the response arrived, the replay really can run a non-idempotent operation twice. This also sits in direct tension with the #2884 fix, which added the optimistic reconnect-and-retry to ride over transient disconnects.I'd like to take this. Proposal that keeps both properties: gate the replay on the tool's advertised
ToolAnnotations. Servers can declareidempotentHint/readOnlyHint; for those tools a replay after connection loss is safe and the #2884 behavior stays. For everything else the call is not replayed:call_toolraisesToolExecutionExceptionstating that the connection was lost, the remote execution outcome is unknown, and the call was deliberately not retried, with the original connection error as the inner exception so applications can reconnect and reconcile under their own policy.tools/list,get_prompt, and the long-running task path keep their current retry behavior since reads and prompts have no side effects.Regression coverage: the issue's duplication repro (state increments once, exception surfaces), the default no-replay path, and the annotated-idempotent path still reconnecting.
Yufeng He (@he-yufeng), the no-replay path and preserved connection exception fit the failure here. One boundary to keep explicit in the annotated path: the current MCP ToolAnnotations contract says annotations are advisory and must not drive tool-use decisions for untrusted servers. Advertising
readOnlyHintoridempotentHintshould therefore not, by itself, enable replay after a lost response.I would keep no replay as the default and make any exception client-owned opt-in for a trusted server/tool contract. This also matches the issue author's proposed policy.
One additional negative regression would make that boundary reviewable: a server advertises
idempotentHint=True(orreadOnlyHint=True), increments state, then loses the response; without client opt-in, assert one invocation/one effect and the unknown-outcome exception. Keep the separate later-call recovery control alongside it. The positive replay case can explicitly configure trusted policy and a genuinely idempotent fake operation, instead of treating the annotation as that policy.suhasagg commented
on Oct 10, 2026 ContributorAuthorMore actionsThanks Yufeng He (@he-yufeng) and gomission.ai (@gomission) for the detailed investigation and thoughtful feedback.
I appreciate the independent reproduction of the issue and the discussion around the reconnect-and-retry behavior introduced in #2884. The points raised about tool annotations, connection recovery, and session ownership are important for arriving at a safe and maintainable fix.
I'd like to share the current implementation status and propose a path forward.
1. Existing implementation and validation
As the original reporter of #9204, I have already prepared a local fix for the duplicate-execution scenario and completed initial regression testing.
Local implementation:
c10b7ef28Previously completed validation:
- Full MCP test suite: 395 passed, 2 skipped
- Duplicate-side-effect regression test: passed
- Ruff lint and formatting: passed
git diff --check: passed
The existing patch addresses the behavior in
MCPTool._call_tool_with_retries()where a transport failure can cause the sametools/callrequest to be replayed even though the remote operation may already have completed.The fix prevents automatic replay in this ambiguous-outcome scenario and preserves the underlying transport exception.
These results reflect the initial local implementation. I will rerun the relevant checks against the final patch before PR submission.
2. Retry safety and MCP tool annotations
I agree with gomission.ai (@gomission) that
idempotentHintandreadOnlyHintshould not, by themselves, authorize automatic replay following an ambiguous connection failure.MCP tool annotations provide useful descriptive metadata, but they do not independently establish a guaranteed idempotency or deduplication contract.
A server may have committed an operation before the client receives a transport error. In that situation, replaying the operation based solely on an annotation could retain the duplicate-execution risk described in this issue.
My preferred default is therefore:
- Do not automatically replay a
tools/callrequest when its remote execution outcome is unknown. - Propagate an exception clearly communicating the uncertain outcome while preserving the original transport error.
- Keep any future opt-in replay policy separate, with appropriate application-controlled trust or provider-supported idempotency guarantees.
- Preserve existing behavior for successful calls and ordinary MCP tool-error responses.
This approach prevents framework-induced replay without claiming exactly-once execution semantics.
3. Connection recovery and compatibility with #2884
I also agree that preventing replay and restoring connection availability are separate concerns.
The behavior introduced in #2884 was intended to improve resilience to transient disconnections, so I believe the fix should preserve that objective wherever it can be done safely.
Based on the discussion, the preferred direction appears to be:
- Ambiguous operation: Return an unknown-outcome exception without replaying the original request.
- Framework-owned sessions: Allow a subsequent independent operation to establish a fresh connection where appropriate.
- Caller-managed sessions: Respect external session ownership and avoid unexpected reconnection or resets.
- Concurrency: Ensure recovery does not interfere with active operations or introduce connection lifecycle races.
- Compatibility: Minimize unrelated changes to MCP behavior.
My initial patch addresses the unsafe replay behavior. I would welcome maintainer guidance before extending the connection recovery policy, particularly where it affects session lifecycle semantics.
4. Additional regression coverage
I agree with the additional tests suggested in the discussion and propose incorporating the following cases:
- A server completes an operation but the response is lost.
- The original
tools/callis attempted exactly once. - The caller receives an exception indicating that the remote execution outcome is unknown.
- The original transport exception remains available.
- An advertised
idempotentHint=TrueorreadOnlyHint=Truedoes not independently authorize replay without an explicit trusted policy. - A subsequent independent invocation can recover appropriately for framework-owned sessions.
- Caller-managed sessions are not unexpectedly reset.
- Existing successful-call and normal tool-error behavior remains unchanged.
These tests should make the intended safety and recovery boundaries explicit and reviewable.
5. Proposed contribution and coordination
Since I have already prepared and validated the initial fix, I would be glad to submit and maintain a PR based on the existing implementation, incorporating the additional regression coverage and any adjustments requested during review.
I hope this can provide a useful starting point for resolving the issue while avoiding duplicated implementation effort.
Yufeng He (@he-yufeng), thank you for independently confirming the behavior and highlighting the compatibility considerations with #2884. Your feedback on the proposed patch would be very welcome.
gomission.ai (@gomission), thank you for clarifying the annotation trust boundary and the distinction between recovery and replay.
Eduard van Valkenburg (@eavanvalkenburg), could you please advise whether the no-automatic-replay default aligns with the intended framework behavior.
I am happy to adjust the implementation and tests to align with the maintainers' preferred design and contribution workflow.
Thanks everyone for the constructive discussion and guidance.
Metadata
Metadata
Labels
Type
Projects
- StatusShow more project fieldsNo status
Observed Behavior
The Python MCPTool implementation automatically reconnects and retries
session.call_tool()when an MCP connection fails withanyio.ClosedResourceErroror anMcpErrorindicating that the session has terminated.This can result in duplicate execution of non-idempotent MCP tools.
An MCP server may successfully perform an operation before the client receives its response. If the connection closes after execution but before the response reaches the client, the framework cannot determine whether the operation succeeded.
The existing retry logic reconnects and issues the same
tools/callrequest again, potentially repeating the operation.Example:
ClosedResourceError.The same issue can affect payment processing, database writes, ticket creation, email delivery, and other operations with side effects.
Affected code:
python/packages/core/agent_framework/_mcp.pyMCPTool._call_tool_with_retries()The original implementation retries ordinary
tools/calloperations without knowing whether the previous execution completed.This is a correctness and data-integrity issue, rather than simply a connection-recovery failure.
Expected Behavior
When an MCP
tools/callrequest fails because the connection was lost and the remote execution outcome cannot be determined:ToolExecutionExceptionclearly indicating that the remote execution outcome is unknown.The framework should avoid introducing at-least-once execution semantics for ordinary MCP tool calls without an explicit idempotency guarantee.
Steps to Reproduce
MCPTool._call_tool_with_retries().call_tool()invocation performs the side effect and then raisesanyio.ClosedResourceError, simulating a lost response.CallToolResulton the next invocation.connect(reset=True)to restore the connection without clearing the simulated server state.await tool.call_tool("charge_customer").Actual result: The simulated remote side effect executes twice for a single logical tool invocation.
Expected result: The remote operation is invoked only once, and the caller receives an exception indicating that the execution outcome is unknown.
The regression test
test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnectreproduces the failure scenario using a simulated server-side side effect.Minimal Reproduction
The following is a minimal mock-based reproduction of the lost-response scenario. It requires the Python MCP dependencies and an Agent Framework revision containing the original automatic retry implementation.
Behavior on the original implementation:
Behavior expected after the fix:
The invocation raises
ToolExecutionExceptioninstead of replaying the request, and both the side-effect counter and MCP call count remain at 1.This reproduction simulates a lost response; it does not require a live MCP server.
Error Messages and Stack Traces
Package Versions
agent-framework-core==1.21.0 mcp==1.30.0 Source: microsoft/agent-framework (local development checkout) Baseline commit: 91ab44f
Python Version
Python 3.14.8
Operating System
Linux
Regression
Unknown
Additional Context
Root cause
The original
MCPTool._call_tool_with_retries()performs up to twosession.call_tool()attempts after connection loss.The retry assumes that a failed client request implies that the remote operation did not complete. That assumption is unsafe for non-idempotent operations.
Suggested resolution
Avoid automatic replay of ordinary MCP
tools/callrequests when the remote execution outcome is unknown.Return a
ToolExecutionExceptionthat preserves the original exception and clearly identifies the uncertain outcome.Connection recovery for subsequent independent operations should remain separate from retrying the failed operation.
Validation
A local fix and regression tests have been implemented.
git diff --check: passed.The regression test verifies that a completed simulated server-side operation is not executed again after its response is lost.
The fix also updates existing MCP reconnection tests to reflect the no-replay behavior.
Local fix commit:
c10b7ef28Compatibility consideration
This changes the previous automatic-retry behavior for ordinary MCP tool calls. Maintainer feedback is welcome on the desired retry and reconnection policy, particularly for idempotent tools and caller-managed MCP sessions.
Acknowledgements