Skip to content

Python: [Bug]: MCP tools/call retry can duplicate non-idempotent side effects after connection loss #9204

Description

Observed Behavior

The Python MCPTool implementation automatically reconnects and retries session.call_tool() when an MCP connection fails with anyio.ClosedResourceError or an McpError indicating that the session has terminated.

This can result in duplicate execution of non-idempotent MCP tools.

An MCP server may successfully perform an operation before the client receives its response. If the connection closes after execution but before the response reaches the client, the framework cannot determine whether the operation succeeded.

The existing retry logic reconnects and issues the same tools/call request again, potentially repeating the operation.

Example:

  1. An agent invokes an MCP tool to charge a customer.
  2. The MCP server successfully processes the charge.
  3. The connection closes before the response reaches the client.
  4. The framework catches ClosedResourceError.
  5. The framework reconnects and automatically invokes the same tool again.
  6. The customer may be charged twice.

The same issue can affect payment processing, database writes, ticket creation, email delivery, and other operations with side effects.

Affected code:

python/packages/core/agent_framework/_mcp.py

MCPTool._call_tool_with_retries()

The original implementation retries ordinary tools/call operations without knowing whether the previous execution completed.

This is a correctness and data-integrity issue, rather than simply a connection-recovery failure.

Expected Behavior

When an MCP tools/call request fails because the connection was lost and the remote execution outcome cannot be determined:

  1. The framework should not automatically replay the potentially non-idempotent operation.
  2. The caller should receive a ToolExecutionException clearly indicating that the remote execution outcome is unknown.
  3. The original connection error should remain available as the underlying exception.
  4. Applications should be able to explicitly reconnect and initiate a new operation according to their own retry or reconciliation policy.
  5. Successful tool calls and normal MCP tool-error responses should retain their existing behavior.

The framework should avoid introducing at-least-once execution semantics for ordinary MCP tool calls without an explicit idempotency guarantee.

Steps to Reproduce

  1. Use the Python Microsoft Agent Framework implementation containing MCPTool._call_tool_with_retries().
  2. Configure an MCP tool with a non-idempotent side effect, such as incrementing a counter or creating a record.
  3. Mock the MCP session so that the first call_tool() invocation performs the side effect and then raises anyio.ClosedResourceError, simulating a lost response.
  4. Configure the mock session to return a successful CallToolResult on the next invocation.
  5. Mock connect(reset=True) to restore the connection without clearing the simulated server state.
  6. Invoke the tool once using await tool.call_tool("charge_customer").
  7. Observe that the original implementation automatically retries the tool after reconnecting.

Actual result: The simulated remote side effect executes twice for a single logical tool invocation.

Expected result: The remote operation is invoked only once, and the caller receives an exception indicating that the execution outcome is unknown.

The regression test test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnect reproduces the failure scenario using a simulated server-side side effect.

Minimal Reproduction

The following is a minimal mock-based reproduction of the lost-response scenario. It requires the Python MCP dependencies and an Agent Framework revision containing the original automatic retry implementation.

import asyncio
from unittest.mock import AsyncMock, MagicMock, patch

from anyio import ClosedResourceError
from mcp.types import CallToolResult

from agent_framework._mcp import MCPStdioTool


async def main():
    tool = MCPStdioTool(
        name="duplicate_execution_test",
        command="test_command",
        load_tools=True,
    )

    session = MagicMock()
    side_effect_count = 0

    async def server_call_tool(*args, **kwargs):
        nonlocal side_effect_count

        # Simulate a completed server-side operation.
        side_effect_count += 1

        if side_effect_count == 1:
            # Server executed the operation, but response was lost.
            raise ClosedResourceError()

        return CallToolResult(content=[])

    session.call_tool = AsyncMock(side_effect=server_call_tool)

    tool.session = session
    tool.is_connected = True
    tool._tools_loaded = True

    async def reconnect(*, reset=False):
        assert reset is True
        tool.session = session
        tool.is_connected = True

    with patch.object(tool, "connect", side_effect=reconnect):
        await tool.call_tool("charge_customer")

    print("Remote side effects:", side_effect_count)
    print("MCP call count:", session.call_tool.await_count)

    assert side_effect_count == 1, "Duplicate execution detected!"


asyncio.run(main())

Behavior on the original implementation:

Remote side effects: 2
MCP call count: 2
AssertionError: Duplicate execution detected!

Behavior expected after the fix:

The invocation raises ToolExecutionException instead of replaying the request, and both the side-effect counter and MCP call count remain at 1.

This reproduction simulates a lost response; it does not require a live MCP server.

Error Messages and Stack Traces

No underlying MCP server error is required to reproduce this bug.

The simulated transport raises:


anyio.ClosedResourceError


The original implementation catches this exception, reconnects, and retries the same `tools/call` operation.

The resulting failure is a duplicate remote side effect, which may occur without an application-visible exception if the second request succeeds.

The mock reproduction reports:


Remote side effects: 2
MCP call count: 2
AssertionError: Duplicate execution detected!

Package Versions

agent-framework-core==1.21.0 mcp==1.30.0 Source: microsoft/agent-framework (local development checkout) Baseline commit: 91ab44f

Python Version

Python 3.14.8

Operating System

Linux

Regression

Unknown

Additional Context

Root cause

The original MCPTool._call_tool_with_retries() performs up to two session.call_tool() attempts after connection loss.

The retry assumes that a failed client request implies that the remote operation did not complete. That assumption is unsafe for non-idempotent operations.

Suggested resolution

Avoid automatic replay of ordinary MCP tools/call requests when the remote execution outcome is unknown.

Return a ToolExecutionException that preserves the original exception and clearly identifies the uncertain outcome.

Connection recovery for subsequent independent operations should remain separate from retrying the failed operation.

Validation

A local fix and regression tests have been implemented.

  • Full MCP test suite: 395 passed, 2 skipped.
  • Focused duplicate-execution regression test: passed.
  • Ruff lint: passed.
  • Ruff formatting: passed.
  • git diff --check: passed.

The regression test verifies that a completed simulated server-side operation is not executed again after its response is lost.

The fix also updates existing MCP reconnection tests to reflect the no-replay behavior.

Local fix commit: c10b7ef28

Compatibility consideration

This changes the previous automatic-retry behavior for ordinary MCP tool calls. Maintainer feedback is welcome on the desired retry and reconnection policy, particularly for idempotent tools and caller-managed MCP sessions.

Acknowledgements

  • I searched existing issues and did not find a duplicate.
  • I personally verified this behavior and the reproduction details are authentic.
  • I will wait for explicit maintainer agreement before starting implementation of a non-trivial change.

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 8, 2026
  2. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Oct 8, 2026
  3. github-actions commented on Oct 8, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    Repro: python/packages/core/agent_framework/_mcp.py::MCPTool._call_tool_with_retries around lines 2875-2909 replays ordinary tools/call requests after ClosedResourceError or a session-terminated McpError. The trigger is a server completing a non-idempotent operation before the response connection closes. Minimal repro: mock session.call_tool() to increment persistent state then raise ClosedResourceError on its first call and succeed on its second; one tool.call_tool() invocation increments the state twice.

    • Failing test: python/packages/core/tests/core/test_mcp.py::test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnect
    • Files examined: python/packages/core/agent_framework/_mcp.py, python/packages/core/tests/core/test_mcp.py, python/packages/core/agent_framework/exceptions.py, python/packages/core/pyproject.toml
    • Tests run: test_mcp_tool_reconnects_after_session_terminated_error, test_mcp_tool_connection_properly_invalidated_after_closed_resource_error, test_mcp_regular_tool_call_does_not_duplicate_side_effect_after_disconnect
    • Reported version: 1.21.0
    • Current version: 1.21.0
  4. added
    agentsUsage: [Issues, PRs], Target: Single agent
    mcpUsage: [Issues, PRs], Target: MCP
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 8, 2026
  5. gomission commented on Oct 8, 2026

    @gomission

    One retry-policy detail: I would not use a remote idempotentHint alone to authorize replay after a lost response. MCP tool annotations are hints, not guarantees of runtime behavior; the failure here is precisely that the server may have committed before the response was lost. See the MCP guidance on tool annotations: https://blog.modelcontextprotocol.io/posts/2026-03-16-tool-annotations/

    No automatic replay on that ambiguous path seems like the right default. If opt-in retries are added later, I would gate them on client-owned trust policy or a provider idempotency/reconciliation contract, not the annotation alone.

    A useful availability control alongside the existing lost-response test: after the first call commits and loses its response, assert one effect and one tools/call plus the unknown-outcome exception; then issue a distinct second call and assert that it reconnects and succeeds without reissuing the first. That tests recovery separately from replay.

  6. suhasagg commented on Oct 8, 2026

    @suhasagg
    ContributorAuthor

    Thanks gomission.ai (@gomission) for the detailed feedback.

    I agree with the key distinction: preventing replay of an operation with an unknown outcome and restoring connection availability for future operations are separate concerns.

    1. Retry safety and execution semantics

    The underlying issue is that a transport failure does not necessarily imply that the remote MCP operation failed.

    If the server commits a non-idempotent operation but the response is lost, automatically replaying the same tools/call request can produce duplicate side effects.

    I agree that idempotentHint alone should not be treated as sufficient authorization for replay. Tool annotations are advisory and do not provide a guaranteed deduplication or exactly-once execution contract.

    My proposed default behavior is therefore:

    • Do not automatically replay a tools/call request when its remote execution outcome is unknown.
    • Propagate a ToolExecutionException that clearly communicates the uncertain outcome and preserves the underlying transport exception.
    • Keep retry and reconciliation decisions under application control unless an explicit, trustworthy idempotency contract exists.
    • Preserve existing behavior for successful calls and ordinary MCP tool-error responses.

    This does not guarantee exactly-once execution across distributed systems; it specifically prevents the framework from introducing duplicate execution through an automatic replay.

    2. Connection recovery must remain independent of replay

    I also agree with your suggestion to validate availability after the ambiguous failure.

    The local fix currently prevents replay and supports explicit reconnection. However, it does not yet establish automatic reconnection for a subsequent independent tool call.

    There are two possible recovery policies:

    Option A — Explicit recovery: The failed operation raises an unknown-outcome exception, and the application explicitly reconnects before initiating another operation.

    Option B — Recovery on the next independent call: The failed operation raises an unknown-outcome exception without replay, while a later, distinct tool invocation can automatically restore the connection and execute normally.

    I lean toward Option B for usability, provided reconnection can be implemented safely without interfering with concurrent calls, connection lifecycle ownership, or caller-managed MCP sessions.

    Importantly, reconnecting must never implicitly replay the original operation.

    3. Proposed regression coverage

    Once the preferred recovery policy is confirmed, I can extend the tests to verify the complete failure-and-recovery sequence:

    1. The first tools/call produces a simulated server-side effect and then raises ClosedResourceError, representing a lost response.
    2. The original call is attempted exactly once, with no automatic replay.
    3. The caller receives a ToolExecutionException indicating that the remote execution outcome is unknown.
    4. The underlying connection exception remains available.
    5. A distinct second tool invocation is initiated.
    6. The second invocation succeeds after the appropriate reconnection process.
    7. The first operation's side-effect count remains exactly one.

    I would also retain coverage for session-terminated McpError, successful tool responses, and normal MCP tool errors.

    4. Current implementation and validation

    I have prepared a local fix, but have not opened a PR.

    Local commit: c10b7ef28

    Current validation:

    • Full MCP test suite: 395 passed, 2 skipped
    • Focused duplicate-side-effect regression test: passed
    • Ruff lint and formatting: passed
    • git diff --check: passed

    The implementation intentionally avoids introducing a new automatic reconnection policy until the expected lifecycle behavior is agreed upon.

    5. Maintainer guidance requested

    Eduard van Valkenburg (@eavanvalkenburg) — could you please confirm whether Option A (explicit reconnection) or Option B (automatic recovery on a subsequent independent call) is the preferred behavior for Python MCP tools?

    It would also be helpful to confirm whether retrying a potentially non-idempotent tools/call after an ambiguous transport failure should be disabled by default, with any future opt-in retry policy requiring stronger guarantees than MCP tool annotations alone.

    I am happy to adjust the implementation and regression tests to align with the framework's intended semantics.

    I will wait for maintainer agreement on the recovery policy before opening the PR.

    Thanks again for the review and suggestions.

  7. gomission commented on Oct 8, 2026

    @gomission

    Thanks, that distinction helps. I recommend Option B when the framework owns the session, while keeping caller-managed sessions under Option A.

    After an ambiguous call, return the unknown-outcome exception and never replay it. A later, independent tool call can establish a fresh session and proceed. The reconnect should be serialized with connection state so it cannot race active calls or reset a caller-owned session.

    The regression should assert one side effect and one tools/call for the ambiguous operation, then a separate second call that reconnects and succeeds. That restores availability without implying the first operation failed or promising exactly-once execution.

  8. he-yufeng commented on Oct 9, 2026

    @he-yufeng
    Contributor

    Verified on current main (_call_tool_with_retries): the first ClosedResourceError / session-terminated McpError triggers connect(reset=True) and an unconditional replay of the same tools/call. Since the error can surface after the server executed the request but before the response arrived, the replay really can run a non-idempotent operation twice. This also sits in direct tension with the #2884 fix, which added the optimistic reconnect-and-retry to ride over transient disconnects.

    I'd like to take this. Proposal that keeps both properties: gate the replay on the tool's advertised ToolAnnotations. Servers can declare idempotentHint / readOnlyHint; for those tools a replay after connection loss is safe and the #2884 behavior stays. For everything else the call is not replayed: call_tool raises ToolExecutionException stating that the connection was lost, the remote execution outcome is unknown, and the call was deliberately not retried, with the original connection error as the inner exception so applications can reconnect and reconcile under their own policy. tools/list, get_prompt, and the long-running task path keep their current retry behavior since reads and prompts have no side effects.

    Regression coverage: the issue's duplication repro (state increments once, exception surfaces), the default no-replay path, and the annotated-idempotent path still reconnecting.

  9. gomission commented on Oct 10, 2026

    @gomission

    Yufeng He (@he-yufeng), the no-replay path and preserved connection exception fit the failure here. One boundary to keep explicit in the annotated path: the current MCP ToolAnnotations contract says annotations are advisory and must not drive tool-use decisions for untrusted servers. Advertising readOnlyHint or idempotentHint should therefore not, by itself, enable replay after a lost response.

    I would keep no replay as the default and make any exception client-owned opt-in for a trusted server/tool contract. This also matches the issue author's proposed policy.

    One additional negative regression would make that boundary reviewable: a server advertises idempotentHint=True (or readOnlyHint=True), increments state, then loses the response; without client opt-in, assert one invocation/one effect and the unknown-outcome exception. Keep the separate later-call recovery control alongside it. The positive replay case can explicitly configure trusted policy and a genuinely idempotent fake operation, instead of treating the annotation as that policy.

  10. suhasagg commented on Oct 10, 2026

    @suhasagg
    ContributorAuthor

    Thanks Yufeng He (@he-yufeng) and gomission.ai (@gomission) for the detailed investigation and thoughtful feedback.

    I appreciate the independent reproduction of the issue and the discussion around the reconnect-and-retry behavior introduced in #2884. The points raised about tool annotations, connection recovery, and session ownership are important for arriving at a safe and maintainable fix.

    I'd like to share the current implementation status and propose a path forward.

    1. Existing implementation and validation

    As the original reporter of #9204, I have already prepared a local fix for the duplicate-execution scenario and completed initial regression testing.

    Local implementation: c10b7ef28

    Previously completed validation:

    • Full MCP test suite: 395 passed, 2 skipped
    • Duplicate-side-effect regression test: passed
    • Ruff lint and formatting: passed
    • git diff --check: passed

    The existing patch addresses the behavior in MCPTool._call_tool_with_retries() where a transport failure can cause the same tools/call request to be replayed even though the remote operation may already have completed.

    The fix prevents automatic replay in this ambiguous-outcome scenario and preserves the underlying transport exception.

    These results reflect the initial local implementation. I will rerun the relevant checks against the final patch before PR submission.

    2. Retry safety and MCP tool annotations

    I agree with gomission.ai (@gomission) that idempotentHint and readOnlyHint should not, by themselves, authorize automatic replay following an ambiguous connection failure.

    MCP tool annotations provide useful descriptive metadata, but they do not independently establish a guaranteed idempotency or deduplication contract.

    A server may have committed an operation before the client receives a transport error. In that situation, replaying the operation based solely on an annotation could retain the duplicate-execution risk described in this issue.

    My preferred default is therefore:

    • Do not automatically replay a tools/call request when its remote execution outcome is unknown.
    • Propagate an exception clearly communicating the uncertain outcome while preserving the original transport error.
    • Keep any future opt-in replay policy separate, with appropriate application-controlled trust or provider-supported idempotency guarantees.
    • Preserve existing behavior for successful calls and ordinary MCP tool-error responses.

    This approach prevents framework-induced replay without claiming exactly-once execution semantics.

    3. Connection recovery and compatibility with #2884

    I also agree that preventing replay and restoring connection availability are separate concerns.

    The behavior introduced in #2884 was intended to improve resilience to transient disconnections, so I believe the fix should preserve that objective wherever it can be done safely.

    Based on the discussion, the preferred direction appears to be:

    1. Ambiguous operation: Return an unknown-outcome exception without replaying the original request.
    2. Framework-owned sessions: Allow a subsequent independent operation to establish a fresh connection where appropriate.
    3. Caller-managed sessions: Respect external session ownership and avoid unexpected reconnection or resets.
    4. Concurrency: Ensure recovery does not interfere with active operations or introduce connection lifecycle races.
    5. Compatibility: Minimize unrelated changes to MCP behavior.

    My initial patch addresses the unsafe replay behavior. I would welcome maintainer guidance before extending the connection recovery policy, particularly where it affects session lifecycle semantics.

    4. Additional regression coverage

    I agree with the additional tests suggested in the discussion and propose incorporating the following cases:

    • A server completes an operation but the response is lost.
    • The original tools/call is attempted exactly once.
    • The caller receives an exception indicating that the remote execution outcome is unknown.
    • The original transport exception remains available.
    • An advertised idempotentHint=True or readOnlyHint=True does not independently authorize replay without an explicit trusted policy.
    • A subsequent independent invocation can recover appropriately for framework-owned sessions.
    • Caller-managed sessions are not unexpectedly reset.
    • Existing successful-call and normal tool-error behavior remains unchanged.

    These tests should make the intended safety and recovery boundaries explicit and reviewable.

    5. Proposed contribution and coordination

    Since I have already prepared and validated the initial fix, I would be glad to submit and maintain a PR based on the existing implementation, incorporating the additional regression coverage and any adjustments requested during review.

    I hope this can provide a useful starting point for resolving the issue while avoiding duplicated implementation effort.

    Yufeng He (@he-yufeng), thank you for independently confirming the behavior and highlighting the compatibility considerations with #2884. Your feedback on the proposed patch would be very welcome.

    gomission.ai (@gomission), thank you for clarifying the annotation trust boundary and the distinction between recovery and replay.

    Eduard van Valkenburg (@eavanvalkenburg), could you please advise whether the no-automatic-replay default aligns with the intended framework behavior.

    I am happy to adjust the implementation and tests to align with the maintainers' preferred design and contribution workflow.

    Thanks everyone for the constructive discussion and guidance.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

agentsUsage: [Issues, PRs], Target: Single agentmcpUsage: [Issues, PRs], Target: MCPpythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions