Skip to content

Streamable HTTP session lifecycle can retain unreachable sessions and serialize session creation behind slow requests #3605

Description

@fegloff

Summary

While implementing a real MCP readiness probe against the Python SDK's stateful Streamable HTTP transport, we found several related session-lifecycle behaviors that can leave server-side sessions retained even though no client can use or close them.

Observed with mcp==1.27.1.

The cases are especially visible for short-lived probes, but they apply to ordinary clients as well.

Observed behaviors

1. Closed sessions remain retained

After a client closes a session with DELETE, the session is closed but remains present in the session manager's internal session registry.

For a long-running server with repeated short-lived sessions, closed entries can therefore accumulate unless the embedding application adds its own cleanup.

2. A session can be registered before the initiating request is ultimately accepted

A new session may be registered server-side before later request validation rejects the request.

We observed this around cases such as an invalid Host request. Similar orphaning is possible when initialization is interrupted after the server has registered the session but before the client receives the session id.

The important lifecycle problem is:

  1. the server creates/registers a session;
  2. the request fails or is interrupted before the client learns the session id;
  3. the client therefore cannot address that session;
  4. the client cannot send DELETE;
  5. the server retains a session that no client can clean up.

This also matters when a readiness probe times out or is killed during initialization.

3. Session creation can be serialized behind a slow request

The session-creation path uses shared synchronization such that a slow/stalled opening request can hold the creation path long enough to delay unrelated clients attempting to create sessions.

A client that uploads or stalls slowly during this path can therefore affect other clients' ability to initialize.

Expected behavior

Ideally the SDK would provide lifecycle semantics such that:

  • a session successfully closed with DELETE is removed from the session registry when it is no longer needed;
  • if an opening request creates/registers a session but fails before the client can receive/use its session id, that session is automatically closed and removed;
  • interrupted initialization does not leave indefinitely retained unreachable sessions;
  • a slow opening request does not unnecessarily serialize unrelated session creation;
  • applications do not need to infer whether a session was ever successfully exposed to a client in order to garbage-collect it safely.

If intentional retention is required for protocol reasons, an SDK-owned bounded cleanup/reaping mechanism would avoid each embedding server having to implement its own session lifecycle policy.

Current application workaround

Our server currently distinguishes sessions that a client has demonstrated knowledge of from sessions whose id was never presented back by any request.

For an unclaimed session, a grace period starts only after the request that opened it finishes. If no request ever presents that session id, a later request can close/remove the abandoned session after the grace period.

Once any request presents the session id, the session is treated as claimed and is not subject to this orphan cleanup, even after a long idle period.

This avoids expiring legitimate long-lived MCP sessions while bounding sessions that became unreachable during initialization.

We also added a regression test specifically detecting the SDK's retention of closed sessions so the workaround can be removed when upstream lifecycle behavior changes.

This is only orphan cleanup; it is not an authentication or resource-quota mechanism. A caller that can reach an unauthenticated MCP endpoint can intentionally establish and keep valid sessions open.

Why this matters for readiness probes

A useful readiness check should exercise the real MCP surface:

initialize -> initialized -> tools/list -> DELETE

Normally that cleans up correctly. But the probe process may be timed out/killed, or initialization may be rejected after the server has already registered state. In those cases the client may never receive the session id and therefore cannot perform cleanup itself.

That turns routine readiness probing into a source of retained server-side state unless the application compensates for the SDK lifecycle behavior.

Environment

  • Python MCP SDK: mcp==1.27.1
  • Stateful Streamable HTTP session manager
  • Reproduced with local deterministic tests around abandoned initialization, rejected requests, delayed opening requests, explicit DELETE, and long-lived claimed sessions.

I can provide a minimal reproduction or point to the application-side tests/workaround if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    v1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions