Skip to content

Python: [Bug]: Background session release waits past its finite timeout during cancellation cleanup #9231

Description

@ktz03

Observed Behavior

BackgroundAgentsProvider.release_session(session, timeout=0.03) waits for a background agent to finish deferred cancellation cleanup instead of returning at the configured timeout. A real Agent.run with an offline BaseChatClient that finishes its cleanup after 0.35 seconds makes release take about 0.35–0.36 seconds. A cooperative cancellation control returns promptly with the same timeout.

The deferred-cleanup client catches repeated CancelledError while finishing a finite cleanup operation. The provider requests cancellation, reaches its timeout, then continues waiting for that task to settle. This extends the host's session teardown beyond its configured bound. The reproduction finishes on its own; it does not run a permanently hanging task.

Expected Behavior

A finite release timeout should bound the provider's wait for cancellation, with unfinished tasks handled using the existing abandonment/exception-observation policy. Cooperative task cancellation, timeout=None, and the cancel_running=False guard should retain their behavior.

Steps to Reproduce

  1. Run the code below with current agent-framework-core source.
  2. It creates a real Agent and starts its run through the public background-task tool, first with cooperative cancellation and then with finite deferred cleanup.
  3. Both releases use timeout=0.03; only the deferred-cleanup case waits until the 0.35-second cleanup finishes.

Minimal Reproduction

import asyncio
import time

from agent_framework import (
    Agent, AgentSession, BackgroundAgentsProvider, BaseChatClient,
    ChatResponse, Message, SessionContext,
)


class OfflineClient(BaseChatClient):
    def __init__(self, defer_cleanup):
        super().__init__()
        self.defer_cleanup = defer_cleanup
        self.started = asyncio.Event()

    async def _inner_get_response(self, *, messages, options, **kwargs):
        self.started.set()
        finish_by = asyncio.get_running_loop().time() + 0.35
        cancelled = False
        while True:
            try:
                if cancelled:
                    await asyncio.sleep(max(0, finish_by - asyncio.get_running_loop().time()))
                    return ChatResponse(messages=[Message(role='assistant', contents=['cleaned up'])])
                await asyncio.sleep(10)
            except asyncio.CancelledError:
                if not self.defer_cleanup:
                    raise
                cancelled = True

    async def _inner_get_streaming_response(self, *, messages, options, **kwargs):
        raise RuntimeError('This reproduction uses nonstreaming Agent.run only.')
        yield


async def reproduce(defer_cleanup):
    client = OfflineClient(defer_cleanup)
    agent = Agent(client=client, name='worker')
    provider = BackgroundAgentsProvider([agent])
    session = AgentSession()
    context = SessionContext(session_id=session.session_id, input_messages=[])
    await provider.before_run(agent=None, session=session, context=context, state={})
    tools = {tool.name: tool for tool in context.tools}
    await tools['background_agents_start_task'].invoke(
        arguments={'agent_name': 'worker', 'input': 'start', 'description': 'cleanup'},
        skip_parsing=True,
    )
    await client.started.wait()
    start = time.monotonic()
    await provider.release_session(session, timeout=0.03)
    print({'deferred_cleanup': defer_cleanup, 'timeout': 0.03, 'elapsed': time.monotonic() - start})


async def main():
    await reproduce(False)
    await reproduce(True)


asyncio.run(main())

Representative actual output on Windows/Python 3.14.3:

{'deferred_cleanup': False, 'timeout': 0.03, 'elapsed': 0.000076}
{'deferred_cleanup': True, 'timeout': 0.03, 'elapsed': 0.360132}

Error Messages and Stack Traces

No exception is raised by release_session. After waiting for cleanup to finish, it logs:

Session release timed out waiting for 0 task(s). They will be abandoned.

The experimental-provider and offline-client function-invocation warnings are retained; no model or network client is used.

Package Versions

Executed agent-framework-core source from current main 2d9cc3f8ae465baf1aed3022f89b6848db4ffa7b (pyproject.toml version 1.21.0). All 84 core-package Python files match that commit. The existing environment has cached agent-framework-core 1.20.0 dependency metadata; this is a source-based reproduction, without a fresh workspace-lock install.

Python Version

Python 3.14.3

Operating System

Windows

Regression

Unknown

Additional Context

The current _drain_runtime wraps asyncio.gather(*pending, return_exceptions=True) in asyncio.wait_for. Its timeout cancels the gather and then waits for cancellation to finish, so the catch/abandonment branch is not reached at the deadline when the background task defers cancellation cleanup.

This follows the existing release contract introduced by PratikWayase in #7450, addressing antsok's #7385. moonbox3's review asked for a bounded shutdown policy; the accepted implementation added the timeout. The Foundry cleanup guidance in eavanvalkenburg's merged #8899 also relies on the finite provider bound. This report concerns the local release wait; the distributed runtime/lease work in #8760 and karthik-0306's #9040 remains separate.

Additional public controls exercised cooperative cancellation, finite deferred cleanup, timeout=0, timeout=None, and rejecting a release with cancel_running=False while a task is running. Both a protocol-compatible offline agent and a real Agent.run reproduce the finite-timeout discrepancy. No full core suite, external model/service, other operating system, or performance validation was run. Please confirm the scoped direction of enforcing the existing finite release wait before implementation.

AI Assistance

AI-assisted analysis, reproduction, and writing.

Acknowledgements

  • I searched existing issues and did not find a duplicate.
  • I personally verified this behavior and the reproduction details are authentic.
  • I will wait for explicit maintainer agreement before starting implementation of a non-trivial change.

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 9, 2026
  2. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Oct 9, 2026
  3. github-actions commented on Oct 9, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    Repro: python/packages/core/agent_framework/_harness/_background_agents.py::BackgroundAgentsProvider._drain_runtime at lines 405-458 exceeds a finite timeout when a tracked task repeatedly catches cancellation during finite cleanup. The public trigger is release_session(session, timeout=0.03) after starting such a task through background_agents_start_task. Minimal repro: run the added test, which defers cleanup for 0.35 seconds and observes release returning after approximately 0.351 seconds.

    • Failing test: python/packages/core/tests/core/test_harness_background_agents.py::test_release_session_timeout_bounds_deferred_cancellation_cleanup
    • Files examined: python/AGENTS.md, python/.github/skills/python-testing/SKILL.md, python/packages/core/AGENTS.md, python/packages/core/pyproject.toml, python/packages/core/agent_framework/_harness/_background_agents.py, python/packages/core/tests/core/test_harness_background_agents.py
    • Tests run: test_release_session_timeout_bounds_deferred_cancellation_cleanup, test_release_session_times_out_if_task_ignores_cancellation, test_release_session_cancels_and_clears, test_release_session_raises_if_cancel_running_false
    • Reported version: 1.21.0
    • Current version: 1.21.0
  4. added
    harness[Issues, PRs], Target: harness-level items
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 9, 2026
  5. charan-rathore commented on Oct 10, 2026

    @charan-rathore
    Contributor

    I reproduced this on current main (fd52de7) with the issue's script: release_session(session, timeout=0.03) blocked 0.351s against the deferred-cleanup client, while the cooperative-cancellation control returned promptly. Cross-checked on Linux against the Windows/Python 3.14 report - same behavior.

    One scoping note from tracing _drain_runtime: the never-settling task case is already bounded - what escapes the deadline is specifically a task that catches CancelledError and resumes a finite cleanup. wait_for(gather(...)) cancels at the deadline and then waits for the gather to settle, and a cancel-swallowing task keeps that gather pending until its cleanup completes, so the timeout stops being a bound. A deadline-based abandon (e.g. asyncio.wait with the remaining budget, then dropping and observing the stragglers under the existing abandonment policy) would keep the bound without changing the cooperative, timeout=None, or cancel_running=False paths.

    I see this is assigned to westey (@westey-m) - if there's no internal fix already in flight, I'd be glad to open a PR along those lines. Happy to hold off if you'd rather take it internally.

    Instinct assisted with the reproduction and analysis.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

harness[Issues, PRs], Target: harness-level itemspythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions