Skip to content

Python: feat(core): add max_duration_seconds bound to tool invocation loop - #7772

Merged
Eduard van Valkenburg (eavanvalkenburg) merged 12 commits into
microsoft:mainfrom
karthik-0306:fix-issue-7587
Sep 8, 2026
Merged

Eduard van Valkenburg (eavanvalkenburg) merged 12 commits into
microsoft:mainfrom
karthik-0306:fix-issue-7587

Conversation

@karthik-0306

@karthik-0306 Karthik Thota (karthik-0306) commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

Motivation & Context

Function invocation loops in FunctionInvocationLayer (_tools.py) currently allow capping LLM roundtrips via max_iterations and total function calls via max_function_calls, but lack a wall-clock time limit. Unattended or complex agent runs can execute tools repeatedly and stall for long periods without a bounded total duration.

This PR addresses #7587 by introducing max_duration_seconds to FunctionInvocationConfiguration.

The issue's motivating scenario is a single, continuous, unattended loop execution. This PR implements that core case, and additionally extends the duration bound to persist across human-approval round-trips.

Description & Review Guide

What are the major changes?

  • max_duration_seconds Config Field: Added max_duration_seconds: float | None to FunctionInvocationConfiguration (TypedDict) and normalized validation (> 0 or None).
  • Graceful Degradation Path: When max_duration_seconds is exceeded mid-loop (checked after each tool batch), further tool calls are disabled (tool_choice = "none") and the model is forced to produce a final text response, reusing the established max_function_calls degradation path.
  • Shared Precedence Logic: A single _apply_batch_limit_decision helper decides tool-disable state for both streaming and non-streaming loops, so the two paths can't independently disagree on precedence (duration → consecutive-errors → call-count).
  • Wall-Clock Budget Tracking Across Approval Round-Trips: budget_state persists in AgentSession.state (via ToolApprovalMiddleware) so duration is measured cumulatively even when a run pauses for human approval and resumes in a separate agent.run() call.

What is the impact of these changes?

  • Provides a wall-clock safeguard against runaway function invocation loops.
  • Backward-compatible: default max_duration_seconds is None (unlimited); approval-persistence logic is a no-op for callers not using ToolApprovalMiddleware.

Related Issue

Fixes #7587

Contribution Checklist

  • The code builds clean without any errors or warnings
  • All unit tests pass, and I have added new tests where possible
  • The PR follows the Contribution Guidelines
  • This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
  • This is not a breaking change.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds duration bounds and machine-readable termination reasons to Python function-invocation loops.

Changes:

  • Adds and validates max_duration_seconds.
  • Tracks stop reasons across streaming and non-streaming paths.
  • Adds tests and changelog documentation.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 5 comments.

File Description
python/packages/core/agent_framework/_tools.py Implements duration tracking and stop reasons.
python/packages/core/tests/core/test_function_invocation_logic.py Tests duration limits and termination signals.
python/CHANGELOG.md Documents the new behavior.
Suppressed comments (2)

python/packages/core/agent_framework/_tools.py:3528

  • The streaming path has the same enforcement gap: approved calls are replayed before this check, while the call-dropping/fallback logic at lines 3449-3466 recognizes only max_function_calls. Consequently, an expired approval or a provider-emitted call despite tool_choice="none" can still execute. Include duration expiry in a shared pre-execution and fallback predicate.
                if (
                    max_duration_seconds is not None
                    and (perf_counter() - budget_state["start_time"]) >= max_duration_seconds
                ):

python/packages/core/agent_framework/_tools.py:3518

  • The streaming branch also leaks the internal action name "stop" as a public stop reason. This is outside the documented value set and differs from approval-time error exhaustion, which reports completed. Use the same documented semantic reason for consecutive-error exhaustion in both paths.
                budget_state.setdefault("stop_reason", "stop")

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/packages/core/agent_framework/_tools.py Outdated
Comment thread python/packages/core/agent_framework/_tools.py Outdated
Comment thread python/packages/core/agent_framework/_tools.py Outdated
Comment thread python/packages/core/tests/core/test_function_invocation_logic.py Outdated
Comment thread python/packages/core/agent_framework/_tools.py
@github-actions github-actions Bot changed the title feat(core): add max_duration_seconds bound and stop_reason signal to tool loop (#7587) Python: feat(core): add max_duration_seconds bound and stop_reason signal to tool loop (#7587) Aug 19, 2026
Comment thread python/packages/core/agent_framework/_tools.py Outdated
Comment thread python/packages/core/agent_framework/_tools.py Outdated
- **agent-framework-core**: Refactored _apply_batch_limit_decision to compute perf_counter exactly once per decision point, eliminating the structural fragility where the threshold check and log message used separate clock samples.
- **agent-framework-core**: Re-ordered limit checking in Phase 1 to execute before approval response resolution, successfully preventing execution during approved replays when limits are reached. A post-approval check ensures consecutive error limits (�ction == stop) remain handled.
- **agent-framework-core**: Rewrote 6 tests in 	est_function_invocation_logic.py that used a fragile call_count mock. The tests now use a mutable clock array that advances directly during the tool execution semantic step, providing true robustness against internal engine refactors.

Note: The fallback response trigger (_ensure_function_invocation_limit_fallback_response) remains scoped strictly to the function call limit, preserving pre-existing behavior. Expanding this to cover consecutive errors (�ction == stop) or duration timeouts is intentionally left out of scope for this fix.
… streaming limit decision, fix double-counted approval calls against max_function_calls
Comment thread python/CHANGELOG.md
Comment thread python/packages/core/agent_framework/_agents.py Outdated
@github-actions

Copy link
Copy Markdown
Contributor

Python Test Coverage

Python Test Coverage Report •
FileStmtsMissCoverMissing
packages/core/agent_framework
   _agents.py4724490%600, 655, 1225, 1270, 1365–1369, 1468, 1498, 1535, 1630, 1658, 1671, 1720, 1722, 1731–1736, 1741, 1743, 1749–1750, 1757, 1759–1760, 1768–1769, 1772–1774, 1784–1789, 1793, 1798, 1800
   _tools.py15739593%232–233, 410, 412, 425, 450–452, 460, 478, 492, 499, 506, 529, 531, 538, 546, 681, 720–722, 730, 785–787, 813, 839, 843, 881–883, 887, 1060, 1072, 1079–1082, 1103, 1111, 1125–1127, 1520, 1605, 1718–1719, 1776, 1823, 1830–1831, 1951, 2028, 2124, 2138, 2141, 2148, 2151, 2157, 2169, 2186, 2195, 2203, 2207, 2227, 2229, 2236, 2294, 2297, 2320, 2327, 2332–2333, 2336, 2340, 2343, 2365, 2399, 2467, 2496–2497, 2594, 2622, 2662, 2665, 2722, 2816, 2928, 3037, 3210, 3213, 3223, 3240–3241, 3763
packages/core/agent_framework/_harness
   _tool_approval.py3654089%69, 72, 77, 114, 131, 134, 137, 194–195, 214, 243, 251, 262, 280–285, 297, 311, 334, 385, 409, 411–412, 444, 454, 469, 471–472, 474–475, 530–532, 584–585, 635, 644
TOTAL48219449390% 

Python Unit Test Overview

Tests Skipped Failures Errors Time
9763 36 💤 0 ❌ 0 🔥 2m 38s ⏱️

@eavanvalkenburg

Copy link
Copy Markdown
Member

Karthik Thota (@karthik-0306) I replied on the comment thread, please have a look

Comment thread python/packages/core/agent_framework/_tools.py Outdated
@eavanvalkenburg

Copy link
Copy Markdown
Member

Overal, this look good Karthik Thota (@karthik-0306) I am a bit on the fence myself on whether we should or should not include waits between runs, the pro, as you call out, is that waiting long for a approval can be part of the run, thereby controlling the overall throughpuyt, however the con of that is that we designed approval to exit the agent run altogether, because it might be offloaded to some other system and not come back until days later, and so there is a meaningful difference between two agents that runs for 3 days, where 1 waits 48 hours for a approval, and the other never waits. And what if the approval wait times range from seconds to days, then you can never really set a meaningful overall timeout. But like I said, I am on the fence... curious about your ideas for this. CC westey (@westey-m)

@karthik-0306

Copy link
Copy Markdown
Contributor Author

Overal, this look good Karthik Thota (Karthik Thota (@karthik-0306)) I am a bit on the fence myself on whether we should or should not include waits between runs, the pro, as you call out, is that waiting long for a approval can be part of the run, thereby controlling the overall throughpuyt, however the con of that is that we designed approval to exit the agent run altogether, because it might be offloaded to some other system and not come back until days later, and so there is a meaningful difference between two agents that runs for 3 days, where 1 waits 48 hours for a approval, and the other never waits. And what if the approval wait times range from seconds to days, then you can never really set a meaningful overall timeout. But like I said, I am on the fence... curious about your ideas for this. CC westey (westey (@westey-m))

Appreciate you laying out both sides rather than deciding fast — worth getting right.

Before this issue there was no wall-clock bound at all, so either direction is a net win for the framework — no stake here in defending the current behavior.

My lean: ship the simpler version now — not less work, but we don't have real data yet on how approval wait times actually distribute (seconds vs. days). Building pause/resume against a guess feels premature; building it once we see real usage doesn't.

If you'd rather go pause/resume from the start:

  • Replace the single start_time with accumulated_active_seconds + a segment_start for the current active stretch.
  • On pause (approval requested): fold the current segment into accumulated_active_seconds, clear segment_start.
  • On resume: open a new segment.
  • Elapsed = accumulated_active_seconds + current open segment, instead of now - start_time.
  • Builds on the existing session-state persistence, not a rewrite.

Whatever you and Westey land on, happy to build it.

Merged via the queue into microsoft:main with commit 0d3ea14 Sep 8, 2026
37 checks passed
@karthik-0306 Karthik Thota (karthik-0306) changed the title Python: feat(core): add max_duration_seconds bound and stop_reason signal to tool loop (#7587) Python: feat(core): add max_duration_seconds bound to tool invocation loop Sep 8, 2026

This branch was previously deployed

1 inactive deployment
github-app-auth — 6c481b04 Deployed Sep 8, 2026 by karthik-0306 via add_label #22154
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Usage: [Issues, PRs], Target: documentation in the code base and learn docs python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: [Feature]: Bound an agent run by duration (and by spend), not only by iteration and call count

4 participants