Skip to content

Python: [Bug]: Mixed Tool Batch Applies Approval Wrapper To All Tool Calls #6385

Description

@tschokokuki

Description

Summary

When the model emits multiple tool calls in one assistant turn, and only one of those tools has approval_mode set to always_require, the framework wraps all function calls in function_approval_request.

In AG-UI this results in confirm_changes being emitted for every tool call in the batch, including tools that should execute without approval.

Impact

  • Non-sensitive tools are incorrectly blocked behind approval dialogs.
  • UI shows duplicate or extra approval prompts.
  • User flow becomes confusing and slower.

Reproduction

  1. Register two tools:
    • add_comment with approval_mode="always_require"
    • search_work_items with default approval_mode (never_require)
  2. Send a prompt that causes the model to emit both function calls in the same assistant message.
  3. Observe AG-UI events and stored messages.

Expected

Only add_comment should be wrapped as function_approval_request and mapped to confirm_changes.
search_work_items should execute directly.

Actual

Both add_comment and search_work_items are wrapped as function_approval_request, producing two confirm_changes approvals.

Root Cause

In _try_execute_function_calls, approval_needed is computed as a single boolean for the whole batch. If any function call targets a tool in approval_tools, the code returns function_approval_request for all function_call items.

Current behavior location:

  • agent_framework/_tools.py
  • Function: _try_execute_function_calls
  • Branch: if approval_needed then return Content.from_function_approval_request for every function_call in function_calls

Proposed Fix

Handle function calls per item instead of all-or-nothing by batch:

  1. Split function_calls into groups:
    • approval_calls: function_call where tool is in approval_tools
    • executable_calls: function_call where tool is known and does not require approval
    • declaration_only_calls: function_call where tool is declaration_only or additional_tool
  2. Return approval request content only for approval_calls.
  3. Execute executable_calls normally.
  4. Preserve existing declaration_only behavior.
  5. Keep terminate_on_unknown_calls behavior unchanged.

Notes

A temporary mitigation is to disable multiple tool calls at model level:

  • default_options={"allow_multiple_tool_calls": False}

This reduces incidence but does not fix the underlying batching logic.

Applied workaround:

I disallow multiple tool calls for now.
default_options={"allow_multiple_tool_calls": False}

Disclaimer: To my understanding non-critical tool calls should not be wrapped. But I may not be in the complete picture of the planned/intended approval concept of the library.

Code Sample

Error Messages / Stack Traces

What gets emitted in MessagesSnapshot:

`
messages:<...>,
"tool_calls": [        {
            "id": "f991b6dc-fa33-4065-9e3c-f7373cc023b7",
            "name": null,
            "role": "assistant",
            "content": null,
                {
                    "id": "call_ZPd6rILF7c5QoA9yqJ3jb9cj",
                    "type": "function",
                    "function": {
                        "name": "add_comment",
                        "arguments": "{\"id\": 69239, \"comment\": \"test\"}"
                    },
                    "encrypted_value": null
                },
                {
                    "id": "call_AgE9A3WI1iu3FrIb7D3ZDYpR",
                    "type": "function",
                    "function": {
                        "name": "search_work_items",
                        "arguments": "{\"work_item_type\": \"Task\", \"assigned_to\": \"@me\", \"top\": 200}"
                    },
                    "encrypted_value": null
                },
                {
                    "id": "1c54f251-4dcf-4d51-b61e-d774fd084842",
                    "type": "function",
                    "function": {
                        "name": "confirm_changes",
                        "arguments": "{\"function_name\": \"add_comment\", \"function_call_id\": \"call_ZPd6rILF7c5QoA9yqJ3jb9cj\", \"function_arguments\": {\"id\": 69239, \"comment\": \"test\"}, \"steps\": [{\"description\": \"Execute add_comment\", \"status\": \"enabled\"}]}"
                    },
                    "encrypted_value": null
                },
                {
                    "id": "2cc96972-0c2b-469f-aa3d-df24efc3336b",
                    "type": "function",
                    "function": {
                        "name": "confirm_changes",
                        "arguments": "{\"function_name\": \"search_work_items\", \"function_call_id\": \"call_AgE9A3WI1iu3FrIb7D3ZDYpR\", \"function_arguments\": {\"work_item_type\": \"Task\", \"assigned_to\": \"@me\", \"top\": 200}, \"steps\": [{\"description\": \"Execute search_work_items\", \"status\": \"enabled\"}]}"
                    },
                    "encrypted_value": null
                }
            ],
            "encrypted_value": null
        }
}

Package Versions

agent-framework-ag-ui

Python Version

No response

Additional Context

No response

Activity

  1. added theissue type on Jun 8, 2026
  2. added
    bugUsage: [Issues], Target: all issues (Legacy, prefer issue type: bug)
    on Jun 8, 2026
  3. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Jun 8, 2026
  4. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Jun 8, 2026
  5. github-actions commented on Jun 8, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    Repro: _try_execute_function_calls in python/packages/core/agent_framework/_tools.py (lines 1699-1729) uses a single boolean approval_needed for the whole batch; when any function_call targets an approval_tools entry, it wraps ALL function_calls as function_approval_request. Trigger: pass a mixed batch where one Content has name in approval_tools and another does not. Minimal repro: call _try_execute_function_calls with tools=[FunctionTool(approval_mode="always_require"), FunctionTool(approval_mode="never_require")] and function_calls containing one call to each; assert only 1 result has type function_approval_request.

    • Failing test: python/packages/core/tests/test_mixed_batch_approval_bug.py
    • Files examined: python/packages/core/agent_framework/_tools.py, python/packages/core/agent_framework/init.py, python/packages/core/tests/workflow/test_agent_executor.py
    • Tests run: python/packages/core/tests/test_mixed_batch_approval_bug.py::test_mixed_batch_only_approval_tool_gets_wrapped
    • Current version: 1.8.0
  6. rpelevin commented on Jun 9, 2026

    @rpelevin

    This looks like the right bug to keep per-call rather than per-turn.

    The important invariant is that approval classification should travel with each function call, not with the assistant message batch. A mixed batch can contain one call that needs approval and one call that can execute immediately.

    I would split the batch into per-call envelopes before producing AG-UI events:

    • tool call id;
    • tool name;
    • validated arguments digest;
    • approval mode for that tool;
    • approval request id only when approval is required;
    • execution status: pending_approval, executing, completed, denied, or cancelled.

    The non-sensitive call should not inherit the approval wrapper just because a sibling call needs it. Conversely, approving the sensitive call should not authorize any sibling call.

    Acceptance tests I would want:

    • one approval-required tool plus one default tool in the same assistant message creates one approval request, not two;
    • the default tool executes or emits normal tool lifecycle events without confirm_changes;
    • approving the gated call resumes only that tool_call_id;
    • denial/cancel produces a terminal status for only that call;
    • reordered or retried batch items cannot consume the wrong approval.

    That would make AG-UI display the real control boundary instead of a batch-level approximation.

  7. removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Jun 9, 2026
  8. eavanvalkenburg commented on Jun 9, 2026

    @eavanvalkenburg
    Member

    This is on my list, we have done a similar update in dotnet, to allow some scenario's like this.

  9. rpelevin commented on Jun 9, 2026

    @rpelevin

    That dotnet precedent sounds like the right anchor.

    For Python, I would probably make the parity test very small and explicit:

    • create two function calls in the same assistant message;
    • one tool has approval required;
    • one sibling tool does not;
    • assert only the approval-required call becomes a function_approval_request;
    • assert the non-gated sibling stays on the normal execution path;
    • assert approving or denying the gated call cannot affect the sibling call.

    The key is that the approval wrapper should be a property of the individual tool call envelope, not a property of the whole mixed batch.

    If the dotnet change already has that shape, mirroring the same regression vocabulary in Python would make the expected behavior very clear across both implementations.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugUsage: [Issues], Target: all issues (Legacy, prefer issue type: bug)pythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions