Skip to content

Python: [Bug]: Streaming tool calls from Claude deployments in Foundry are silently dropped (missing output_item.added, function_call_arguments.done ignored) #9083

Description

Description

When FoundryChatClient streams from an Anthropic Claude deployment in Azure AI Foundry (claude-opus-5), the tool calls the model makes are silently dropped. The service's Responses stream for Claude doesn't follow the format the OpenAI Responses parser expects, and the parser ignores the events that hold the complete arguments.

What Foundry sends for a single tool call (raw openai SDK, responses.create(..., stream=True), full output below):

response.output_item.added              0 'message'                                  <- no output_item.added for the function_call
response.content_part.added             0 None
response.output_text.delta              0 'I'
response.output_text.delta              0 "'ll search the documentation for"
response.output_text.delta              0 ' that.'
response.function_call_arguments.delta  0 'ut of account", "max_results": 8}'         <- only a tail fragment (sometimes '')
response.function_call_arguments.done   0 '{"query": "reset password when locked out of account", "max_results": 8}'   <- complete
response.output_item.done               0 'function_call'                            <- complete item, same output_index as the message

How agent_framework_openai handles it (_chat_client.py, streaming event parser):

  • function_call_ids (output_index → call_id/name) is only filled in from response.output_item.added, so it never gets an entry for this function call.
  • response.function_call_arguments.delta is only used when that lookup succeeds, so every delta is dropped.
  • response.function_call_arguments.done isn't handled. It gets logged as Unparsed event of type: response.function_call_arguments.done and the complete arguments are thrown away.

Actual behavior: the function call never reaches the function-invocation loop. The tool isn't invoked, and the agent returns only the model's preamble (for example "I'll search the documentation for this.") as its final answer, with no error or warning. This reproduced on 3 of 3 runs with agent.run(..., stream=True) and also when the agent is hosted with ResponsesHostServer, on both the latest packages and older ones (versions below).

Expected behavior: streaming should produce the same complete function call as the non-streaming path, and the tool should be invoked with the arguments from response.function_call_arguments.done / response.output_item.done. Possible approaches:

  • Register the call_id/name from response.output_item.done (or create the call there) when no output_item.added was seen for that item_id.
  • Treat the .done arguments as authoritative, or at least fall back to them when the joined deltas are missing or don't parse.
  • Never drop a function call silently. If it can't be rebuilt, log a warning or raise.

The non-conformant stream also needs fixing on the Foundry side (I'm reporting that separately through Azure support). But since Foundry is a first-class target for FoundryChatClient and ResponsesHostServer, the parser shouldn't quietly lose tool calls when the stream is incomplete.

Steps to reproduce

  1. Deploy claude-opus-5 (format: Anthropic, version 2, SKU GlobalStandard) in an Azure AI Foundry project.
  2. Set FOUNDRY_PROJECT_ENDPOINT and AZURE_AI_MODEL_DEPLOYMENT_NAME, then run the sample below.
  3. There are no CALL/RESULT lines in the output, so the tool was never invoked. Running the same agent without streaming (agent.run(prompt)) invokes the tool correctly.

Code Sample

import asyncio
import os
from typing import Annotated

from agent_framework import Agent, tool
from agent_framework.foundry import FoundryChatClient
from azure.identity import DefaultAzureCredential
from pydantic import Field


@tool(approval_mode="never_require")
def search_docs(
    query: Annotated[str, Field(description="The search query.")],
    max_results: Annotated[int, Field(description="Number of results.")] = 5,
) -> str:
    """Search the documentation."""
    return f"No documents found for query: {query}"


async def main() -> None:
    client = FoundryChatClient(
        project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
        model=os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"],  # Claude deployment, e.g. claude-opus-5
        credential=DefaultAzureCredential(),
    )
    agent = Agent(
        client=client,
        instructions="Always call search_docs with max_results=8 before answering.",
        tools=[search_docs],
        default_options={"store": False},
    )
    prompt = "How do I reset my password if I'm locked out of my account?"

    stream = agent.run(prompt, stream=True)
    async for _ in stream:
        pass
    response = await stream.get_final_response()
    for message in response.messages:
        for content in message.contents:
            if content.type == "function_call":
                print("CALL  ", content.name, repr(content.arguments))
            elif content.type == "function_result":
                print("RESULT", repr(content.result))
    print("TEXT  ", repr(response.text[:200]))


asyncio.run(main())

Output (3 runs, latest packages):

TEXT   "I'll search the documentation for information on password resets and account lockouts."
TEXT   "I'll look that up in the documentation."
TEXT   "I'll search the documentation for information on password resets and account lockouts."

Raw event dump without agent-framework, which shows the service-side stream above:

import asyncio, os
from openai import AsyncOpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider

tools = [{"type": "function", "name": "search_docs", "description": "Search the documentation.",
          "parameters": {"type": "object", "properties": {"query": {"type": "string"}, "max_results": {"type": "integer"}},
                         "required": ["query"]}}]

async def main():
    token = get_bearer_token_provider(DefaultAzureCredential(), "https://ai.azure.com/.default")
    client = AsyncOpenAI(base_url=os.environ["FOUNDRY_PROJECT_ENDPOINT"].rstrip("/") + "/openai/v1",
                         api_key=await asyncio.to_thread(token))
    stream = await client.responses.create(
        model=os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"], tools=tools, stream=True, store=False,
        input="Search the docs for how to reset a password when locked out. Use max_results 8.")
    async for e in stream:
        if e.type not in ("response.created", "response.in_progress", "response.completed"):
            print(e.type, getattr(e, "output_index", None), repr(getattr(e, "delta", None) or getattr(e, "arguments", None)
                  or (getattr(e, "item", None) and e.item.type)))

asyncio.run(main())

Error Messages / Stack Traces

# No error is raised. With DEBUG logging on the agent_framework logger, the complete arguments are visibly discarded:
DEBUG agent_framework.openai: Unparsed event of type: response.output_item.added: ResponseOutputItemAddedEvent(item=ResponseOutputMessage(... type='message' ...), output_index=0, ...)
DEBUG agent_framework.openai: Unparsed event of type: response.function_call_arguments.done: ResponseFunctionCallArgumentsDoneEvent(arguments='{"query": "account lockout recovery unlock", "max_results": 8}', item_id='fc_...', output_index=0, sequence_number=10, type='response.function_call_arguments.done')

Package Versions

agent-framework-core: 1.20.0, agent-framework-openai: 1.15.0, agent-framework-foundry: 1.14.0, agent-framework-foundry-hosting: 1.0.0b261002, openai: 3.24.0. Also reproduced with agent-framework-core: 1.16.0, agent-framework-openai: 1.14.1, agent-framework-foundry: 1.11.0, agent-framework-foundry-hosting: 1.0.0b260821, openai: 2.54.0.

Python Version

Python 3.13 (Linux x86_64; also in a Foundry hosted-agent container)

Additional Context

  • Second symptom from the same stream: in a hosted agent with a longer system prompt (agent-framework-core 1.16.0 under ResponsesHostServer), the call wasn't dropped. Its arguments came out as a complete JSON object with a delta fragment appended ({...valid args...}<fragment>", "<key>": <value>}), so each call failed with Error: Argument parsing failed. After three model retries the request ended with Maximum consecutive function call errors reached (3), and the user got Function invocation limit reached before a final answer could be produced. The model's own .done arguments were valid every time.
  • Related but different: Python: [Bug]: Streamed parallel tool calls get merged into the wrong call when their argument deltas interleave #8336 (interleaved parallel tool-call deltas merged into the wrong call; fixed in Python: Fix streamed parallel tool calls merging into the wrong call #8337) and Python: [Bug]: Duplicate function call in Foundry Hosted agent when streaming #7485 (duplicate declaration-only function call in Foundry Hosted streaming). Here there is only one tool call, and the problem is a missing output_item.added plus incomplete deltas, with the complete arguments only in the .done events.
  • Workaround: don't stream (agent.run(prompt)). That works, but it isn't an option behind ResponsesHostServer when the client asks for streaming.
  • Regression? No. Neither version tested handles this stream.

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 5, 2026
  2. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Oct 5, 2026
  3. github-actions commented on Oct 5, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    The bug reproduces in python/packages/openai/agent_framework_openai/_chat_client.py::_parse_chunk_from_openai when a stream has no function-call output_item.added event but later includes complete function_call_arguments.done and function-call output_item.done events. Minimal repro: feed that sequence to _parse_chunk_from_openai and assert that one complete function-call content is emitted; the current parser emits none.

    • Failing test: python/packages/openai/tests/openai/test_openai_chat_client.py::test_parse_chunk_recovers_function_call_when_added_event_is_missing
    • Files examined: python/AGENTS.md, python/packages/openai/agent_framework_openai/_chat_client.py, python/packages/openai/tests/openai/test_openai_chat_client.py, python/packages/foundry/agent_framework_foundry/_chat_client.py, python/packages/foundry/tests/foundry/test_foundry_chat_client.py, python/packages/openai/pyproject.toml, python/packages/foundry/pyproject.toml
    • Tests run: test_parse_chunk_from_openai_function_call_is_actionable, test_parse_chunk_recovers_function_call_when_added_event_is_missing
    • Reported version: agent-framework-openai 1.15.0
    • Current version: agent-framework-openai 1.15.0
  4. added
    agentsUsage: [Issues, PRs], Target: Single agent
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 5, 2026
  5. Rabib001 commented on Oct 6, 2026

    @Rabib001

    I'd like to pick this one up if the core team is ok with it (the function-calling spec asks outside contributors to check first, so asking before I write anything).

    I went through _parse_chunk_from_openai on main and it lines up with the report:

    • function_call_ids only gets filled from response.output_item.added, so without that event every function_call_arguments.delta is dropped
    • response.function_call_arguments.done isn't handled and ends up in the "Unparsed event" debug log
    • response.output_item.done has no function_call branch, so the complete item gets thrown away too

    What I had in mind:

    • add a function_call branch to response.output_item.done that emits the call from the done item, built the same way as the non-streaming path (call_id, name, arguments, fc_id)
    • only do that when the call wasn't already streamed through .added + deltas, checked by call_id rather than output_index since Foundry reuses index 0 for the message here. Normal OpenAI streams wouldn't change and wouldn't get duplicated arguments
    • FoundryChatClient passes everything through the base parser, so the fix only touches the openai package

    Tests: the event order from this issue produces one complete call that actually gets invoked, a normal .added + deltas + .done stream still produces exactly one call, and the streamed result matches the non-streaming one for the same item.

    Does that sound like the right direction? And are there other rows in the spec matrix you'd want covered?

  6. self-assigned this
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

agentsUsage: [Issues, PRs], Target: Single agentpythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions