You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Python: [Bug]: Streaming tool calls from Claude deployments in Foundry are silently dropped (missing output_item.added, function_call_arguments.done ignored) #9083
When FoundryChatClient streams from an Anthropic Claude deployment in Azure AI Foundry (claude-opus-5), the tool calls the model makes are silently dropped. The service's Responses stream for Claude doesn't follow the format the OpenAI Responses parser expects, and the parser ignores the events that hold the complete arguments.
What Foundry sends for a single tool call (raw openai SDK, responses.create(..., stream=True), full output below):
response.output_item.added 0 'message' <- no output_item.added for the function_call
response.content_part.added 0 None
response.output_text.delta 0 'I'
response.output_text.delta 0 "'ll search the documentation for"
response.output_text.delta 0 ' that.'
response.function_call_arguments.delta 0 'ut of account", "max_results": 8}' <- only a tail fragment (sometimes '')
response.function_call_arguments.done 0 '{"query": "reset password when locked out of account", "max_results": 8}' <- complete
response.output_item.done 0 'function_call' <- complete item, same output_index as the message
How agent_framework_openai handles it (_chat_client.py, streaming event parser):
function_call_ids (output_index → call_id/name) is only filled in from response.output_item.added, so it never gets an entry for this function call.
response.function_call_arguments.delta is only used when that lookup succeeds, so every delta is dropped.
response.function_call_arguments.done isn't handled. It gets logged as Unparsed event of type: response.function_call_arguments.done and the complete arguments are thrown away.
Actual behavior: the function call never reaches the function-invocation loop. The tool isn't invoked, and the agent returns only the model's preamble (for example "I'll search the documentation for this.") as its final answer, with no error or warning. This reproduced on 3 of 3 runs with agent.run(..., stream=True) and also when the agent is hosted with ResponsesHostServer, on both the latest packages and older ones (versions below).
Expected behavior: streaming should produce the same complete function call as the non-streaming path, and the tool should be invoked with the arguments from response.function_call_arguments.done / response.output_item.done. Possible approaches:
Register the call_id/name from response.output_item.done (or create the call there) when no output_item.added was seen for that item_id.
Treat the .done arguments as authoritative, or at least fall back to them when the joined deltas are missing or don't parse.
Never drop a function call silently. If it can't be rebuilt, log a warning or raise.
The non-conformant stream also needs fixing on the Foundry side (I'm reporting that separately through Azure support). But since Foundry is a first-class target for FoundryChatClient and ResponsesHostServer, the parser shouldn't quietly lose tool calls when the stream is incomplete.
Steps to reproduce
Deploy claude-opus-5 (format: Anthropic, version 2, SKU GlobalStandard) in an Azure AI Foundry project.
Set FOUNDRY_PROJECT_ENDPOINT and AZURE_AI_MODEL_DEPLOYMENT_NAME, then run the sample below.
There are no CALL/RESULT lines in the output, so the tool was never invoked. Running the same agent without streaming (agent.run(prompt)) invokes the tool correctly.
Code Sample
importasyncioimportosfromtypingimportAnnotatedfromagent_frameworkimportAgent, toolfromagent_framework.foundryimportFoundryChatClientfromazure.identityimportDefaultAzureCredentialfrompydanticimportField@tool(approval_mode="never_require")defsearch_docs(
query: Annotated[str, Field(description="The search query.")],
max_results: Annotated[int, Field(description="Number of results.")] =5,
) ->str:
"""Search the documentation."""returnf"No documents found for query: {query}"asyncdefmain() ->None:
client=FoundryChatClient(
project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
model=os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"], # Claude deployment, e.g. claude-opus-5credential=DefaultAzureCredential(),
)
agent=Agent(
client=client,
instructions="Always call search_docs with max_results=8 before answering.",
tools=[search_docs],
default_options={"store": False},
)
prompt="How do I reset my password if I'm locked out of my account?"stream=agent.run(prompt, stream=True)
asyncfor_instream:
passresponse=awaitstream.get_final_response()
formessageinresponse.messages:
forcontentinmessage.contents:
ifcontent.type=="function_call":
print("CALL ", content.name, repr(content.arguments))
elifcontent.type=="function_result":
print("RESULT", repr(content.result))
print("TEXT ", repr(response.text[:200]))
asyncio.run(main())
Output (3 runs, latest packages):
TEXT "I'll search the documentation for information on password resets and account lockouts."
TEXT "I'll look that up in the documentation."
TEXT "I'll search the documentation for information on password resets and account lockouts."
Raw event dump without agent-framework, which shows the service-side stream above:
importasyncio, osfromopenaiimportAsyncOpenAIfromazure.identityimportDefaultAzureCredential, get_bearer_token_providertools= [{"type": "function", "name": "search_docs", "description": "Search the documentation.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}, "max_results": {"type": "integer"}},
"required": ["query"]}}]
asyncdefmain():
token=get_bearer_token_provider(DefaultAzureCredential(), "https://ai.azure.com/.default")
client=AsyncOpenAI(base_url=os.environ["FOUNDRY_PROJECT_ENDPOINT"].rstrip("/") +"/openai/v1",
api_key=awaitasyncio.to_thread(token))
stream=awaitclient.responses.create(
model=os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"], tools=tools, stream=True, store=False,
input="Search the docs for how to reset a password when locked out. Use max_results 8.")
asyncforeinstream:
ife.typenotin ("response.created", "response.in_progress", "response.completed"):
print(e.type, getattr(e, "output_index", None), repr(getattr(e, "delta", None) orgetattr(e, "arguments", None)
or (getattr(e, "item", None) ande.item.type)))
asyncio.run(main())
Error Messages / Stack Traces
# No error is raised. With DEBUG logging on the agent_framework logger, the complete arguments are visibly discarded:
DEBUG agent_framework.openai: Unparsed event of type: response.output_item.added: ResponseOutputItemAddedEvent(item=ResponseOutputMessage(... type='message' ...), output_index=0, ...)
DEBUG agent_framework.openai: Unparsed event of type: response.function_call_arguments.done: ResponseFunctionCallArgumentsDoneEvent(arguments='{"query": "account lockout recovery unlock", "max_results": 8}', item_id='fc_...', output_index=0, sequence_number=10, type='response.function_call_arguments.done')
Python 3.13 (Linux x86_64; also in a Foundry hosted-agent container)
Additional Context
Second symptom from the same stream: in a hosted agent with a longer system prompt (agent-framework-core 1.16.0 under ResponsesHostServer), the call wasn't dropped. Its arguments came out as a complete JSON object with a delta fragment appended ({...valid args...}<fragment>", "<key>": <value>}), so each call failed with Error: Argument parsing failed. After three model retries the request ended with Maximum consecutive function call errors reached (3), and the user got Function invocation limit reached before a final answer could be produced. The model's own .done arguments were valid every time.
🤖 Automated triage reproduction notes(agent-authored — trust but verify)
Agent analysis
The bug reproduces in python/packages/openai/agent_framework_openai/_chat_client.py::_parse_chunk_from_openai when a stream has no function-call output_item.added event but later includes complete function_call_arguments.done and function-call output_item.done events. Minimal repro: feed that sequence to _parse_chunk_from_openai and assert that one complete function-call content is emitted; the current parser emits none.
I'd like to pick this one up if the core team is ok with it (the function-calling spec asks outside contributors to check first, so asking before I write anything).
I went through _parse_chunk_from_openai on main and it lines up with the report:
function_call_ids only gets filled from response.output_item.added, so without that event every function_call_arguments.delta is dropped
response.function_call_arguments.done isn't handled and ends up in the "Unparsed event" debug log
response.output_item.done has no function_call branch, so the complete item gets thrown away too
What I had in mind:
add a function_call branch to response.output_item.done that emits the call from the done item, built the same way as the non-streaming path (call_id, name, arguments, fc_id)
only do that when the call wasn't already streamed through .added + deltas, checked by call_id rather than output_index since Foundry reuses index 0 for the message here. Normal OpenAI streams wouldn't change and wouldn't get duplicated arguments
FoundryChatClient passes everything through the base parser, so the fix only touches the openai package
Tests: the event order from this issue produces one complete call that actually gets invoked, a normal .added + deltas + .done stream still produces exactly one call, and the streamed result matches the non-streaming one for the same item.
Does that sound like the right direction? And are there other rows in the spec matrix you'd want covered?
Description
When
FoundryChatClientstreams from an Anthropic Claude deployment in Azure AI Foundry (claude-opus-5), the tool calls the model makes are silently dropped. The service's Responses stream for Claude doesn't follow the format the OpenAI Responses parser expects, and the parser ignores the events that hold the complete arguments.What Foundry sends for a single tool call (raw
openaiSDK,responses.create(..., stream=True), full output below):How
agent_framework_openaihandles it (_chat_client.py, streaming event parser):function_call_ids(output_index→call_id/name) is only filled in fromresponse.output_item.added, so it never gets an entry for this function call.response.function_call_arguments.deltais only used when that lookup succeeds, so every delta is dropped.response.function_call_arguments.doneisn't handled. It gets logged asUnparsed event of type: response.function_call_arguments.doneand the complete arguments are thrown away.Actual behavior: the function call never reaches the function-invocation loop. The tool isn't invoked, and the agent returns only the model's preamble (for example
"I'll search the documentation for this.") as its final answer, with no error or warning. This reproduced on 3 of 3 runs withagent.run(..., stream=True)and also when the agent is hosted withResponsesHostServer, on both the latest packages and older ones (versions below).Expected behavior: streaming should produce the same complete function call as the non-streaming path, and the tool should be invoked with the arguments from
response.function_call_arguments.done/response.output_item.done. Possible approaches:call_id/namefromresponse.output_item.done(or create the call there) when nooutput_item.addedwas seen for thatitem_id..donearguments as authoritative, or at least fall back to them when the joined deltas are missing or don't parse.The non-conformant stream also needs fixing on the Foundry side (I'm reporting that separately through Azure support). But since Foundry is a first-class target for
FoundryChatClientandResponsesHostServer, the parser shouldn't quietly lose tool calls when the stream is incomplete.Steps to reproduce
claude-opus-5(format: Anthropic, version 2, SKU GlobalStandard) in an Azure AI Foundry project.FOUNDRY_PROJECT_ENDPOINTandAZURE_AI_MODEL_DEPLOYMENT_NAME, then run the sample below.CALL/RESULTlines in the output, so the tool was never invoked. Running the same agent without streaming (agent.run(prompt)) invokes the tool correctly.Code Sample
Output (3 runs, latest packages):
Raw event dump without agent-framework, which shows the service-side stream above:
Error Messages / Stack Traces
Package Versions
agent-framework-core: 1.20.0, agent-framework-openai: 1.15.0, agent-framework-foundry: 1.14.0, agent-framework-foundry-hosting: 1.0.0b261002, openai: 3.24.0. Also reproduced with agent-framework-core: 1.16.0, agent-framework-openai: 1.14.1, agent-framework-foundry: 1.11.0, agent-framework-foundry-hosting: 1.0.0b260821, openai: 2.54.0.
Python Version
Python 3.13 (Linux x86_64; also in a Foundry hosted-agent container)
Additional Context
ResponsesHostServer), the call wasn't dropped. Its arguments came out as a complete JSON object with a delta fragment appended ({...valid args...}<fragment>", "<key>": <value>}), so each call failed withError: Argument parsing failed.After three model retries the request ended withMaximum consecutive function call errors reached (3), and the user gotFunction invocation limit reached before a final answer could be produced.The model's own.donearguments were valid every time.output_item.addedplus incomplete deltas, with the complete arguments only in the.doneevents.agent.run(prompt)). That works, but it isn't an option behindResponsesHostServerwhen the client asks for streaming.