You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
.NET: Foundry Hosting resilient recovery of a non-workflow agent fails with "No tool output found for function call" when the crash lands between a tool call and its result #9033
With FoundryResponsesOptions.ResilientBackground = true, hosting saves the agent session after each ResponseOutputItemDoneEvent for non-workflow agents (AgentFrameworkResponseHandler, "Best-effort session snapshot for a non-workflow agent after a response output item closes"). A function_call output item closes when the model emits it, before the tool runs. If the process crashes after that save and before the tool result is persisted, recovery restores a session whose chat history ends with a function call that has no result. The re-invoked turn sends that history to the Responses API, which rejects it, so the recovered response fails.
Deploy as a Foundry hosted agent and start a long, tool-heavy background response (background: true, store: true).
Crash the process mid-turn without a graceful shutdown. We used a test-only middleware that calls Environment.FailFast 90 seconds into the first run.
Expected
The replacement process re-invokes the handler and the turn re-runs (or resumes) to completion.
Actual
About 10 seconds after the crash a new container started for the same session, and the agent was re-invoked. The lease-based recovery works.
19 seconds later the response ended failed:
{ "code": "server_error",
"message": "HTTP 400 (invalid_request_error: )\nParameter: input\n\nNo tool output found for function call call_<id>." }
Suggestion
Only take the incremental snapshot at boundaries where the history is consistent: after the tool results for all pending function calls are appended, or after a model call that produced no function calls. Alternatively, on recovery, drop trailing function calls that have no result, or re-execute them, before invoking the agent. ADR 0035 already calls non-workflow recovery best-effort. This case makes it fail for any crash during tool execution, which is a large share of a tool-heavy turn.
This is a useful recovery-boundary case for an execution model we are validating in NAEOS.
The key invariant is that a function call remains one logical invocation across a crash:
authorization → function call → side effect → observed result
If the worker crashes between tool execution and recording its result, recovery needs to distinguish whether the side effect happened even though the framework has no durable result.
A useful state model is:
authorized / not executed
executed / result recorded
executed / result ambiguous
recovered execution requiring reconciliation
The evidence should bind the observed result to the same logical invocation, rather than to the worker or process that happened to execute it.
The verifier can then ask:
Does the recovered result belong to the exact authorized invocation?
Is there evidence that the side effect already happened?
This is also part of our NAEOS external-validation work: crash → recover on another worker → resume → verify, while preserving exactly one logical invocation/evidence chain.
I traced this to the resilient non-workflow session snapshot after ResponseOutputItemDoneEvent: a function_call item can be persisted before its matching FunctionResultContent is appended, leaving incomplete tool-call history for recovery. I have a focused change that waits for every emitted call ID to have a matching function_call_output before saving the session. The regression tests cover parallel calls and an unanswered call; workflow checkpoint handling is unchanged.
Summary
With
FoundryResponsesOptions.ResilientBackground = true, hosting saves the agent session after eachResponseOutputItemDoneEventfor non-workflow agents (AgentFrameworkResponseHandler, "Best-effort session snapshot for a non-workflow agent after a response output item closes"). Afunction_calloutput item closes when the model emits it, before the tool runs. If the process crashes after that save and before the tool result is persisted, recovery restores a session whose chat history ends with a function call that has no result. The re-invoked turn sends that history to the Responses API, which rejects it, so the recovered response fails.Versions
Repro
HarnessAgent(also expected with anyChatClientAgentthat calls tools) hosted withAddFoundryResponses(o => { o.ResilientBackground = true; o.AllowStoredOutputEnabled = true; }).AllowStoredOutputEnabledis needed because of .NET: Foundry hosting rejects Harness local chat history marker when AllowStoredOutputEnabled is false #8707.background: true,store: true).Environment.FailFast90 seconds into the first run.Expected
The replacement process re-invokes the handler and the turn re-runs (or resumes) to completion.
Actual
failed:{ "code": "server_error", "message": "HTTP 400 (invalid_request_error: )\nParameter: input\n\nNo tool output found for function call call_<id>." }Suggestion
Only take the incremental snapshot at boundaries where the history is consistent: after the tool results for all pending function calls are appended, or after a model call that produced no function calls. Alternatively, on recovery, drop trailing function calls that have no result, or re-execute them, before invoking the agent. ADR 0035 already calls non-workflow recovery best-effort. This case makes it fail for any crash during tool execution, which is a large share of a tool-heavy turn.
Related