Skip to content

.NET: Foundry Hosting resilient recovery of a non-workflow agent fails with "No tool output found for function call" when the crash lands between a tool call and its result #9033

Description

Summary

With FoundryResponsesOptions.ResilientBackground = true, hosting saves the agent session after each ResponseOutputItemDoneEvent for non-workflow agents (AgentFrameworkResponseHandler, "Best-effort session snapshot for a non-workflow agent after a response output item closes"). A function_call output item closes when the model emits it, before the tool runs. If the process crashes after that save and before the tool result is persisted, recovery restores a session whose chat history ends with a function call that has no result. The re-invoked turn sends that history to the Responses API, which rejects it, so the recovered response fails.

Versions

  • Microsoft.Agents.AI.Foundry.Hosting 1.23.0-preview.260928.1, Microsoft.Agents.AI.Harness 1.23.0
  • Azure.AI.AgentServer.Core 1.0.0-beta.29, Azure.AI.AgentServer.Responses 1.0.0-beta.8
  • Foundry hosted agent (Responses container protocol 2.0), model gpt-5.4, .NET 10

Repro

  1. A HarnessAgent (also expected with any ChatClientAgent that calls tools) hosted with AddFoundryResponses(o => { o.ResilientBackground = true; o.AllowStoredOutputEnabled = true; }). AllowStoredOutputEnabled is needed because of .NET: Foundry hosting rejects Harness local chat history marker when AllowStoredOutputEnabled is false #8707.
  2. Deploy as a Foundry hosted agent and start a long, tool-heavy background response (background: true, store: true).
  3. Crash the process mid-turn without a graceful shutdown. We used a test-only middleware that calls Environment.FailFast 90 seconds into the first run.

Expected

The replacement process re-invokes the handler and the turn re-runs (or resumes) to completion.

Actual

  • About 10 seconds after the crash a new container started for the same session, and the agent was re-invoked. The lease-based recovery works.
  • 19 seconds later the response ended failed:
{ "code": "server_error",
  "message": "HTTP 400 (invalid_request_error: )\nParameter: input\n\nNo tool output found for function call call_<id>." }

Suggestion

Only take the incremental snapshot at boundaries where the history is consistent: after the tool results for all pending function calls are appended, or after a model call that produced no function calls. Alternatively, on recovery, drop trailing function calls that have no result, or re-execute them, before invoking the agent. ADR 0035 already calls non-workflow recovery best-effort. This case makes it fail for any crash during tool execution, which is a large share of a tool-heavy turn.

Related

Activity

  1. added
    .NETUsage: [Issues, PRs], Target: .Net
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 4, 2026
  2. bayu911-source commented on Oct 4, 2026

    @bayu911-source

    This is a useful recovery-boundary case for an execution model we are validating in NAEOS.

    The key invariant is that a function call remains one logical invocation across a crash:

    authorization → function call → side effect → observed result

    If the worker crashes between tool execution and recording its result, recovery needs to distinguish whether the side effect happened even though the framework has no durable result.

    A useful state model is:

    • authorized / not executed
    • executed / result recorded
    • executed / result ambiguous
    • recovered execution requiring reconciliation

    The evidence should bind the observed result to the same logical invocation, rather than to the worker or process that happened to execute it.

    The verifier can then ask:

    1. Does the recovered result belong to the exact authorized invocation?
    2. Is there evidence that the side effect already happened?

    This is also part of our NAEOS external-validation work: crash → recover on another worker → resume → verify, while preserving exactly one logical invocation/evidence chain.

  3. jstar0 commented on Oct 5, 2026

    @jstar0
    Contributor

    I traced this to the resilient non-workflow session snapshot after ResponseOutputItemDoneEvent: a function_call item can be persisted before its matching FunctionResultContent is appended, leaving incomplete tool-call history for recovery. I have a focused change that waits for every emitted call ID to have a matching function_call_output before saving the session. The regression tests cover parallel calls and an unanswered call; workflow checkpoint handling is unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

.NETUsage: [Issues, PRs], Target: .NettriageUsage: [Issues], Target: All issues that still need to be triaged

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions