Skip to content

Python: [Bug]: evaluate_agent() omits ContextProvider/SkillsProvider tools from tool_definitions, breaking tool_selection/tool_input_accuracy #9160

Description

Description

evaluate_agent() (and agent_framework._evaluation._to_eval_item, which builds the tool_definitions sent to Microsoft Foundry's builtin agent evaluators) only collects tool definitions from two sources: agent.default_options["tools"] and agent.mcp_tools[*].functions. Any tool exposed through a ContextProvider/SkillsProvider (for example the framework's own load_skill tool, added when a SkillsProvider is attached to the agent) is never included.

When the agent's response messages contain a tool_call to one of these context-provider-exposed tools (because the model legitimately called it mid-conversation), the Foundry Evaluations API's tool_selection and tool_input_accuracy evaluators return score: null, passed: null, with the reason:

Not applicable: Tool definitions for all tool calls must be provided.

This happens even though the tool call is real, succeeded, and is fully present in the conversation's response_messages — the evaluator is correctly enforcing its own precondition ("all tool calls must have a matching definition"), but the SDK never gave it one for this tool.

What I expected: tool_definitions to include every tool the agent can actually call, regardless of which provider (MCP, default_options["tools"], or a ContextProvider/SkillsProvider) exposed it to the model — so these two evaluators can score any conversation that uses a skill-provided tool, not just conversations that happen to avoid them entirely.

What happens instead: in a test suite with a SkillsProvider attached (even one that only exposes a single load_skill tool, auto-approved, no user-facing side effects), any dataset row whose conversation calls load_skill before its real MCP/function tool permanently loses tool_selection/tool_input_accuracy signal. In our own regression dataset, this affected 5 of 6 rows — only the one row whose conversation never touched load_skill got scored.

Steps to reproduce

  1. Build an Agent with at least one MCP tool and a ContextProvider/SkillsProvider that exposes an additional tool not sourced from default_options["tools"] or agent.mcp_tools (e.g. a framework-provided load_skill tool).
  2. Run a query that causes the model to call the context-provider tool first, then a real MCP tool.
  3. Call agent_framework.evaluate_agent(agent=agent, queries=[...], ...) against Foundry's tool_selection and tool_input_accuracy evaluators.
  4. Inspect the run's per-item output (output_items on the Evaluations API, or the Foundry portal run detail) for that row.

Code Sample

from agent_framework import evaluate_agent

# agent has an MCP tool source AND a SkillsProvider exposing `load_skill`
results = await evaluate_agent(
    agent=agent,
    queries=["<a query that makes the model call load_skill, then a real MCP tool>"],
    evaluators=["tool_selection", "tool_input_accuracy", "tool_call_success"],
    model=judge_model,
)

for item in results[0].items:
    for score in item.scores:
        print(score.name, score.score, score.passed)
# tool_selection None None
# tool_input_accuracy None None
# tool_call_success 1.0 True   <- only this one, which doesn't depend on tool_definitions coverage, scores normally

Relevant framework code (confirmed in the installed agent_framework 1.20.0 wheel, agent_framework/_evaluation.py, function _to_eval_item):

elif agent:
    raw_tools = getattr(agent, "default_options", {}).get("tools", [])
    typed_tools = [tool for tool in normalize_tools(raw_tools) if isinstance(tool, FunctionTool)]
    seen = {tool.name for tool in typed_tools}
    for mcp in getattr(agent, "mcp_tools", []):
        for tool in getattr(mcp, "functions", []):
            if isinstance(tool, FunctionTool) and tool.name not in seen:
                typed_tools.append(tool)
                seen.add(tool.name)

This never visits agent.context_providers (or however the attached SkillsProvider's own tool(s) are exposed), so a tool like load_skill never makes it into tool_definitions.

Error Messages / Stack Traces

Not an exception; the evaluator returns a normal (non-error) result with a null score and this reason string (confirmed via the raw Evaluations API output_items response for the affected rows):

Not applicable: Tool definitions for all tool calls must be provided.

Package Versions

agent-framework-core: 1.20.0, agent-framework-foundry: 1.14.0, agent-framework-openai: 1.15.0

Python Version

Python 3.12.6

Additional Context

  • We worked around this on our side by excluding tool_selection/tool_input_accuracy from our own retry-loop's failure-detection logic, since the None score is reproducible on every attempt (not a transient judge/model error) and was making our test's retry loop exhaust all attempts on an otherwise-clean run.
  • This is specifically about tools exposed through a ContextProvider/SkillsProvider attached to the Agent (the framework's own skill-loading mechanism). We have not checked whether other non-MCP, non-default_options tool sources have the same gap, but the fix likely generalizes to "collect tool definitions from every tool source the agent actually exposes to the model," not just MCP and default_options["tools"].
  • Workaround suggestions for other evaluators facing the same tool is documented for unsupported provider-hosted tools (Azure AI Search, Bing Grounding, etc. — see https://learn.microsoft.com/azure/foundry/concepts/evaluation-evaluators/agent-evaluators#composite-evaluators-preview), but that guidance assumes the caller supplies tool_definitions directly; it doesn't cover the evaluate_agent(agent=...) path where the SDK builds tool_definitions for the caller and silently omits an entire tool source.

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 7, 2026
  2. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Oct 7, 2026
  3. github-actions commented on Oct 7, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    Repro: python/packages/core/agent_framework/_evaluation.py::_to_eval_item at lines 877-885 omits tools injected by agent.context_providers, so a response calling load_skill receives an EvalItem.tools list containing only default/MCP tools. Minimal repro: attach a real SkillsProvider and one MCP FunctionTool, run the provider hook to establish load_skill is available, then call _to_eval_item() with a response containing a load_skill function call and assert both definitions are present.

    • Failing test: python/packages/core/tests/core/test_evaluation.py::TestToEvalItem::test_with_agent_context_provider_tools
    • Files examined: python/packages/core/agent_framework/_evaluation.py, python/packages/core/agent_framework/_skills.py, python/packages/core/agent_framework/_sessions.py, python/packages/core/agent_framework/_agents.py, python/packages/core/tests/core/test_evaluation.py, python/packages/core/pyproject.toml
    • Tests run: python/packages/core/tests/core/test_evaluation.py, TestToEvalItem::test_with_agent_context_provider_tools
    • Reported version: 1.20.0
    • Current version: 1.20.0
  4. added
    agentsUsage: [Issues, PRs], Target: Single agent
    foundryUsage: [Issues, PRs], Target: all Foundry integrations
    skillsUsage: [Issues, PRs], Target: skills related features
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 7, 2026
  5. zhengguangzhuo commented on Oct 8, 2026

    @zhengguangzhuo

    Hi Eduard van Valkenburg (@eavanvalkenburg) — I would like to take this issue. I have reproduced the omission locally and prepared a focused change to include tools exposed by the agent ContextProvider/SkillsProvider in evaluation items, alongside the existing default and MCP tools, without changing response serialization. I see the issue is currently assigned to you; could you confirm whether this scope is available for a community contribution or already planned? I will wait for your confirmation before opening a PR.

  6. self-assigned this
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

agentsUsage: [Issues, PRs], Target: Single agentfoundryUsage: [Issues, PRs], Target: all Foundry integrationspythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflowskillsUsage: [Issues, PRs], Target: skills related features

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions