Repository navigation
Python: [Bug]: evaluate_agent() omits ContextProvider/SkillsProvider tools from tool_definitions, breaking tool_selection/tool_input_accuracy #9160
Description
Activity
- addedpythonUsage: [Issues, PRs], Target: PythonUsage: [Issues, PRs], Target: PythontriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 7, 2026 - addedreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflowUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
on Oct 7, 2026 🤖 Automated triage reproduction notes (agent-authored — trust but verify)
Agent analysis
Repro:
python/packages/core/agent_framework/_evaluation.py::_to_eval_itemat lines 877-885 omits tools injected byagent.context_providers, so a response callingload_skillreceives anEvalItem.toolslist containing only default/MCP tools. Minimal repro: attach a realSkillsProviderand one MCPFunctionTool, run the provider hook to establishload_skillis available, then call_to_eval_item()with a response containing aload_skillfunction call and assert both definitions are present.- Failing test:
python/packages/core/tests/core/test_evaluation.py::TestToEvalItem::test_with_agent_context_provider_tools - Files examined: python/packages/core/agent_framework/_evaluation.py, python/packages/core/agent_framework/_skills.py, python/packages/core/agent_framework/_sessions.py, python/packages/core/agent_framework/_agents.py, python/packages/core/tests/core/test_evaluation.py, python/packages/core/pyproject.toml
- Tests run: python/packages/core/tests/core/test_evaluation.py, TestToEvalItem::test_with_agent_context_provider_tools
- Reported version:
1.20.0 - Current version:
1.20.0
- Failing test:
- addedagentsUsage: [Issues, PRs], Target: Single agentUsage: [Issues, PRs], Target: Single agentfoundryUsage: [Issues, PRs], Target: all Foundry integrationsUsage: [Issues, PRs], Target: all Foundry integrationsskillsUsage: [Issues, PRs], Target: skills related featuresUsage: [Issues, PRs], Target: skills related featuresand removedtriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 7, 2026 Hi Eduard van Valkenburg (@eavanvalkenburg) — I would like to take this issue. I have reproduced the omission locally and prepared a focused change to include tools exposed by the agent ContextProvider/SkillsProvider in evaluation items, alongside the existing default and MCP tools, without changing response serialization. I see the issue is currently assigned to you; could you confirm whether this scope is available for a community contribution or already planned? I will wait for your confirmation before opening a PR.
Metadata
Metadata
Labels
Type
Projects
- StatusShow more project fieldsNo status
Description
evaluate_agent()(andagent_framework._evaluation._to_eval_item, which builds thetool_definitionssent to Microsoft Foundry's builtin agent evaluators) only collects tool definitions from two sources:agent.default_options["tools"]andagent.mcp_tools[*].functions. Any tool exposed through aContextProvider/SkillsProvider(for example the framework's ownload_skilltool, added when aSkillsProvideris attached to the agent) is never included.When the agent's response messages contain a
tool_callto one of these context-provider-exposed tools (because the model legitimately called it mid-conversation), the Foundry Evaluations API'stool_selectionandtool_input_accuracyevaluators returnscore: null,passed: null, with the reason:This happens even though the tool call is real, succeeded, and is fully present in the conversation's
response_messages— the evaluator is correctly enforcing its own precondition ("all tool calls must have a matching definition"), but the SDK never gave it one for this tool.What I expected:
tool_definitionsto include every tool the agent can actually call, regardless of which provider (MCP,default_options["tools"], or aContextProvider/SkillsProvider) exposed it to the model — so these two evaluators can score any conversation that uses a skill-provided tool, not just conversations that happen to avoid them entirely.What happens instead: in a test suite with a
SkillsProviderattached (even one that only exposes a singleload_skilltool, auto-approved, no user-facing side effects), any dataset row whose conversation callsload_skillbefore its real MCP/function tool permanently losestool_selection/tool_input_accuracysignal. In our own regression dataset, this affected 5 of 6 rows — only the one row whose conversation never touchedload_skillgot scored.Steps to reproduce
Agentwith at least one MCP tool and aContextProvider/SkillsProviderthat exposes an additional tool not sourced fromdefault_options["tools"]oragent.mcp_tools(e.g. a framework-providedload_skilltool).agent_framework.evaluate_agent(agent=agent, queries=[...], ...)against Foundry'stool_selectionandtool_input_accuracyevaluators.output_itemson the Evaluations API, or the Foundry portal run detail) for that row.Code Sample
Relevant framework code (confirmed in the installed
agent_framework1.20.0 wheel,agent_framework/_evaluation.py, function_to_eval_item):This never visits
agent.context_providers(or however the attachedSkillsProvider's own tool(s) are exposed), so a tool likeload_skillnever makes it intotool_definitions.Error Messages / Stack Traces
Not an exception; the evaluator returns a normal (non-error) result with a null score and this reason string (confirmed via the raw Evaluations API
output_itemsresponse for the affected rows):Package Versions
agent-framework-core: 1.20.0, agent-framework-foundry: 1.14.0, agent-framework-openai: 1.15.0
Python Version
Python 3.12.6
Additional Context
tool_selection/tool_input_accuracyfrom our own retry-loop's failure-detection logic, since theNonescore is reproducible on every attempt (not a transient judge/model error) and was making our test's retry loop exhaust all attempts on an otherwise-clean run.ContextProvider/SkillsProviderattached to theAgent(the framework's own skill-loading mechanism). We have not checked whether other non-MCP, non-default_optionstool sources have the same gap, but the fix likely generalizes to "collect tool definitions from every tool source the agent actually exposes to the model," not just MCP anddefault_options["tools"].tool_definitionsdirectly; it doesn't cover theevaluate_agent(agent=...)path where the SDK buildstool_definitionsfor the caller and silently omits an entire tool source.