Problem
The tool-use suite passes 100% on six easy cases, which measures little. The
behaviors that actually break agents are disambiguation, refusal, empty results,
and long dependency chains. None are covered.
Proposal
Add five adversarial scenarios to agent-tool-use/scenarios.ts:
no-tool-needed — answer directly; call nothing even though tools exist
disambiguate-similar-tools — forecast tool vs current-weather tool
empty-result-no-hallucination — empty search; must not fabricate a policy
long-chain-dependency — four tools, each depending on the previous
near-duplicate-names — get_user vs get_user_settings
Scripted expectations keep them deterministic in CI; the same cases run in the
live and model-comparison suites.
Acceptance criteria
Problem
The tool-use suite passes 100% on six easy cases, which measures little. The
behaviors that actually break agents are disambiguation, refusal, empty results,
and long dependency chains. None are covered.
Proposal
Add five adversarial scenarios to
agent-tool-use/scenarios.ts:no-tool-needed— answer directly; call nothing even though tools existdisambiguate-similar-tools— forecast tool vs current-weather toolempty-result-no-hallucination— empty search; must not fabricate a policylong-chain-dependency— four tools, each depending on the previousnear-duplicate-names—get_uservsget_user_settingsScripted expectations keep them deterministic in CI; the same cases run in the
live and model-comparison suites.
Acceptance criteria