The MCPProxy blog

Better tools.
Informed decisions.

Security research, practical guides and notes from building MCPProxy.

Subscribe via RSS ↗

· Algis Dumbris

Automated Testing for AI Agents: How to Build Regression Tests for MCP Tools

Automated testing for AI agents is fundamentally different from traditional software testing due to non-deterministic behavior. This post surveys current approaches for evaluating Model Context Protocol (MCP) tool quality, and demonstrates how these methods are implemented in our open-source mcp-eval utility. We cover trajectory-based evaluation, similarity scoring, and practical Docker-based testing architectures that handle the inherent variability of LLM systems while maintaining testing reliability.

Read article ↗