Agent Seer: Synthesizing Scenarios from Specification Understanding

| Source: Apple ML Research

Tags: Apple, Agent Seer, MCP, agentic AI, AI evaluation, tool-calling, benchmarks

Apple's Agent Seer auto-generates graded evaluation scenarios for tool-calling agents from MCP specifications alone — no manual curation or live tool access needed — with strong quality across 7 diverse domains. Key finding: argument value accuracy, not tool selection, is the failure mode that coarse benchmarks miss.

Details

Evaluating AI agents that call external tools has required either expensive manual scenario construction or static benchmarks that can't track evolving APIs. Apple ML Research's Agent Seer solves this by generating complete, graded evaluation scenarios directly from Model Context Protocol (MCP) specifications — using only function names, natural-language descriptions, and typed parameter schemas. The pipeline operates with no examples, no live tool execution, and no domain-specific tuning. From a single MCP spec, it enriches raw schemas, generates scenarios with synthetic tool outputs, and expands them into multi-turn dialogues exhibiting realistic tool-calling patterns and conversational coherence. Evaluated across 7 diverse MCP specifications spanning different domains and tool-suite sizes, the pipeline achieves strong tool-calling correctness and coherence, with complete tool coverage on small and medium specifications. Two findings emerge. First, parameter schema complexity is the strongest predictor of quality variation — how many parameters a tool has and how complex their types are matters more than how many tools exist in the suite. Second, argument value accuracy is the dominant failure mode among imperfect scenarios — the agent selects the right tool but fills argument values incorrectly. This dimension is invisible to name-match metrics that most agent benchmarks use, meaning many existing evaluations are over-counting successes. For practitioners building MCP-connected agents or tool-augmented LLMs, Agent Seer offers a scalable path to generating evaluation suites that stay current as APIs change, without manual work per endpoint.