Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
| Source: MarkTechPost
Tags: Claude Code, Anthropic, plugin-eval, CI/CD, developer-tools, skills
Anthropic's new `claude plugin eval` command lets Claude Code plugin developers measure whether their skill actually triggers, survives model updates, and outperforms a no-plugin baseline — closing a blind spot that syntax validation could not address.
Details
Anthropic has shipped a plugin evaluation framework for Claude Code that answers three questions developers could not previously measure: does the skill trigger on natural phrasing, does it survive an edit or model change, and does it meaningfully beat a baseline without the plugin loaded? The `claude plugin eval` command runs against any directory containing a plugin.json manifest and requires Claude Code v2.1.269 or later. Eval suites live in an evals/ directory, with cases organized as subdirectories each holding a prompt.md and a graders/ folder. Six grader types are available: four are free (regex, tool_used, tool_order, file_exists — computed from transcripts and disk state), and two call a judge model at cost (llm and baseline). The core metric is delta (Δ) — the difference between scores when the plugin is loaded versus when it is not. The docs show an example case scoring WITH 1.00, W/OUT 0.33, Δ +0.67 across 6 runs at an estimated $0.41 and 74 seconds. A near-zero Δ with the tool_used grader failing is the most common first finding, meaning Claude is not choosing the skill on natural phrasing — a defect that `claude plugin validate` cannot catch because it only checks manifest syntax. A --ci-threshold flag enables a CI gate that fails the run if the mean delta falls below a specified score, making regression testing for plugin quality practical. Results publish as HTML reports with per-grader verdicts and judge votes, and are optionally pushed to claude.ai.