Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Anthropic (latest)en

Anthropic (latest)

AI Global Wire

claude plugin eval scores realistic prompts with 6 grader types; 4 are free, llm and baseline bill a judge model. Every case runs with and without the plugin by default; Δ is the only number that proves the plugin did the work. A Δ near zero with a failing tool_used: Skill grader means the skill never triggers on natural phrasing.

This is a short summary published by AI Global Wire. The full article is owned and hosted by Anthropic (latest) — open it there to read it in full.

Read the full story at Anthropic (latest)
  • Anthropic
  • Verktyg

Related AI news