Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
Anthropic (latest)en
Anthropic (latest)
AI Global Wireclaude plugin eval scores realistic prompts with 6 grader types; 4 are free, llm and baseline bill a judge model. Every case runs with and without the plugin by default; Δ is the only number that proves the plugin did the work. A Δ near zero with a failing tool_used: Skill grader means the skill never triggers on natural phrasing.
This is a short summary published by AI Global Wire. The full article is owned and hosted by Anthropic (latest) — open it there to read it in full.
Read the full story at Anthropic (latest)- Anthropic
- Verktyg
Related AI news
- NYC-based Luminary, which develops AI-powered workflow tools for estate planning and wealth transfer management, raised a $22M Series A led by Ten Coves Capital (Davis Janowski/Wealth Management)Techmeme · September 12, 2026
- Nvidia in talks to invest $10b in Anthropic IPO: sourcesTech in Asia · September 12, 2026
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series ForecastingarXiv cs.AI · September 12, 2026
- Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM WorkflowsarXiv cs.AI · September 12, 2026
- Defining AI Agents: A Compendium of Criteria, Metrics, and BenchmarksarXiv cs.AI · September 12, 2026
- MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAGarXiv cs.AI · September 12, 2026