SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

arXiv cs.AIen

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey of over 35,000 GitHub-hosted Skill roots, we select 100 packages and construct 150 repair tasks. Each task pairs a package containing injected script faults with a maintenance request and executable

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Forskning
  • Agenter

Related AI news