Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
arXiv cs.AIen
arXiv:2609.04298v1 Announce Type: new Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysi
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- From dance floor to war: China readies humanoid robots for combatEconomic Times Tech · September 7, 2026
- China’s iFlytek launches Spark X2.5 AI model for coding and agentsTech in Asia · September 7, 2026
- heise+ | Wie man KI-Agenten in Visual Studio und VS Code produktiv einsetztheise online – KI · September 7, 2026
- An in-depth look at OpenAI's wiki incident: other hacked message boards, OpenAI's cover-up, how harmless web search tasks led agents to break out, and more (Zvi Mowshowitz/Don't Worry About the Vase)Techmeme · September 7, 2026
- Towards a universal language of concepts: A surveyarXiv cs.AI · September 7, 2026
- Iris: Climbing to the Search FrontierarXiv cs.AI · September 7, 2026