BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.16305v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge only after multiple turns, yet existing evaluations often reduce agent behavior to task or attack success, obscuring whether an agent acts, refuses, or remains appropriately calibrated as the interaction evolves. We introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents. Blindspot evaluates complete user-agent-environment trajectories through adaptive adversarial intera
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Företag
Related AI news
- Huawei sees AI agents driving 90% of token traffic by 2035DIGITIMES · September 17, 2026
- Snap targets enterprises with Salesforce, Nvidia AI tools for augmented-reality glassesEconomic Times Tech · September 17, 2026
- Dassault Systemes shifts to AI-native platforms, stakes its next phase on TaiwanDIGITIMES · September 17, 2026
- Open-weight model developer Arcee AI reaches $1B-plus valuation with new fundingSiliconANGLE · September 17, 2026
- Open-weight model developer Arcee AI reaches $1B-plus valuation with undisclosed Series B fundingSiliconANGLE · September 17, 2026
- ByteDance's Anew Labs reportedly raises US$290M as AI drug discovery gains momentumDIGITIMES · September 17, 2026