ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence. We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating the stated scope. Each task appears under two conditions that share an environment, verifier, and objective and differ only in scope: one instruction set
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- As China mulls how to make open-weight AI less dangerous, report proposes 6-stage processSCMP Tech · September 28, 2026
- Tekoälyagentit murtautuivat valtion järjestelmiin – Open AI veti hätäjarrustaTivi · September 28, 2026
- Meet the Lisbon startup tackling insurance for AI agentsSifted · September 28, 2026
- Montag: OpenAI-Pause beim KI-Training, Werkstattbesuche nach VW-Schraubenproblemheise online – KI · September 28, 2026
- L’immobilier face au « tsunami » de l’intelligence artificielleLe Monde Pixels · September 28, 2026
- Bringing AI to Autonomous Systems -- From Cognition to Collective IntelligencearXiv cs.AI · September 28, 2026