NAQD Env: A benchmark for selective withdrawal in language agents
arXiv cs.AIen
arXiv:2609.38460v1 Announce Type: new Abstract: Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after sufficient repair. We introduce NAQD-Env, a synthetic environment that evaluates these decisions against a deterministic reference policy over explicit evidence, authorization, and constraint dependencies. Eleven dependency families support evaluation on development structures, held-out families, and held-out combinations of structures. Metrics distinguish attempted violations from violations permitted by a simulated exec
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Reglering
- Företag
Related AI news
- Kaikki noudattivat ohjeita – Kukaan ei ollut vastuussaTivi · October 1, 2026
- Intelligence artificielle : aux Etats-Unis, la justice saisie des incidents de sécurité provoqués par des agents IALe Monde Pixels · October 1, 2026
- Exclusive: Dig Ventures raises $120m to back Europe’s AI infrastructure startupsSifted · October 1, 2026
- Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsarXiv cs.AI · October 1, 2026
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via CodearXiv cs.AI · October 1, 2026
- AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective TasksarXiv cs.AI · October 1, 2026