Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. In this work, we introduce CONFLICTGUI, a benchmark covering instruction-internal conflicts and instruction-GUI context conflicts to study conflict-aware termination. Our evaluation reveals severe execution-biased overcompliance: agents that perform well on feasible tasks often continue to execute blindly under conflicting instructions. To mitigate this behavior, we propose CONFL
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- Vertex Ventures leads $6m round in SG robotics firm HiveboticsTech in Asia · September 4, 2026
- Chinese AI firm Moonshot files confidentially for Hong Kong IPOEconomic Times Tech · September 4, 2026
- AI startup Crusoe valued at $30 billion after new fundingEconomic Times Tech · September 4, 2026
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency PenaltyarXiv cs.AI · September 4, 2026
- GPS-Bench: A Governance Policy Benchmark for Automating Policy AnalysisarXiv cs.AI · September 4, 2026
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer ReviewsarXiv cs.AI · September 4, 2026