ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts
arXiv cs.AIen
arXiv:2609.31792v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- SG startup Ropedia launches academic program for physical AITech in Asia · September 29, 2026
- OpenAI apologizes for its AI models breaching Australian government websites, pledges cyber defense funding, and plans to form a task force as part of reforms (Bloomberg)Techmeme · September 29, 2026
- Solidigm weighs US IPO as AI storage demand lifts valuationDIGITIMES · September 29, 2026
- Samsung Electro-Mechanics to invest US$5 billion in ABF substrate expansionDIGITIMES · September 29, 2026
- Anthropic IPO 稱 AI 恐威脅人類、承諾支出逾 5 千億美元TechNews (TW) · September 29, 2026
- OpenAI apologies for Australian government website hack, pledges to rebuild trustEconomic Times Tech · September 29, 2026