From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization
arXiv cs.AIen
arXiv:2609.19630v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two l
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Företag
Related AI news
- Security researchers in an OpenAI bug bounty program hacked OpenAI, accessing its "monorepo" on GitHub, using a cybersecurity version of Opus 4.8 and Opus 5 (Robert McMillan/Wall Street Journal)Techmeme · September 18, 2026
- Zero trust har ett stort AI-problemComputer Sweden · September 18, 2026
- What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered AnalysisarXiv cs.AI · September 18, 2026
- Do AI Agents Understand Computer Architecture?arXiv cs.AI · September 18, 2026
- Self Improvement via Fast Tree-searcharXiv cs.AI · September 18, 2026
- When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e ScreeningarXiv cs.AI · September 18, 2026