Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.26167v1 Announce Type: new Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models were evaluated on the TAME Pain speech corpus. Participants read phonetically balanced Harvard Sentences while one hand was immersed in cold or warm water and reported pain only during periodic pain statements. This protocol generated 5,750 no signal Harvard Sentence utterances whose transcripts contained no lexical pain information and 1,294 signal pain statement utterances in which the pain rating was explicitly spoke
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- SoftBank-backed surgical robotics firm eyes Hong Kong listing to power mainland China pushSCMP Tech · August 28, 2026
- Anthropic unveils new framework allowing AI agents to operate physical devicesEconomic Times Tech · August 28, 2026
- A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem SolvingarXiv cs.AI · August 28, 2026
- The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement LearningarXiv cs.AI · August 28, 2026
- Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM IntegrationarXiv cs.AI · August 28, 2026
- A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in AnianarXiv cs.AI · August 28, 2026