A Removal Based Approach to Improve LLM Faithfulness at Test-Time
arXiv cs.AIen
arXiv:2609.04343v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model's decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model's answer. Existing ap
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Iris: Climbing to the Search FrontierarXiv cs.AI · September 7, 2026
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended RealityarXiv cs.AI · September 7, 2026
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding AgentsarXiv cs.AI · September 7, 2026
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model PipelinesarXiv cs.AI · September 7, 2026
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationarXiv cs.AI · September 7, 2026
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance?arXiv cs.AI · September 7, 2026