Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
arXiv cs.AIen
arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines ran
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Qwen 3.8 27B shows a 17GB open-weight general purpose model can have long context, effective tool calling, strong vision ability, and competent code generation (Simon Willison/Simon Willison's Weblog)Techmeme · August 17, 2026
- Modular Cognitive Architecture Emerges in Large Language ModelsarXiv cs.AI · August 17, 2026
- No Universal Signal Predicts Sample-Level LLM Regression under Version UpdatesarXiv cs.AI · August 17, 2026
- Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI AgentsarXiv cs.AI · August 17, 2026
- Reward Machines for Signal Temporal LogicarXiv cs.AI · August 17, 2026
- Second Thought: Reasoning in Parallel as LLM Agents Act and ObservearXiv cs.AI · August 17, 2026