Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.04561v1 Announce Type: new Abstract: Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder can produce fluent hallucinated transcripts for inputs containing little or no speech. We propose a training-free, inference-time method to reduce these hallucinations using low-rank projection of decoder activations. A compact hallucination-associated subspace is estimated from non-speech calibration data, and decoder hidden states are projected away from this subspace during inference. We evaluate two variants: always-on, which applies projection to all inputs, and gated, which applies it only when Whisper predicts that an input is lik
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Röst-AI
- Verktyg
- Forskning
Related AI news
- Iris: Climbing to the Search FrontierarXiv cs.AI · September 7, 2026
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended RealityarXiv cs.AI · September 7, 2026
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding AgentsarXiv cs.AI · September 7, 2026
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model PipelinesarXiv cs.AI · September 7, 2026
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationarXiv cs.AI · September 7, 2026
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance?arXiv cs.AI · September 7, 2026