LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
arXiv cs.AIen
arXiv:2608.27580v1 Announce Type: new Abstract: Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistically analyzes, and mitigates long-context guardrail failure. We formulate the task as Safety Needle-in-a-Haystack (SafetyNIAH) over a 0.25k-32k length grid; across 15 mainstream guardrails, unsafe recall drops monotonically by more than 50% on average, and a paired Benign-Fill vs. Needle-Repeat design attributes the failure to proportional dilution of the unsafe needle rather than to absolute length
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Urheberrechtsklage gegen KI-Entwickler: Sony und Warner verklagen AnthropicGolem.de · August 31, 2026
- Big Tech reported Q2 "other income" rose significantly to $160B+, driven by investments in AI companies, raising concerns of paper gains overstating the AI boom (Financial Times)Techmeme · August 31, 2026
- Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea PlantationsarXiv cs.AI · August 31, 2026
- Thinking Costs Tokens: When More Structure is Worth the PricearXiv cs.AI · August 31, 2026
- Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety ModeratorarXiv cs.AI · August 31, 2026
- Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative ControlsarXiv cs.AI · August 31, 2026