Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
arXiv cs.AIen
arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, a
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Meta
- Forskning
Related AI news
- Security researchers in an OpenAI bug bounty program hacked OpenAI, accessing its "monorepo" on GitHub, using a cybersecurity version of Opus 4.8 and Opus 5 (Robert McMillan/Wall Street Journal)Techmeme · September 18, 2026
- What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered AnalysisarXiv cs.AI · September 18, 2026
- Do AI Agents Understand Computer Architecture?arXiv cs.AI · September 18, 2026
- Self Improvement via Fast Tree-searcharXiv cs.AI · September 18, 2026
- When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e ScreeningarXiv cs.AI · September 18, 2026
- Reach or Solve? Attributing Agentic RL Gains with Checkpoint HandoffsarXiv cs.AI · September 18, 2026