Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.36254v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Acro
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Chinese firms trail global peers on profits, but AI power boom offers bright spot: NatixisSCMP Tech · September 30, 2026
- More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with JevarXiv cs.AI · September 30, 2026
- GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience AnalysisarXiv cs.AI · September 30, 2026
- The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfacearXiv cs.AI · September 30, 2026
- An Empirical Study and Assessment of EU AI Act Compliance CheckersarXiv cs.AI · September 30, 2026
- Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language ModelsarXiv cs.AI · September 30, 2026