Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Ro
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- AI demand drives Episil SiC V-shaped recovery as GaN and TVS capacity reaches full utilizationDIGITIMES · October 8, 2026
- Singtel signs MOU with SIT on sovereign AITech in Asia · October 8, 2026
- Google debuts SynthID Detector tool for flagging AI-generated content, but it’s far from perfectSiliconANGLE · October 8, 2026
- Chang Wah Technology posts record revenue as demand broadens across electronics and AIDIGITIMES · October 8, 2026
- Pricey local AI machines arrive as memory costs threaten PC shipmentsDIGITIMES · October 8, 2026
- At WebexOne, Cisco’s Jeetu Patel argues the agent era will be won on context, cost and controlSiliconANGLE · October 8, 2026