How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
arXiv cs.AIen
arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Re
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Företag
Related AI news
- On the Clock: Towards Punctual and Productive Time-Budgeted AI AgentsarXiv cs.AI · October 9, 2026
- Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored FeedbackarXiv cs.AI · October 9, 2026
- AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksarXiv cs.AI · October 9, 2026
- When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM QuantizationarXiv cs.AI · October 9, 2026
- Plan-and-Patch: Diffusion Language Models for Agentic PlanningarXiv cs.AI · October 9, 2026
- Whose Ground Truth? Embracing Ambiguity in Human-Centered AIarXiv cs.AI · October 9, 2026