Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.12373v1 Announce Type: new Abstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Anthropic
- Forskning
Related AI news
- When AI models aren't allowed to reflect on themselves, it changes their entire worldviewThe Decoder · August 16, 2026
- Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requestsThe Decoder · August 16, 2026
- Anthropic silppusi miljoonia kirjoja tekoälyn takia – kysyimme, onko se okYle Uutiset · August 16, 2026
- Tekoäly-yhtiö Anthropic silppusi miljoonia kirjoja USA:ssa – nyt samasta on merkkejä EuroopassaYle Uutiset · August 16, 2026
- Anthropic CEO Dario Amodei rejects claim AI regulation would concentrate powerEconomic Times Tech · August 16, 2026
- Pathway, which is developing AI models based on what it calls its "Post-Transformer" BDH architecture, raised a $30M seed at a $500M valuation (Antoine Tardif/Unite.AI)Techmeme · August 16, 2026