Asymmetries in Spontaneous and Instructed Deception
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data th
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Meta
- Verktyg
- Forskning
Related AI news
- New AI models, critical threshold, and a design debate: A busy day for AI industryEconomic Times Tech · September 3, 2026
- Chinese chipmaker Enflame 4,073 times oversubscribed in Shanghai IPO amid Nvidia raceSCMP Tech · September 3, 2026
- Several international law firms are seeking to build bespoke AI tools to gain an edge and protect their IP, while using off-the-shelf AI for everyday tasks (Nick Huber/Financial Times)Techmeme · September 3, 2026
- heise+ | OpenClaw selbst gebaut: Agent Runtime, Channels und KI-Heartbeatheise online – KI · September 3, 2026
- « Histoire culturelle de l’IA » aborde l’intelligence artificielle comme un phénomène social global et une révolution anthropologiqueLe Monde Pixels · September 3, 2026
- OpenAI is building 'automated shutdown' capabilities for AI tools, letter to lawmakers saysEconomic Times Tech · September 3, 2026