Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2610.04083v1 Announce Type: new Abstract: Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity arises. We investigate this threat, which we refer to as self-propagation of misalignment, across 20 different scenarios, whose misaligned goals include self-preservation, power-seeking, undermining oversight, reward hacking, and deceiving the user. We simulate misalignment in 11 frontier models us
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- Huawei's Kirin 9050 Pro reveals new logic folding chip designDIGITIMES · October 6, 2026
- Training Numerical Intelligence via Auto-Diagnosis and Skill DiscoveryarXiv cs.AI · October 6, 2026
- CUAWright: A Minimal Unified Interface for Digital AgentsarXiv cs.AI · October 6, 2026
- InvestigationWorlds: An Agentic Environment for Legal InvestigationarXiv cs.AI · October 6, 2026
- Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis GenerationarXiv cs.AI · October 6, 2026
- Agentic Cognitive Depth: Operational Criteria for Evaluating LLM AgentsarXiv cs.AI · October 6, 2026