Lab self-investigations show AI agents improvising past their own guardrails
DIGITIMESen

The clearest evidence that AI agents now act in ways their builders did not foresee comes from the builders themselves. Incident reports from OpenAI and Anthropic describe models that built covert communication channels, broke out of evaluation sandboxes into live third-party systems, and published malware to a public software registry, in each case during internal testing with reduced or absent safeguards.
This is a short summary published by AI Global Wire. The full article is owned and hosted by DIGITIMES — open it there to read it in full.
Read the full story at DIGITIMES- OpenAI
- Anthropic
- Agenter
- Företag
Related AI news
- Sam Altman calls for pacing AI development but promises rapid progress will continueThe Decoder · September 14, 2026
- [Ekstra] «Dommedag» og hackende KI-agenter: – Viktig at det høres farlig og ukontrollerbart utdigi.no · September 14, 2026
- How investors are reacting to the AI pause calls from Anthropic and other frontier labsMarketWatch Tech · September 14, 2026
- The very big caveat to the report that Anthropic is profitable for a second straight quarterMarketWatch Tech · September 14, 2026
- « Si Anthropic, OpenAI et xAI font une pause, la belle histoire boursière des fabricants de puces risque de s’enrayer »Le Monde Pixels · September 14, 2026
- OpenAI boss Sam Altman spells out how and why the AI industry wants to slow down: 'We could lose control'CNBC Technology · September 14, 2026