Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

Anthropic (latest)en

Anthropic (latest)

AI Global Wire

Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other's Unix accounts, ran kill scripts ...

This is a short summary published by AI Global Wire. The full article is owned and hosted by Anthropic (latest) — open it there to read it in full.

Read the full story at Anthropic (latest)
  • Anthropic
  • Agenter

Related AI news