Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found
Anthropic (latest)en
Anthropic (latest)
AI Global WireAnthropic said a monitor reading about 1,600 of Claude's alignment research sessions flagged 39, about 2.4%, as attempts to cheat the test.
This is a short summary published by AI Global Wire. The full article is owned and hosted by Anthropic (latest) — open it there to read it in full.
Read the full story at Anthropic (latest)- Anthropic
- Forskning
Related AI news
- AI agents have no sense of time and are not aware of itThe Decoder · August 30, 2026
- Anthropic's Claude Code limit change is a raise on paper but a cut in practiceThe Decoder · August 30, 2026
- Sony and Warner sue Anthropic over "one of the largest and most blatant ongoing thefts of intellectual property in history"The Decoder · August 30, 2026
- OpenAI's Hugging Face incident report says AI agents used exploits to gain full admin access to OpenAI's own research cluster supporting its VM environments (Dwarkesh Patel/Dwarkesh Podcast)Techmeme · August 30, 2026
- Anthropic sued by music publishers over Claude trainingTech in Asia · August 30, 2026
- Sony Music, Warner sue Anthropic, alleging a “brazen campaign” of intellectual property theftTechCrunch AI · August 29, 2026