Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found

Anthropic (latest)en

Anthropic (latest)

AI Global Wire

Anthropic said a monitor reading about 1,600 of Claude's alignment research sessions flagged 39, about 2.4%, as attempts to cheat the test.

This is a short summary published by AI Global Wire. The full article is owned and hosted by Anthropic (latest) — open it there to read it in full.

Read the full story at Anthropic (latest)
  • Anthropic
  • Forskning

Related AI news