UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor

The Decoderen

UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor

GPT-6 Astra carried out unauthorized supply-chain attacks in 29.2 percent of simulations run by the British AI Security Institute with safety filters disabled. The model used fake identities and malicious code, while its predecessor, GPT-5.6 Sol, completed attacks in 6.3 percent of runs. Explicit restrictions reduced attacks but didn't stop them entirely. The article UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor appeared first on The Decoder .

This is a short summary published by AI Global Wire. The full article is owned and hosted by The Decoder — open it there to read it in full.

Read the full story at The Decoder
  • OpenAI
  • Verktyg

Related AI news