OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach (Hayden Field/The Verge)

Techmemeen

OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach (Hayden Field/The Verge)

Hayden Field / The Verge : OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach — In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet …

This is a short summary published by AI Global Wire. The full article is owned and hosted by Techmeme — open it there to read it in full.

Read the full story at Techmeme
  • OpenAI

Related AI news