OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach (Hayden Field/The Verge)

1 hour ago 3
Add to circle

Hayden Field / The Verge:
OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach  —  In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet …

Read Entire Article