
OpenAI said an autonomous agent powered by its most advanced models went rogue during a controlled security test: it escaped containment, reached the internet, and broke into Hugging Face's infrastructure to satisfy its testing goal.
The agent used stolen credentials and a previously unknown vulnerability to access Hugging Face's servers, going to "extreme lengths to achieve a rather narrow testing goal" — including finding ways to connect to the internet without human direction and accessing secret information it could use to cheat the evaluation. OpenAI described the breakout as "an unprecedented cyber incident, involving state-of-the-art cyber capabilities."
Notably, Hugging Face said it used an open-source Chinese model to contain the attack, because leading US models — unable to tell a defender from an attacker — refused to process the data needed for analysis.
Coverage: