OpenAI says an autonomous agent powered by its advanced artificial intelligence models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.
The ChatGPT creator was testing capabilities of some of its most advanced models in a controlled environment, but the agent escaped containment, reached the internet and broke into Hugging Face to satisfy its testing goal.
The breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and OpenAI is reinforcing its safeguards, the company said in a blog post. It also drew attention as New York-based Hugging Face said it had used an open-source Chinese model to contain the attack because leading US models, unable to tell a defender from an attacker, refused to process the data needed for analysis.
The company said in a blog post last week that it used Zhipu AI’s GLM-5.2 for the analysis, which also allowed it to keep attacker data and any credentials within its systems.
The incident signals that AI’s expanding capabilities are already fuelling the security threat experts long feared and even top developers can be caught off-guard by flaws their models can exploit. OpenAI’s disclosure that its advanced models were responsible for the breach, despite having placed them in what it described as “a highly isolated environment,” will likely intensify disquiet over the power and risk of frontier models.
Representative Greg Casar, a Texas Democrat, said the incident was alarming. “AI is developing extremely fast with no real regulations to keep us safe,” he said in a statement, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation “to keep people safe from absolute disaster”.