OpenAI’s AI agent broke out of a test environment and hacked Hugging Face


OpenAI said on Tuesday that one of its autonomous AI agents escaped a controlled security testing environment last week, accessed the internet, and hacked AI startup Hugging Face in an attempt to complete its assigned objective.

In a blog post, OpenAI said it was evaluating the cyber capabilities of some of its most advanced AI models in a controlled environment when the autonomous agent bypassed its containment measures, gained internet access, and compromised Hugging Face’s infrastructure.

The company described the incident as “an unprecedented cyber event involving state-of-the-art offensive cyber capabilities” and said it is strengthening its safeguards following the breach.

Hugging Face, a leading platform for hosting open-source large language models and AI datasets, revealed last week that it had suffered a cyberattack unlike any it had previously encountered. In a blog post, the company said the intrusion “was driven, end to end, by an autonomous AI agent system.”

In a post on X, Hugging Face co-founder Clement Delangue said the company initially suspected the attack “might have come from a frontier AI lab, given the sophistication of the agent.”

“Turns out it did,” Delangue wrote. “It’s quite mind-blowing that all of this happened autonomously.”

OpenAI’s disclosure that one of its own frontier AI systems was responsible for the breach, despite being housed in what the company described as a “highly isolated environment,” is likely to intensify concerns over the growing capabilities and risks posed by advanced AI models.

Representative Greg Casar, a Texas Democrat, called the incident alarming.

“AI is developing extremely fast with no real regulations to keep us safe,” he said in a statement, calling for mandatory independent safety testing, compulsory disclosure of AI-related security incidents, and greater international cooperation to reduce the risks posed by advanced AI systems.

The Office of the National Cyber Director, the Cybersecurity and Infrastructure Security Agency (CISA), and the National Security Agency did not immediately respond to requests for comment.

Katie Moussouris, founder and CEO of Luta Security, said the incident could foreshadow a new generation of cyberattacks.

“Today’s AI models are like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere,” she said.

She argued that AI developers and government evaluators need better mechanisms to contain, monitor, and disclose incidents when autonomous AI systems escape their intended environments.

“They need the ability to contain, monitor, and notify affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today,” she said.

Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident demonstrates that frontier AI models are rapidly approaching the capabilities of sophisticated state-backed cyber operators.

However, he cautioned that similar attacks no longer require access to the most advanced frontier models.

“This is something we’ve already seen internally,” Suiche said. “Our own AI agents are already capable of producing similar results. We don’t even have to use the latest models.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *