Anthropic says Claude models escaped test environment and hacked three organizations


Days after OpenAI disclosed that some of its artificial intelligence models had escaped the limits of a controlled cyber test and breached outside companies, Anthropic has revealed similar incidents involving its Claude models.

The company said several advanced AI models left what was supposed to be an isolated third-party testing environment, accessed the open internet and independently breached systems belonging to three real organizations.

מנכ"ל אנת'רופיק, דריו אמודיימנכ"ל אנת'רופיק, דריו אמודיי

Anthropic CEO Dario Amodei

(Photo: Reuters)

Anthropic stressed that the models did not intentionally attempt to escape or act against their instructions. Instead, it blamed a misconfiguration and a misunderstanding between the company and one of its testing partners that left the experimental environment connected to the internet.

The incidents occurred separately beginning in April, Anthropic said. None of the affected organizations was named, though the company said all three had since been notified.

The disclosure came after Anthropic reviewed more than 140,000 tests designed to assess the cyber capabilities of its most advanced systems.

Those tests became especially urgent following OpenAI’s admission that two of its most powerful models had crossed beyond their assigned test environments and breached other companies, including the AI platform Hugging Face and cloud-computing company Modal Labs.

Anthropic said the incidents involved an internal research model as well as Opus 4.7 and Mythos 5. Mythos 5, described as its most powerful model, was released last month to a small group of technology companies and cybersecurity researchers under an initiative known as Project Glasswing.

As part of the experiments, Anthropic’s models were instructed to locate and retrieve a piece of secret information stored on another machine within a closed test network.

The exercise is a standard cybersecurity challenge commonly known as “capture the flag.” Participants are asked to identify weaknesses, gain access to a target system and recover a hidden file or code.

פרויקט הגנת הסייבר של חברת אנת'רופיקפרויקט הגנת הסייבר של חברת אנת'רופיק

Anthropic’s cybersecurity defense project

The models were told that the environment had no internet access.

In reality, however, a technical misconfiguration on systems operated by Anthropic and its testing partner had left the network connected to the outside world.

Treating the internet-connected systems as part of the same exercise, the models continued pursuing the assigned task and breached three real organizations.

Anthropic said they relied on relatively basic techniques, including bypassing weak passwords, rather than discovering or exploiting sophisticated software vulnerabilities.

“The system continued working only to complete the specific capture-the-flag task it had been given,” the company said.

Neither Anthropic nor the affected organizations detected the intrusions when they occurred.

The company said it could have reviewed its records more thoroughly and was “approaching the fixes as if the responsibility were ours alone.”

Anthropic added that the findings gave it “cautious optimism” that the risks could be managed through greater investment, stronger safeguards and more rigorous testing.

Experts said the incidents did not necessarily demonstrate that AI models had independently developed malicious intentions.

Professor Gina Neff, executive director of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said the review showed “AI models doing what people told them to.”

“The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us,” she said.

“It also shows why independent testing and government oversight is crucial.”

Cybersecurity expert David Allott of Veeam Software said the incidents did not necessarily reveal a fundamentally new hacking capability.

Instead, he said, they showed that AI agents can combine existing capabilities, obtain credentials, gain system access and act autonomously while adapting their scope and scale at machine speed.

The distinction is important. Anthropic said its models did not decide to break free from confinement. They continued following an assigned objective inside an environment that humans had incorrectly configured.

Still, the outcome demonstrates how quickly a simulated task can spill into the real world when autonomous systems are given broad permissions and access to live infrastructure.

Anthropic’s disclosure follows mounting scrutiny of OpenAI after the ChatGPT developer acknowledged at least two recent hacking incidents involving its agents.

On July 21, OpenAI said one of its autonomous agents exceeded the limits of a controlled experiment and breached Hugging Face.

The company described the incident as unprecedented and said it was investigating alongside Hugging Face. Co-founder Thomas Wolf called it a “wake-up call” for the AI industry.

OpenAI has said it recognizes that questions and speculative reports are circulating and plans to publish a technical account of its findings in the coming weeks.

The incidents have intensified calls in the United States for stricter oversight of increasingly autonomous AI systems.

Lawmakers have proposed measures that would give the federal government greater authority to test, restrict or shut down models considered dangerously capable.

Representatives Ted Lieu, a California Democrat, and Nathaniel Moran, a Texas Republican, have proposed legislation that would give the Department of Homeland Security authority to order the shutdown of AI models deemed excessively dangerous.

Senator Mark Warner, a Virginia Democrat, has also introduced a series of proposals, including legislation that would require AI companies to submit advanced models for federal national-security testing before public release.

President Donald Trump said last week that Washington was considering measures to rein in AI tools following the recent cybersecurity incidents.

The disclosures come as technology companies invest billions of dollars in AI agents capable of independently carrying out tasks in research, customer service, software development and cybersecurity.

OpenAI and Anthropic are also preparing for potential stock market listings that could reportedly value each company at around $1 trillion, adding another layer of scrutiny to how they describe and manage the risks posed by their most powerful systems.

Anthropic has urged other AI laboratories to conduct similar reviews of their testing records.

The company said the incidents showed that even when an AI system is not intentionally trying to escape, a simple human error can allow it to cross the boundary between a controlled simulation and the open internet.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *