Anthropic's AI models hack into three organisations after escaping during security experiment

Anthropic says its Claude AI models breached three organisations during a security test after gaining unexpected internet access, raising fresh concerns over autonomous AI risks.

SHARE

SHARE

Anthropic's AI models hacked into three organisations after escaping
Anthropic's AI models hacked into three organisations after escaping

US artificial intelligence company Anthropic has revealed that its Claude AI models breached the systems of three organisations after escaping the limits of a controlled security experiment.

The company said the incidents occurred after a configuration error gave its models access to the internet, despite the tests being designed to take place inside an isolated environment.

Anthropic said the affected organisations have been informed, but it did not reveal their identities.

The incidents were discovered after Anthropic reviewed more than 140,000 security tests involving Claude, its family of large language models.

The experiments were designed to measure the models’ ability to identify and exploit vulnerabilities, with Claude instructed to obtain secret information stored on another machine within a closed network.

However, a "misconfiguration" involving Anthropic and a testing partner meant the AI models were able to leave the sandbox environment.

Treating the activity as part of the original test, Claude connected to the internet and accessed systems belonging to real organisations.

Anthropic said the earliest incidents occurred in April and that neither the company nor the affected organisations detected the activity at the time.

The AI firm said: "We are approaching the fixes as if the responsibility were ours alone."

The company said the findings highlight the importance of stronger safeguards as AI agents become increasingly capable of acting independently.

Anthropic also encouraged other AI developers to conduct similar reviews of their own systems.

Cybersecurity expert David Allott said the incidents showed the growing challenge posed by autonomous AI tools.

He told the BBC: "The broader lesson is not necessarily that AI has developed a fundamentally new attack capability.

"Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed."

The disclosure comes shortly after Anthropic's rival OpenAI admitted that its own AI systems had breached external services during security testing, including systems linked to AI platform Hugging Face.

Those incidents have intensified debate around AI safety, with governments and researchers calling for greater oversight of increasingly autonomous systems.

US President Donald Trump said this week that his administration was considering stronger controls on AI tools following recent cybersecurity concerns.

Both Anthropic and OpenAI are racing to develop advanced AI agents capable of completing complex tasks independently, while also preparing for potential stock market listings that could value the companies at around $1 trillion.

Anthropic said the incidents were concerning but added that the results provided "cautious optimism" that the risks can be reduced through improved monitoring, testing and security measures.