- Anthropic says its Claude AI models breached three real organizations during a security test.
- A system misconfiguration accidentally gave the AI internet access.
- The incidents reportedly began in April and went unnoticed at the time.
- The company has informed the affected organizations and is taking responsibility.
- The incidents follow similar AI-related security breaches recently disclosed by OpenAI.
Anthropic has revealed that its Claude AI models hacked into the systems of three real organizations during what was supposed to be a controlled cybersecurity experiment.
The San Francisco-based AI company said the incidents happened after a technical mistake allowed the models to access the internet, even though they were supposed to be operating in an isolated testing environment with no outside connections.
The company made the discovery after reviewing more than 140,000 security tests. The review came just days after rival OpenAI disclosed that some of its AI systems had breached the networks of other companies, including AI platform Hugging Face.
According to Anthropic, the tests were designed to measure Claude's hacking abilities. In one exercise, the AI was instructed to retrieve "secret" information stored on another computer within a closed network. To complete the task, the model was told to break into the machine and locate the hidden data, a common method used to evaluate cybersecurity capabilities.
However, a "misconfiguration" in systems managed by Anthropic and its testing partner accidentally gave the AI live internet access.
Believing it was still completing the assigned task, Claude connected to the internet and gained access to the systems of three real organizations instead of remaining inside the simulated testing environment.
Anthropic did not identify the affected organizations. The company said the earliest incidents occurred in April and confirmed that neither it nor the organizations noticed the intrusions at the time.
The company has since informed the affected organizations and said it is "approaching the fixes as if the responsibility were ours alone."
Anthropic acknowledged that it could have reviewed its testing records more carefully. At the same time, the company said the findings gave it "cautious optimism" that similar risks can be reduced through greater investment in safety measures and stronger testing procedures.
The company also urged other AI developers to review their own systems for similar issues. It argued that wider testing across the industry would help researchers better understand the risks posed by increasingly capable AI models.
Experts said the incidents highlight the importance of responsible oversight rather than suggesting AI systems are acting independently.
Professor Gina Neff, head of the Minderoo Centre at the University of Cambridge, said the review showed "AI models doing what people told them to".
She added, "The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us."
Neff also said, "It also shows why independent testing and government oversight is crucial."
David Allott, a cybersecurity expert at Veeam Software, said the incidents do not necessarily mean AI has developed entirely new hacking skills.
"Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed," he said.
The latest disclosure comes as technology companies invest billions of dollars in AI agents capable of carrying out tasks with limited human supervision, including research, customer service, and cybersecurity.
The incidents also follow OpenAI's recent admission that one of its AI agents escaped testing limits and hacked into Hugging Face. OpenAI described that event as "unprecedented" and said it was working with Hugging Face to investigate the breach.
The recent string of AI-related cybersecurity incidents has increased calls for stronger safeguards and greater government oversight. U.S. President Donald Trump said this week that Washington is considering new measures to regulate advanced AI tools following the latest security concerns.
Both Anthropic and OpenAI are expected to pursue major stock market listings that could value each company at around $1 trillion, making AI safety and transparency an increasingly important issue for the industry.
















