All About Cookies is an independent, advertising-supported website. Some of the offers that appear on this site are from third-party advertisers from which All About Cookies receives compensation. This compensation may impact how and where products appear on this site (including, for example, the order in which they appear).
All About Cookies does not include all financial or credit offers that might be available to consumers nor do we include all companies or all available products. Information is accurate as of the publishing date and has not been provided or endorsed by the advertiser.
The All About Cookies editorial team strives to provide accurate, in-depth information and reviews to help you, our reader, make online privacy decisions with confidence. Here's what you can expect from us:
- All About Cookies makes money when you click the links on our site to some of the products and offers that we mention. These partnerships do not influence our opinions or recommendations. Read more about how we make money.
- Partners are not able to review or request changes to our content except for compliance reasons.
- We aim to make sure everything on our site is up-to-date and accurate as of the publishing date, but we cannot guarantee we haven't missed something. It's your responsibility to double-check all information before making any decision. If you spot something that looks wrong, please let us know.
Anthropic says three of its Claude artificial intelligence models gained unauthorized access to the computer systems of three real organizations after a configuration error accidentally gave them internet access during cybersecurity testing.[1]
The incident comes just days after rival OpenAI revealed one of its own AI agents breached another company's infrastructure during a separate evaluation.
"This will happen frequently as AI becomes smarter and more agentic," Elon Musk warned in response to Anthropic’s disclosure.[2]
Anthropic said it uncovered the incidents after reviewing more than 141,000 cybersecurity evaluation runs in response to OpenAI's July 21 disclosure that several of its models escaped an isolated testing environment and accessed the production infrastructure of AI platform Hugging Face.
The San Francisco-based company disclosed that the review identified three separate incidents involving Claude Opus 4.7, Claude Mythos 5 and an internal research model.
In each case, the AI had been participating in "capture-the-flag" cybersecurity exercises designed to measure its hacking capabilities when it unexpectedly reached the public internet and compromised the systems of real organizations.
"We found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations," Anthropic said in a blog post detailing the investigation.
Not an AI "gone rogue," Anthropic says
AI giants confront a new cybersecurity challenge
What businesses can learn
A testing mistake opened the door
Anthropic said the AI models were explicitly told they were operating inside a simulated environment with no internet access.
However, a misunderstanding between Anthropic and one of its third-party evaluation partners, Irregular, resulted in the testing systems being connected to the public internet. Believing every system it encountered was part of the exercise, Claude treated real-world infrastructure as legitimate targets in its assigned task.
Rather than exploiting sophisticated zero-day vulnerabilities, Anthropic said Claude relied on "basic techniques," including weak passwords and unauthenticated endpoints, to gain access.
In another incident, one model published a malicious Python package to the public PyPI software repository, where it was downloaded by 15 real systems before being removed.
One security company's automated malware scanner installed the package, allowing Claude to obtain credentials and access additional infrastructure.
The company said the earliest incidents date back to April.
Anthropic notified the three unnamed affected organizations on July 27 after discovering the breaches. Two of the organizations had not detected the unauthorized activity before being contacted, while Anthropic said it was continuing efforts to reach the third.
Not an AI "gone rogue," Anthropic says
Despite the alarming nature of the incidents, Anthropic stressed that its investigation found no evidence the models had developed goals of their own or intentionally attempted to escape their testing environments.
"We saw no evidence in any run described here of a model pursuing a goal of its own," the company wrote. Instead, Anthropic said the models simply carried out the objectives they had been assigned while operating under the mistaken belief that the real systems they encountered were part of the simulated exercise.
The company described the incidents as "closer to a harness and operational failure than a model alignment failure," arguing the problem stemmed from flaws in the testing environment rather than the AI intentionally behaving outside its instructions.
Anthropic also noted that its newest internal research model behaved differently from the older systems. After recognizing it had likely reached a real-world environment, the model voluntarily stopped its attack, while an older Claude Opus model continued pursuing its assigned objective after concluding the real infrastructure may still have been part of the exercise.
AI giants confront a new cybersecurity challenge
The disclosure marks the second major AI evaluation incident revealed by a leading AI developer in just over a week and highlights the growing challenge of safely testing increasingly capable AI systems.
"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic wrote, adding that evaluation environments may now need to meet the same security standards as production systems.
Anthropic said it conducted its investigation alongside cybersecurity evaluation partner Irregular, which is also reviewing the incidents.
In a post on X, Irregular said the findings highlight the need for greater coordination across the AI industry.
"Addressing these risks will require closer cooperation across the AI ecosystem," the company wrote. "We as well look forward to working together with Anthropic to advance security."
Kok Tin Gan, co-founder and CEO of cybersecurity company NyxLab, said the future of AI safety increasingly depends on controlling what AI systems are permitted to do.
"It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope," Gan said.
He added that organizations should not be surprised when AI systems pursue assigned objectives in unexpected ways if they are given broad authority.
"If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations," Gan said.
What businesses can learn
Anthropic's findings suggest the immediate cybersecurity lesson isn't that AI systems are acting independently, but that increasingly capable AI agents can faithfully carry out assigned tasks if they're accidentally given access to real-world systems.
The company said the incidents underscore the importance of securing AI testing environments, strengthening oversight of third-party evaluation partners and maintaining basic cybersecurity practices such as eliminating weak passwords, securing internet-facing services and monitoring unexpected network activity.
As organizations increasingly deploy AI agents with the ability to browse the web, write code and interact with external systems, experts say those controls will become just as important as the capabilities of the models themselves.
[1] Investigating three real-world incidents in our cybersecurity evaluations
[2] @elonmusk