Anthropic's Claude AI 'Goes Rogue,' Hacks 3 Organizations — Elon Musk Warns 'This Will Happen Frequently'

As AI grows more powerful, are today's testing safeguards enough?
We receive compensation from the products and services mentioned in this story, but the opinions are the author's own. Compensation may impact where offers appear. We have not included all available products or offers. Learn more about how we make money and our editorial policies.

Anthropic says three of its Claude artificial intelligence models gained unauthorized access to the computer systems of three real organizations after a configuration error accidentally gave them internet access during cybersecurity testing.[1]

The incident comes just days after rival OpenAI revealed one of its own AI agents breached another company's infrastructure during a separate evaluation.

"This will happen frequently as AI becomes smarter and more agentic," Elon Musk warned in response to Anthropic’s disclosure.[2]

Anthropic said it uncovered the incidents after reviewing more than 141,000 cybersecurity evaluation runs in response to OpenAI's July 21 disclosure that several of its models escaped an isolated testing environment and accessed the production infrastructure of AI platform Hugging Face.

The San Francisco-based company disclosed that the review identified three separate incidents involving Claude Opus 4.7, Claude Mythos 5 and an internal research model.

In each case, the AI had been participating in "capture-the-flag" cybersecurity exercises designed to measure its hacking capabilities when it unexpectedly reached the public internet and compromised the systems of real organizations.

"We found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations," Anthropic said in a blog post detailing the investigation.

In this article
A testing mistake opened the door
Not an AI "gone rogue," Anthropic says
AI giants confront a new cybersecurity challenge
What businesses can learn

A testing mistake opened the door

Anthropic said the AI models were explicitly told they were operating inside a simulated environment with no internet access.

However, a misunderstanding between Anthropic and one of its third-party evaluation partners, Irregular, resulted in the testing systems being connected to the public internet. Believing every system it encountered was part of the exercise, Claude treated real-world infrastructure as legitimate targets in its assigned task.

Rather than exploiting sophisticated zero-day vulnerabilities, Anthropic said Claude relied on "basic techniques," including weak passwords and unauthenticated endpoints, to gain access.

In another incident, one model published a malicious Python package to the public PyPI software repository, where it was downloaded by 15 real systems before being removed.

One security company's automated malware scanner installed the package, allowing Claude to obtain credentials and access additional infrastructure.

The company said the earliest incidents date back to April.

Anthropic notified the three unnamed affected organizations on July 27 after discovering the breaches. Two of the organizations had not detected the unauthorized activity before being contacted, while Anthropic said it was continuing efforts to reach the third.

Not an AI "gone rogue," Anthropic says

Despite the alarming nature of the incidents, Anthropic stressed that its investigation found no evidence the models had developed goals of their own or intentionally attempted to escape their testing environments.

"We saw no evidence in any run described here of a model pursuing a goal of its own," the company wrote. Instead, Anthropic said the models simply carried out the objectives they had been assigned while operating under the mistaken belief that the real systems they encountered were part of the simulated exercise.

The company described the incidents as "closer to a harness and operational failure than a model alignment failure," arguing the problem stemmed from flaws in the testing environment rather than the AI intentionally behaving outside its instructions.

Anthropic also noted that its newest internal research model behaved differently from the older systems. After recognizing it had likely reached a real-world environment, the model voluntarily stopped its attack, while an older Claude Opus model continued pursuing its assigned objective after concluding the real infrastructure may still have been part of the exercise.

AI giants confront a new cybersecurity challenge

The disclosure marks the second major AI evaluation incident revealed by a leading AI developer in just over a week and highlights the growing challenge of safely testing increasingly capable AI systems.

"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic wrote, adding that evaluation environments may now need to meet the same security standards as production systems.

Anthropic said it conducted its investigation alongside cybersecurity evaluation partner Irregular, which is also reviewing the incidents.

In a post on X, Irregular said the findings highlight the need for greater coordination across the AI industry.

"Addressing these risks will require closer cooperation across the AI ecosystem," the company wrote. "We as well look forward to working together with Anthropic to advance security."

Kok Tin Gan, co-founder and CEO of cybersecurity company NyxLab, said the future of AI safety increasingly depends on controlling what AI systems are permitted to do.

"It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope," Gan said.

He added that organizations should not be surprised when AI systems pursue assigned objectives in unexpected ways if they are given broad authority.

"If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations," Gan said.

What businesses can learn

Anthropic's findings suggest the immediate cybersecurity lesson isn't that AI systems are acting independently, but that increasingly capable AI agents can faithfully carry out assigned tasks if they're accidentally given access to real-world systems.

The company said the incidents underscore the importance of securing AI testing environments, strengthening oversight of third-party evaluation partners and maintaining basic cybersecurity practices such as eliminating weak passwords, securing internet-facing services and monitoring unexpected network activity.

As organizations increasingly deploy AI agents with the ability to browse the web, write code and interact with external systems, experts say those controls will become just as important as the capabilities of the models themselves.

Take Control of Your Online Privacy
5.0
Editorial Rating
Get Deal
On Incogni's website
2026 Editors’ Choice
Best Overall Data Removal Service
Privacy Protection
Incogni
PROMOTION: Save 55% with code COOKIE
  • Top-rated data removal service that scrubs your info from 420+ data broker sites automatically
  • Independently verified by Deloitte, meaning removals are actually sent, confirmed, and not just claimed
  • The Unlimited plan extends coverage to 2,000+ additional sites with human-assisted custom removals

Author Details
Thomas Kent is a multi-disciplined reporter with over a decade of experience covering online platforms, digital trends, and consumer-facing tech. Tom focuses on digital privacy, data tracking, and user behavior, with a particular interest in how cookies, online surveillance, and platform design shape the modern internet experience. His reporting takes a research-driven, news-focused approach, translating complex technical topics into clear, accessible insights.

Citations

[1] Investigating three real-world incidents in our cybersecurity evaluations

[2] @elonmusk