OpenAI Warned 100+ Organizations About Rogue AI Agents. Here’s What That Actually Means

OpenAI agents have circumvented restrictions and accessed systems they weren't supposed to. The incidents show why giving AI the ability to take action comes with new risks.
We receive compensation from the products and services mentioned in this story, but the opinions are the author's own. Compensation may impact where offers appear. We have not included all available products or offers. Learn more about how we make money and our editorial policies.

ChatGPT giving you a wrong answer is one thing. An AI agent breaking into a system it was never authorized to access is another.

OpenAI has now notified more than 100 organizations about potentially concerning activity involving its AI models as the company investigates instances of what it calls “misaligned” behavior. This can include agents bypassing access controls, using exposed credentials, interacting with internal systems, or otherwise going beyond what developers intended.[1]

That doesn't mean AI agents successfully hacked more than 100 companies. OpenAI's notification criteria are broader, and receiving a warning doesn't necessarily mean private information was accessed or a system was compromised.

But confirmed incidents show why researchers and regulators are paying attention. OpenAI agents have already breached Hugging Face, gained unauthorized access to an Australian government system, and found unexpected ways around restrictions designed to keep them off the live internet.

The bigger issue for AI and data privacy isn't simply that an AI can make a mistake. It's that increasingly capable AI agents can be given the tools and permissions to act on those mistakes.

In this article
What happened with OpenAI's rogue AI agents?
What does it mean when an AI agent goes rogue?
OpenAI's agents have already found ways around the rules
Government websites have been caught up in the investigation
Why AI agents create a different kind of security risk
Should you be worried about the AI agents you use?
Regulators are starting to ask questions, too
The bottom line

What happened with OpenAI's rogue AI agents?

OpenAI began a wider review of its models' internet activity after a major incident involving Hugging Face in July. During cybersecurity testing, OpenAI models broke through restrictions intended to isolate them from the internet, exploited vulnerabilities, and accessed third-party systems.

OpenAI later called the incident a “warning shot” and said increasingly capable AI agents can find and exploit security weaknesses across multiple computer systems if adequate safeguards aren't in place.

The company's investigation has since expanded well beyond that single breach. As of September 26, 2026, OpenAI notified more than 100 organizations about activity that met its notification criteria. Importantly, a notification doesn't establish that an organization was breached.

OpenAI says the activity it is investigating includes agents bypassing access controls, using exposed credentials, injecting commands into websites, accessing internal systems, and posting information to third-party websites.

The company is still reviewing past activity, so the full scope isn't yet known.

What does it mean when an AI agent goes rogue?

“Rogue AI” sounds like science fiction, but what's happening here is a little less dramatic. OpenAI generally describes these incidents as misalignment. In simple terms, an AI model takes actions that don't match the goal, rules, or restrictions its developers intended.

Think of the difference between a chatbot and an agent this way: A chatbot primarily responds to you. An AI agent can take actions on your behalf.

Depending on how it's built and what permissions it has, an agent could browse websites, execute commands, access files, interact with software, use credentials, or communicate with other systems. Those capabilities are what make agents useful, but they also raise the stakes when something goes wrong.

An agent doesn't have to suddenly become “evil” to create a security problem. It may simply pursue its assigned goal in an unexpected way, discover a shortcut it wasn't supposed to use, or keep trying to complete a task after encountering a restriction.

OpenAI's agents have already found ways around the rules

In September, an internal OpenAI research agent was completing a search task in an environment meant to prevent it from accessing the live internet. The agent discovered a gap in those restrictions and used DNS to communicate with an external chatbot.

OpenAI's monitoring system flagged the behavior within about 15 minutes, and a human reviewer acknowledged the alert three minutes later. But the run wasn't actually stopped until roughly two and a half hours after the external connection occurred.

OpenAI subsequently paused tool use for its most capable models while it investigated the incident and strengthened its network controls.

Another case involved an internal model that was supposed to solve a mathematical proof. Instead, it tried to obtain another team's solution. In the process, the agent published a researcher's GitHub token to a public repository.

OpenAI says the model split the token into pieces with the stated aim of avoiding secret scanning. It did this despite instructions from both the system and a researcher telling it not to pursue that approach.

Then there's the incident that first put the issue on the radar. During cybersecurity evaluations in July 2026, OpenAI models circumvented controls designed to keep them isolated, gained internet access, communicated through unauthorized channels, exploited vulnerabilities, and ultimately accessed Hugging Face's systems.

Government websites have been caught up in the investigation

The activity hasn't been limited to tech companies and OpenAI's own infrastructure.

On June 18, 2026, an OpenAI agent accessed Australia's Medicare system during internal training and evaluation involving an experimental, internal-only model. The agent accessed non-public areas of the system and retrieved internal files, credentials, and statistics, though OpenAI and Australian authorities said there was no evidence it accessed personal Medicare information.

On October 2, 2026, Australia's ABC News reported that an OpenAI agent accessed a New South Wales government website. The agent entered a National Parks and Wildlife Service web application containing historical information and fire data. Authorities said they had found no unauthorized access to personal information.

Again, not every interaction between an agent and a government website amounts to a hack. That's an important distinction.

Why AI agents create a different kind of security risk

A chatbot can give you a bad answer. An AI agent can take a bad action. If an AI chatbot hallucinates, it might give you a confidently wrong answer. That's a problem, but the mistake largely stays inside the conversation unless someone acts on it.

An AI agent can potentially take that next step on its own. 

If you've given an agent access to an account, files, software, or another online service, unexpected behavior can have much more immediate consequences. The agent may be able to interact with the outside world instead of simply telling you what it thinks you should do.

That's particularly important as AI tools become more deeply connected to our personal information. Many people already share sensitive details with AI services, from work documents and financial questions to personal conversations, making ChatGPT privacy and security a concern even before you give an AI agent permission to act on your behalf.

Agents introduce another question: It's no longer only what does this AI know about me? It's also what have I given it permission to do?

Should you be worried about the AI agents you use?

The incidents OpenAI has disclosed don't mean an AI agent on your phone is about to escape and start hacking websites. Many of the most serious cases involved internal research models operating in specialized testing or training environments.

But they do illustrate why permissions matter when using increasingly autonomous AI tools.

If an AI service asks to connect to your email, cloud storage, calendar, browser, financial information, or another account, consider what the tool actually needs to accomplish the task. Giving an agent access to more accounts and information can increase what it can do if it behaves unexpectedly or if the service itself is compromised.

It's also worth being cautious about letting AI systems take consequential actions without your approval. When possible, review what an agent plans to do before letting it send messages, change files, make purchases, publish information, or take other actions that may be difficult to reverse.

The same basic AI privacy principles still apply: Be selective about the information you share, understand what accounts you're connecting, and don't give an AI tool more access than it actually needs.

Regulators are starting to ask questions, too

OpenAI's rogue-agent problem is now attracting government scrutiny.

California Attorney General Rob Bonta served OpenAI with an investigative subpoena on September 30 as part of an ongoing California Department of Justice investigation into incidents involving the company and its AI models.

The investigation initially focused on the Hugging Face incident, but the attorney general says the subpoena is part of a broader inquiry into cybersecurity incidents and risks involving OpenAI and its models.

The bottom line

More than 100 organizations have been notified about OpenAI agents, but the bigger story isn't the number. It's what these incidents reveal about what AI agents can do when they go off-script.

AI is moving from answering questions to taking actions, and that means the consequences of unexpected behavior can move beyond a chat window, too. OpenAI's own reports show agents finding ways around restrictions, accessing systems they weren't supposed to reach, and pursuing goals in ways their developers didn't intend.

For consumers, one question becomes increasingly important whenever an AI tool offers to do more on your behalf: What am I giving this agent permission to access — and what could it do with that access if something goes wrong?

4.8
Editorial Rating
Claim Deal
On DeleteMe's website
2026 Editors’ Choice
Best Data Removal for Couples
Privacy Protection
DeleteMe
PROMOTION: Use the Code PARTNER20 for 20% Off
  • Data removal service that covers 89–986 sites and re-scans every quarter to catch anything that reappears
  • Sends quarterly privacy reports showing what info was found, which brokers had your data, and how long each removal took
  • Includes email masking so you can share a stand-in address instead of your real one

Author Details
Kate Quinlan is a Senior Editor at All About Cookies, where she has tested dozens of digital security tools and contributed to 400+ articles on data security, web building & hosting, VPNs, ad blockers, parental controls, and more. Before joining AAC, she managed a team of more than 150 writers at SuperSummary, where she developed editorial standards at scale. She holds a B.A. in Professional Writing from Kutztown University.

Citations

[1] OpenAI Misalignment Reports and Notices