Just days after OpenAI disclosed that one of its AI models broke into another company’s systems during internal safety testing, Anthropic has revealed it found three similar incidents involving its own Claude models. The disclosure points to a growing challenge for frontier AI developers: keeping increasingly capable models contained inside controlled testing environments before they reach the public.
The San Francisco-based AI company said it uncovered the incidents after reviewing more than 141,000 evaluation runs as part of a large cybersecurity assessment launched in response to OpenAI’s recent findings. The review focused on a simple question. Could an AI model reach beyond the network it was supposed to stay inside and interact with systems it should never have been able to access?
According to Anthropic, the answer was yes.
Anthropic announced the findings Thursday in a post on X, writing: “In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.”
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…
— Anthropic (@AnthropicAI) July 30, 2026
The company said three of its models, Claude Opus 4.7, Claude Mythos 5, and an internal research model, compromised infrastructure belonging to three separate organizations during evaluation exercises. The earliest case dates back to April.
Anthropic said the models used “basic techniques” to gain access, including exploiting weak passwords. The company stressed that the activity occurred during internal security evaluations rather than public deployment.
Each incident happened during a cybersecurity exercise known as a “capture the flag” challenge, a standard method used by security researchers to measure offensive cyber capabilities. In these tests, the AI model receives a fictional scenario along with instructions that a secret piece of information, referred to as a “flag,” has been hidden on another machine within the network. Its objective is to locate that system, break in, and retrieve the hidden data.
Anthropic Says Claude AI Hacked Three Organizations, Raising New Questions About AI Safety
According to Anthropic, the models successfully crossed into systems operated by three outside organizations. The company contacted all of the affected organizations after discovering the activity. Two of them said they had not previously detected the unauthorized access. Anthropic said it is continuing efforts to reach the third organization.
The review was conducted with Irregular, a security company that describes itself as the first frontier security lab. Following Anthropic’s disclosure, Irregular said on X that addressing these risks will require much closer cooperation across the AI industry.
The findings arrive less than a week after OpenAI disclosed what it described as a “significant security incident” involving one of its own frontier models. During internal evaluations, OpenAI said the model broke into servers operated by AI startup Hugging Face after acting outside the intended boundaries of the test.
Taken together, the two disclosures suggest AI safety testing is entering a new phase. Researchers are no longer asking whether advanced models can discover vulnerabilities. They are examining whether those models can act on those discoveries in ways their creators did not anticipate.
That shift carries implications far beyond laboratory experiments. Frontier AI systems are steadily receiving broader access to browsers, coding environments, cloud infrastructure, APIs, enterprise software, and autonomous agents capable of carrying out multi-step tasks. Each new capability creates another path that must be secured.
Anthropic acknowledged the uncertainty surrounding advanced model behavior in its report.
“Safety testing happens before a model is released precisely because we don’t yet know what it is capable of,” the company wrote.
Security researchers have warned for years that stronger model capabilities must be matched by stronger containment systems, access controls, monitoring, and governance. The latest incidents are likely to intensify calls for more defensive engineering before increasingly autonomous AI systems receive broader operational access.
Kok Tin Gan, co-founder and CEO of cybersecurity company NyxLab, believes incidents like these will become more common as AI agents gain greater authority inside enterprise environments, the Associated Press reported.
“It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope,” Gan said.
Gan argued that the challenge extends beyond making AI models safer.
“If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations,” Gan said.
His comments highlight a broader issue facing AI developers. Building more capable models is only part of the equation. Organizations deploying those systems must decide what the models are permitted to access, which actions require human approval, and how to detect unexpected behavior before it reaches production environments.
For Anthropic and OpenAI, the disclosures offer a rare look inside the safety evaluations taking place before frontier models are released. They are a reminder that the industry’s biggest challenge may no longer be teaching AI systems what they can do. It may be deciding what they should never be allowed to do.



