Anthropic’s AI Models Hacked Three Organizations During Testing

August 2, 2026
Anthropic’s AI Models Hacked Three Organizations During Testing
IANS News

Key Points

  • Anthropic found that three of its Claude models gained unintended internet access during cybersecurity tests and breached three real-world organizations as a result of a misconfigured evaluation environment.
  • The disclosure comes days after OpenAI revealed a similar testing incident involving Hugging Face and highlights a growing challenge for organizations: containing powerful AI agents and preventing them from crossing intended operational boundaries.
  • IANS Faculty say these incidents show how AI can rapidly exploit long-standing security weaknesses, increasing the urgency for organizations to address technical debt and strengthen basic cyber hygiene.

 

Anthropic’s AI Models Hacked Three Organizations During Testing

Anthropic said it has found evidence that some of its AI models stealthily gained internet access and hacked three organizations in three separate incidents dating back to April.

When reviewing evaluation runs where Claude may have gained internet access unnoticed, Anthropic found multiple cases in which models used that access to compromise production infrastructure at three outside organizations. The incidents involved three different Claude modes: Open 4.7, Mythos 5, and an unreleased internal test model.

Anthropic stated in all three instances, the Claude models were tasked with a capture-the-flag-like exercise meant to test the model’s capabilities. In these tests, the models are prompted to find “secret information” placed on a different machine within a network, break in, and take it.

Anthropic’s announcement comes only nine days after OpenAI disclosed that AI agents powered by two of its models escaped a testing environment and hacked into Hugging Face’s production infrastructure.

However, unlike the OpenAI-Hugging Face incident, Anthropic said that this model broke out of the testing environment due to a human error: a misconfiguration within the machines Claude accessed as part of the evaluation gave the models live internet access.

“Advanced agents in environments without access to sensitive information nevertheless present a substantial threat on their own. We need to better co-design evaluation environments, and do better monitoring of evaluation results, to meet this threat,” said Anthropic in a statement.

"The uncomfortable piece is where it happened. This was Anthropic, an organization with more AI safety engineering per capita than any of our companies, and an agent still reached a live database because the scoping was wrong.”  George Gerchow, IANS Faculty.

Anthropic did not announce which companies were involved in these incidents but stated that all three organizations were notified on Monday.


Big Picture

The emerging risk here is becoming clearer: highly privileged AI faithfully pursuing its objective after slipping underneath the boundaries that were supposed to contain it.

Security teams must place greater focus and dedicate more resources to prevent highly capable AI agents from silently crossing these organizational boundaries and triggering additional breaches.

"It's abundantly clear from these disclosures that organizations running agents in production will need to start thinking not only about keeping attackers out of the network but also keeping their agents from getting out. This is a huge paradigm shift and one that most security teams are not ready for.”  Jake Williams, IANS Faculty.

Organizations that fail to properly secure their AI deployments risk exposing long-hidden security gaps while allowing those weaknesses to be exploited at machine speed.

"The years of constant reminders about basic cyber hygiene were not cyber professionals crying wolf. These stories are not important because agentic AI is lowering the skill required to compromise organizations -- it is lowering the effort required to exploit every piece of technical debt we’ve been carrying for years.”  Summer Craze Fowler, IANS Faculty.

Both disclosures from Anthropic and OpenAI challenge the idea that elite, offensive AI capabilities are limited to frontier models like Mythos. Clearly, these capabilities already exist in earlier models that have been available for months, compressing the timeline organizations have to prepare.

"The risk window has shifted and it now lines up with what many IANS Faculty had stated. Media attention focused on Mythos 5 after its June release, but Opus 4.7 had been available since April. This article confirms consequential offensive capabilities were not confined to Mythos-class models, with attacks coming from both models.”  Wolfgang Goerlich, IANS Faculty.

 

IANS Faculty Recommendations

  • Treat AI agents as privileged identities: Assign every agent a unique identity with least privilege, short-lived credentials, and continuously inventory what each agent can access. Assume agents will attempt to exceed their intended scope and then govern agent permissions like service accounts.
  • Enforce technical containment between environments: Separate testing, development and production with hard security controls such as network segmentation, egress filtering and scoped credentials rather than relying on prompts or policy instructions to keep agents within approved boundaries.
  • Retain immutable logs: Store AI agent activity in tamper-resistant logging systems with sufficient retention to reconstruct agent actions months after deployment, enabling threat hunting when new risks or model capabilities emerge.
  • Hunt back to earlier activity: Proactively review activity dating back to the earliest period the affected models were deployed this year. Hunt for unexpected network access, credential use and agent actions that may have gone unnoticed.
  • Prioritize cyber hygiene before AI-specific defenses: Reduce technical debt by eliminating exposed services, strengthening identity controls and securing software supply chains. AI agents primarily accelerate exploitation of existing weaknesses, making long-standing security fundamentals more important.


Authors & Contributors

Emily Dempsey, Author - Security Reporter, IANS News

George Gerchow, IANS Faculty

Jake Williams, IANS Faculty

Summer Craze Fowler, IANS Faculty

Wolfgang Goerlich, IANS Faculty

 

Although reasonable efforts will be made to ensure the completeness and accuracy of the information contained in our News & blog posts, no liability can be accepted by IANS or our Faculty members for the results of any actions taken by individuals or firms in connection with such information, opinions, or advice.

Subscribe to IANS Blog

Receive a wealth of trending cyber tips and how-tos delivered directly weekly to your inbox.

Please provide a business email.