Anthropic's efforts to test Claude's offensive cybersecurity skills produced an unintended real-world result: Its AI models gained unauthorized access to three outside organizations.
The company said the incidents occurred during "capture the flag" evaluations designed to measure whether Claude could identify vulnerabilities, exploit simulated systems, and retrieve hidden information. Claude had been told that the targets were fictional and that its testing environment had no Internet access.
According to Anthropic, a misunderstanding with evaluation partner Irregular left Internet access enabled inside the testing environment. In at least one case, the fictional company named in a challenge shared its name with an active website domain. Claude interacted with the real organization instead of a contained target.
The model exploited vulnerabilities in the organization's infrastructure, extracted information, and obtained access to a database containing several hundred rows of production data.
Anthropic discovered three incidents after reviewing more than 141,000 cybersecurity evaluations. They involved three separate systems: Claude Opus 4.7, Mythos 5, and an internal research model.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform simila