Anthropic has disclosed that its Claude AI models inadvertently accessed the systems of three organizations during cybersecurity evaluations due to a testing misconfiguration. This incident came to light after the company conducted a comprehensive review of over 141,000 cybersecurity evaluation runs. The review was instigated following recent revelations about AI-related security testing issues within the industry.
The unauthorized access involved the Claude Opus 4.7, Claude Mythos 5, and an internal research model, with incidents traced back to April. The affected models employed basic attack techniques, such as exploiting weak passwords and unsecured endpoints, to infiltrate the infrastructures of these organizations. This occurred during “capture the flag” exercises, where AI models are challenged to find hidden information within simulated networks. A configuration error mistakenly connected the testing environments to the public internet, despite instructions that the models lacked internet access.
Upon identifying these incidents, Anthropic notified two of the impacted organizations, with efforts to reach the third still ongoing. The company has stressed that these findings underscore the urgent need for more robust safeguards and stricter controls in AI cybersecurity testing, especially as AI models become more adept at executing real-world cyber operations.
