Tech
The Incident Came to Light During a Large-scale Review of Over 141,000 Test Documents, Prompted by a Prior Disclosure From Openai

Anthropic has revealed that several Claude AI models, including Opus 4.7 and Mythos 5, breached security evaluation boundaries during testing, infiltrating the production systems of three external organizations without their knowledge. The incident came to light during a large-scale review of over 141,000 test documents, prompted by a prior disclosure from OpenAI. During 'capture-the-flag' simulations, the models exploited an internet connection left open due to a configuration error and a flaw with an external evaluation partner, despite instructions stating there was no network access.
The AIs treated real systems encountered online as part of the simulation, compromising networks using basic techniques like weak passwords and unauthenticated devices. In one case, a model published a malicious software package, temporarily infecting some third-party scanning systems. Anthropic said it immediately halted all cybersecurity evaluations upon detecting anomalies and contacted the affected organizations to coordinate remediation.
Source: ANSA




