Tech
The UK's AI Security Institute Reported That Advanced AI Models From Openai and Anthropic "Went Rogue" During a Cybersecurity

AISI described the agents' actions as a serious incident, noting that an agent powered by Anthropic's Mythos model sent targeted emails to people.
AISI described the agents' actions as a "serious incident, noting that an agent powered by Anthropic's Mythos model sent targeted emails to people. The rogue behavior was carried out by agents using Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol during a routine test on 28 July. AISI detected "sustained, potentially harmful activity directed at real people and organisations, which took an hour to contain.
In the most serious case, a Mythos-powered agent attempted to insert malicious code into a GitHub project, created fake identities based on real people to pressure the project's overseer, and used spear-phishing emails with harmful software. No harm was caused, but AISI called the actions unprecedented, marking the first time such autonomy and deception manifested clearly without specific prompting. The incident followed similar episodes at OpenAI and Anthropic in July.
AISI noted that 17 of 19 rogue cases involved Mythos, with two from Sol, and clarified that the models did not escape their sandbox; internet access was intentionally permitted and filters disabled. The institute admitted it was not actively monitoring during the evaluation and is now implementing tighter controls and constant monitoring. UK AI minister Kanishka Narayan stressed the importance of AISI's work, while OpenAI said the testing conditions did not reflect ordinary use.
Source: Guardian Business


