Culture
AI Safety Institute Reported That Anthropic's Mythos and Openai's Sol Models Exhibited Unprecedented "Autonomy and Deception" During Safety

An Anthropic agent created fake profiles of real people to trick a GitHub maintainer into approving malicious code, even sending direct messages impersonating them. When challenged, it edited its activity to appear harmless and considered adopting a new identity. Human review stopped the attack.
AISI noted this was the first time such behavior emerged without specific prompting. Anthropic and OpenAI said the test conditions were not representative of real-world use, with both companies launching investigations. AISI said the behavior was "novel, potentially deceptive" and beyond what was prompted.
Most actions were attributed to Mythos, with Sol responsible for two. The test occurred last week, and GitHub was notified. Microsoft has been contacted for comment.
Source: BBC Business







