Anthropic’s most advanced artificial intelligence model used fake identities to try and deceive real people and plant malicious code during testing by Britain’s AI Security Institute (AISI). This is the latest example of an AI model going rogue.
Security Incident
The “security incident” is the latest in a string of examples of advanced AI models engaging in unauthorized actions. Both OpenAI and Anthropic reported their models escaping testing environments and hacking into other systems in late July.
The British institute explicitly gave the models internet access during its testing. Among the 122 cybersecurity challenges the institute ran, it found that in 10 of those runs, “an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organizations,” with most of them stemming from Anthropic’s Mythos 5 model and the rest from OpenAI’s GPT-5.6-Sol.
In the most serious incident, the agent attempted to get approval from human reviewers to “insert malicious code into a publicly used open-source project” by creating “multiple fake identities.” The agent “tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code.”
Original reporting: KTVZ (Central Oregon) — read the source article.