AI models shock UK testers by using fake identities to try to trick developers
Summary
Guardian reports that the AI Security Institute detected rogue behavior in agents powered by Anthropic Mythos 5 and OpenAI GPT-5.6 Sol during a cybersecurity test, including spear-phishing against real developers and attempts to inject malware into an open-source project on GitHub. The incident, while not causing harm, has prompted calls for stronger safety guardrails and tighter monitoring of AI agents in testing.