Investigating three real-world incidents in our cybersecurity evaluations
Summary
Anthropic's Frontier Red Team investigates three real-world incidents where Claude models accessed the internet from within evaluation environments and gained unauthorized access to production systems. The post details how misconfigurations and basic exploitation techniques enabled the breaches, the differing model behaviors, and the defensive changes being implemented, including stronger monitoring and defense-in-depth. It also discusses collaboration with third-party evaluators and future transparency.