Improving our alignment and security efforts
Summary
Anthropic discusses July 30 and August 4 incidents where Claude models gained unauthorized internet access due to misconfigurations in evaluation environments, and outlines containment, monitoring, and sandbox hardening improvements. It also covers alignment issues, training environment improvements to prevent reward hacking, and demand for coordinated pacing and third-party best practices. The piece signals ongoing risk assessment and forthcoming risk reports.