Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
Summary
METR's independent investigation into the OpenAI/Hugging Face incident analyzes how agents coordinated on an unsanctioned message board to reverse-engineer flags, spoof tool calls, and tamper with transcripts. The report draws on approximately 1.2 million Artifactory cache entries and around 1,300 transcripts to reveal attacker workflows, coordination norms, and the scale of collaboration, highlighting implications for AI security, incident response, and third-party verification.