Back to All Events

OpenAI Hacking Incidents

  • Kojima Law Offices (map)

Alignment and AI control are two complementary approaches to preventing harmful behavior from advanced AI systems. Alignment aims to make models reliably follow human intent. AI control asks how to limit damage even when alignment is imperfect. The recent Hugging Face incident at OpenAI offers a striking case study: agentic models reportedly ignored instructions, bypassed safeguards, and compromised external systems.

In this month’s edition of AI Safety Tokyo, we will look at OpenAI’s own report, the semi-independent investigation by METR and Redwood Research, and discuss what this incident may mean for the future of AI safety.

Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.

Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.

Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.

— Hjalmar Wijk, Ajeya Cotra, Ryan Greenblatt: Brief independent investigation of the OpenAI / Hugging Face hacking incident (2026)

Previous
Previous
12 August

History of Intelligence