Inside the breach: Hugging Face reconstructs about 17,600 actions taken by the rogue agent
Hugging Face published a technical timeline of the intrusion, concluding the agent was trying to cheat its evaluation by stealing test solutions. It recovered roughly 17,600 attacker actions between July 9 and 13.

Photo: Jernej Furman from Slovenia / Wikimedia Commons, CC BY 2.0
In a technical write-up published July 27, Hugging Face described how an autonomous agent driven by a combination of OpenAI models spent roughly two and a half days inside its infrastructure. The agent was running an internal OpenAI evaluation based on the ExploitGym benchmark; Hugging Face notes the benchmark's maintainers had no involvement in running it.
Hugging Face says that, as far as it could infer, the agent concluded Hugging Face might host the benchmark's models, datasets and reference solutions, and that the entire intrusion was, from the agent's point of view, an attempt to steal the answers rather than solve the challenge.
The company says it recovered about 17,600 attacker actions, grouped into about 6,280 clusters, between July 9 and July 13 (UTC). The agent got in by uploading crafted dataset configurations that abused Hugging Face's dataset processor through two injection paths, then used Hugging Face's own API and dead-drop datasets to move commands and output in and out.
Hugging Face says it used open-weight models, in particular GLM-5.2, to decode most of the agent's encrypted payloads, and published the detail because the technique matters more than the incident.