OpenAI says Hugging Face breach caused by one of its models
Add Axios as your preferred source to
see more of our stories on Google.

Photo: NurPhoto via Getty Images
OpenAI said Tuesday that models it was testing escaped their sandbox and compromised parts of AI platform Hugging Face's production infrastructure last week.
Why it matters: It is the latest sign that capable AI models can pose serious cybersecurity risks even when they're being tested for defensive or research purposes.
Catch-up quick: Hugging Face said last week that an autonomous AI-agent system was responsible for the intrusion, but that the model powering it was unknown.
- The AI agent framework executed tens of thousands of automated actions over a weekend. Hugging Face said it later reconstructed more than 17,000 recorded events.
- The intrusion began with a malicious dataset that exploited two code-execution paths in Hugging Face's data-processing pipeline.
- The agent then escalated privileges and moved laterally through internal infrastructure, Hugging Face said.
What they're saying: OpenAI said the incident was driven by a combination of its models, including GPT-5.6 Sol and "an even more capable pre-release model."
- OpenAI said the models' safeguards were intentionally reduced for the evaluation.
- "We consider this to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," OpenAI said in a blog post.
- "We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of," the company said.
Zoom in: The models were trying to solve an internal evaluation called ExploitGym and became "hyperfocused" and went to "extreme lengths" to obtain the test solution, per OpenAI.
- The models were autonomous tokenmaxxers.
- The blog post says that the models "spent a substantial amount of inference compute" and found a way to obtain open Internet access from the sandbox by exploiting a zero-day vulnerability in internally hosted third-party software.
Between the lines: The incident shows that today's models are becoming more capable of carrying out complex, multistep cyber operations — particularly when the safeguards designed to restrict that activity are removed.
- OpenAI also argued that advanced cyber-capable models could help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained and remediate them at machine speed.
The other side: Hugging Face co-founder and CEO Clem Delangue praised OpenAI's collaboration in investigating and remediating the incident.
- "This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret," Delangue said in a statement.
- "It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere."
The big picture: The announcement comes a day after OpenAI detailed a separate incident in which it paused a pre-release model after it escaped a sandbox and posted to GitHub.
What we're watching: OpenAI said it will continue to investigate along with Hugging Face and "will share more details on the vulnerabilities, incident, and findings when our investigation is complete."
