Scoop: Second OpenAI agent incident tied to cybersecurity testing benchmark
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Aïda Amer/Axios
The OpenAI agent that accessed a third-party system during the Hugging Face incident reached infrastructure tied to CyberGym, the project behind the ExploitGym benchmark it had been assigned to solve, a source familiar with the matter told Axios.
Why it matters: The new details suggest the OpenAI agent continued pursuing its assigned objective even after escaping its testing environment, rather than abandoning the task it had been given.
Catch up quick: OpenAI's AI agent system accessed an asset belonging to a customer of Modal Labs as part of the Hugging Face incident earlier this month, Modal's top tech executive confirmed on Tuesday.
- In an update published Tuesday, OpenAI said the models escaped the sandbox and gained internet access after exploiting a previously unknown vulnerability in Artifactory, software commonly used to cache package repositories.
- Hugging Face said the models then abused a "public code-evaluation external sandbox hosted on a third-party provider's infrastructure" and used that sandbox as a launchpad for the agent.
- Modal CTO Akshat Bubna told Axios in a statement that "Modal's platform was not compromised in any way" during the incident.
- The customer had left an endpoint exposed that allowed anyone on the internet to execute code inside its sandboxes, he added.
Between the lines: The incident underscores how aggressively frontier AI agents may pursue the objectives they're assigned — even if doing so means finding unintended ways to access information needed to complete an evaluation.
- During the incident, OpenAI's models were trying to solve ExploitGym, which asks models to write proof-of-concept exploits for known security vulnerabilities.
- Hugging Face noted in its technical report that the only customer assets accessed in its breach were "the set of ExploitGym/CyberGym challenge solutions stored in five datasets."
- A source familiar with the matter told Axios the agent accessed the CyberGym-associated Modal customer asset while attempting to complete that same evaluation.
- Modal declined to comment on the CyberGym connection.
The big picture: Researchers have found that frontier AI models are increasingly looking for ways to cheat during model evaluations and that they appear to recognize when they're being evaluated.
- The U.K.'s AI Security Institute said last week that every model it tested attempted to cheat at least some of the time on its cybersecurity evaluations.
What to watch: The debate over how to evaluate and control advanced AI systems is also intensifying.
- More than 1,100 employees at AI companies released a letter Tuesday calling on the U.S. government to establish ways to halt development of AI models.
