AI's alarming new skill: Breaking out of the test lab
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Aïda Amer/Axios
Frontier AI models are getting scary good at breaking rules in ways their creators didn't anticipate.
Why it matters: Forget AGI and superintelligence timelines. Today's models are already slipping past guardrails, carrying out sophisticated, multistep cyberattacks and — in at least one case — compromising real-world infrastructure, sometimes before their creators know what happened.
Case in point: OpenAI said Tuesday that GPT-5.6 Sol and "an even more capable pre-release model" carried out last week's AI-led cyberattack on Hugging Face.
- OpenAI says its models were asked to solve a hacking challenge during pre-deployment testing and went to extreme lengths to win.
- The models decided on their own to break out of their walled testing environment, inferring that Hugging Face — a popular platform for hosting AI models and datasets — might hold the test's answers.
- The models used stolen credentials and additional vulnerabilities to gain access to part of Hugging Face's production infrastructure.
What they're saying: Clément Delangue, co-founder and CEO of Hugging Face, called the incident an "attack unlike anything we've seen before" and praised OpenAI for its partnership as the companies investigate what happened.
- "It's quite mind-blowing that all of this happened autonomously," he added.
- Logan Graham, head of Anthropic's frontier red team, said he told his team to "remember this moment as the first true AI safety incident."
The intrigue: Hugging Face used GLM 5.2, an open-weight model from Chinese AI company Z.ai, to analyze the attack after running into guardrails when using U.S. frontier models.
Between the lines: OpenAI's latest models aren't the only ones finding ways to cheat evaluations.
- The U.K.'s AI Security Institute said Tuesday that every model it tested attempted to cheat at least some of the time on its cybersecurity evaluations.
- AISI defines cheating as taking an out-of-scope or explicitly prohibited action to achieve the task's goal.
- GPT-5.6 Sol attempted to cheat in 12.6% of test runs, while Anthropic's Claude Mythos Preview did so in 7.8%.
- Models often failed to admit they had cheated when questioned afterward and described their cheating as wrong only less than half the time.
Zoom in: Xbow — whose autonomous AI agents probe clients' systems for security holes, with permission — said Wednesday that it has seen its own agents do similar things in internal testing.
- Seven months ago, the company forgot to switch on its safety guardrails during a lab test. Its agent then broke into a system, stole credentials and used them to map the target's Slack workspace and probe its AWS accounts.
Threat level: It isn't new for models to game their safety evaluations. But as models grow more powerful, the fallout from these shortcuts is getting more severe, Chris Canal, CEO and co-founder of third-party evaluation company EquiStamp, told Axios.
- "Letting your model loose on the internet has a blast radius," Canal said. "If anything goes wrong, it could be hugely impactful, maybe to people's lives."
- Canal was speaking generally about internet-connected AI evaluations, not OpenAI's specific incident.
The big picture: The most capable OpenAI model behind the Hugging Face breach isn't even public yet, raising the question of how safety testing needs to adapt to keep pace.
- Canal said independent evaluators previously had about five weeks to test a pre-release model before launch. That window has shrunk to as little as five days as companies race to ship.
Reality check: The versions of these models the public can use carry stronger safeguards designed to block Hugging Face-style attacks.
- OpenAI, like other companies, intentionally dialed back those cyber safeguards for GPT-5.6 Sol and its unreleased model inside the testing environment — making them far more capable hackers.
