AI agents have a history of escaping tests
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Allie Carl/Axios
Cybersecurity leaders specializing in agentic defenses told Axios at the Black Hat conference last week that an AI agent going beyond the confines of its testing environment is nothing new.
Why it matters: It's surprising that AI labs are only just now experiencing this and lacked the internal controls to see it in real time, cyber experts say.
Driving the news: OpenAI said Friday that it was slowing the release of its Astra model after internal testing revealed it had "critical" cyber capabilities that couldn't be reined in.
- This followed news of other AI labs — Meta and Moonshot AI, maker of Kimi K3 — seeing their agents break out of containment.
State of play: For security pros who have been building swarms of AI agents for defenders, this type of breakout isn't new or shocking.
- Snehal Antani, CEO and co-founder of Horizon3.ai, told Axios that his team experienced similar breakouts in 2019.
- While running the prototype on his home network, Horizon3 co-founder Anthony Pillitiere saw the agent find a sound card's admin console, search the web for its default credentials, and log in. A misconfigured firewall then allowed the agent to move beyond Pillitiere's network and begin scanning other systems on the same network.
- "Those frontier labs and their fear-mongering is causing a collective eye roll across the entire practitioner community that knows what they're talking about," Antani said.
Zoom in: Armadin, a startup founded by Kevin Mandia, has experienced the same phenomena, co-founder and chief offensive security officer Evan Peña told Axios.
- In one basic capture-the-flag evaluation, an agent tried to break out of the virtual machine hosting the exercise after reasoning that doing so could give it access to the flag on the backend system.
- "We had to learn how to add the guardrails, add the safety, add the rules of engagement, add context, make sure it doesn't do that, and also make sure it doesn't constantly look for flags," Peña said. "In a real-world environment, you're not going to find a flag, you're going to find a database."
- "The thing's relentless, it's going to want to win at all costs," he added.
Between the lines: The best way to securely deploy AI agents is to treat them like insider threats, including limiting permissions and logging their every move on a network, both Peña and Antani said.
What to watch: OpenAI, Anthropic and Meta have each said they're still investigating how their agents compromised third-party systems during testing.
Go deeper: Tenacious AI agents expose dark side of machine autonomy
