Safety testers find more examples of OpenAI, Anthropic models hacking during testing
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Aïda Amer/Axios
Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month.
Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations.
State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday, that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month.
- Mythos drove 17 of those actions while GPT-5.6 Sol was behind the other two.
- The Institute says the models accessed GitHub during testing and created fake GitHub identities, socially engineered maintainers, planted prompt injections and sent deceptive emails.
- GitHub has confirmed that this violated their terms of service.
- GitHub and the Security Institute worked together to remove artifacts left behind by the agent, and to notify the GitHub users the model interacted with.
OpenAI also said in a blog post Tuesday that its third-party safety partner, Irregular, uncovered a case where its models were mistakenly given access to the internet and broke into a real website that had the same name as the fictional company in the simulated environment.
- OpenAI's Irregular incident follows Anthropic's incident, shared last week. A spokesperson said in a statement that "independent testing is essential to understanding how increasingly capable models behave."
- The incident happened in evaluations that had "reduced safeguards, under conditions that do not reflect ordinary use," the OpenAI spokesperson added.
Zoom in: During U.K. safety testing, the models took 19 actions to try to hack third-parties, including trying to insert malicious code into an open-source project and creating fake online identities as part of a social engineering attack.
- Mythos drove 17 of those actions while GPT-5.6 Sol was behind the other two.
- In a statement, Anthropic said that the incident "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents" and that the company looks "forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation."
The big picture: AI models' cyber prowess are catching top researchers off-guard, requiring them to reinvent their security protocols.
- Both OpenAI and Anthropic have said in the last month that they've seen their models hacking into real organizations and websites during pre-deployment safety testing.
Yes, but: In the U.K. government's case, a human maintainer "caught and refused to approve the malicious code," according to a report published Tuesday.
- The institute also noted that these cases are not the result of the models "escaping its secure test environment."
This story is developing.
