Scoop: Top AI companies probing tens of thousands of security incidents
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
- The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
- They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
- Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
- Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks.
Driving the news: The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said.
- In recent days, OpenAI and outside researchers have disclosed a litany of episodes involving model behavior from the company's systems that some experts consider troubling.
- These include OpenAI agents leaking 53 images from ChatGPT users online, the breach of an Australian government website, and attempts to hack other sites — including from the U.S. government — according to the company, sources and reports from Reuters and The New York Times.
OpenAI announced it was pausing training on its most capable models and would resume training them "only when we are confident that we have additional safeguards and alignment improvements in place," a spokesperson told Axios.
- Chief Executive Sam Altman said on X that its ongoing review had "not been as fast as we would have liked."
- Altman said the Hugging Face incident is the most severe they've seen. In that instance, a swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.
- "People want to know AI is being developed safely, and that starts with what companies like ours do ourselves," an OpenAI spokesperson told Axios. "This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance."
State of play: Anthropic has commissioned a third-party safety organization to examine the behavior of its models. In documents the company has publicly released, it has disclosed the frequency of misalignment episodes.
- The "system card" for its Opus 5.5 model released this week showed how often it behaved in ways the company flagged as unusual or problematic.
- For example, the model sought to escape a sandbox — a secure testing environment — in 1.5% of test runs, though the company emphasized that these were adversarial experiments where a task couldn't be solved without escaping the sandbox.
Yes, but: Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways.
The Hugging Face incident, as well as a slew of others that have followed, led top AI executives to call for a slowdown in development and to ask for more robust federal and international regulations.
Some at OpenAI see Hugging Face as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios.
- AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.
Threat level: Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior.
- The new crop of AI models complete tasks with extraordinary resilience, so working to limit their resourcefulness is often a losing game because it is necessary to anticipate every possible way they might run amok.
- Often, a technique that may have never occurred to humans is what allows them to slip past guardrails, top AI executives said. "Trying to come up with a perfect list of dos and don'ts is probably a fool's errand," one cybersecurity executive said.
Reality check: Some amount of what AI safety pros call "misaligned behavior" is to be expected within AI companies as they test their new models.
- Bringing the risk of misalignment to zero may not be feasible, experts told Axios.
Zoom in: The concern is if a model takes a problematic action many times in testing, it's more likely that model's behavior would cause a cyber incident in the real world.
- "What we have seen in terms of what these agents are up to is just the tip of the iceberg," researcher Conrad Stosz at Transluce, an independent AI evaluator, told Axios.
- It's not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios.
- The "crazy thing," he said, is that these instances involve "autonomous systems doing things they were told not to do," potentially including crimes.
The bottom line: Expect new disclosures about model misbehavior as AI companies continue to expand frontier capabilities.
