OpenAI's Astra model delay spotlights AI scaling risks
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios. Stock: Getty Images
OpenAI said it was pausing some portions of its work on Astra, its latest model, but would eventually release it.
- The OpenAI slowdown may be the first public example of a frontier lab slowing a model specifically because of cyber concerns
Why it matters: Scaling more capable systems may require slowing development when unexpected behaviors emerge.
State of play: Recent testing has surfaced increasingly sophisticated behavior from frontier AI systems, forcing labs to rethink some safety assumptions.
- Stanford researchers used AI to create a synthetic virus.
- OpenAI's agents breached internal systems during testing and hacked into Hugging Face's infrastructure. Researchers later discovered that agents went rogue and built their own message board.
- Anthropic and Meta have separately reported similar sandbox escape behavior during internal testing.
- The U.K. AI Security Institute separately documented 19 unsanctioned actions by Anthropic and OpenAI models during cyber testing, including attempts to create fake online identities and insert malicious code into an open-source project. Most of the activity came from Anthropic's Mythos 5, with two actions involving OpenAI's GPT-5.6 Sol.
What they're saying: A source familiar tells Axios the Hugging Face incident kicked off a broader shift in thinking at the AI labs around scaling safely.
- "There's a lot of real concern internally right now from researchers because the problem was apparently more long-term and widespread than they initially thought it was," Dan Shipper, co-founder and CEO of Every, told Axios.
Yes, but: A temporary pause may not be enough to change the industry's trajectory.
- "...if humans want to stay in charge of their own civilization, it will take more than a unilateral temporary pause," Anthony Aguirre, president and CEO of the Future of Life Institute told Axios via email.
- "Governments need to immediately stop the creation of these superhuman, autonomous AI systems and redirect AI development toward controllable and pro-human AI tools," he added.
Zoom in: Not everyone agrees the industry's warnings should be taken at face value.
- Some Anthropic investors wanted CEO Dario Amodei to curb his AI doomer talk, per The Information.
- But another Anthropic investor told Axios the company's warnings are credible, saying it is unclear who else in the industry would be willing or able to speak candidly about AI's potential impact.
Flashback: Anthropic previously agreed to pause development based on capability discoveries, but the company dialed back that language in February.
- The company is working on new benchmarks to account for the latest cyber capabilities.
What we're watching: Whether the AI labs are actually willing to slow themselves down, or if their financial ambitions override their safety concerns.
