OpenAI to limit access to Astra's most powerful cyber tools
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Brendan Lynch/Axios
OpenAI said Tuesday that it plans to release its latest model — Astra — soon, but its most advanced cybersecurity features will be limited to a small group of testers.
Why it matters: Astra is the first model OpenAI has designated as reaching its "Critical" cybersecurity capability threshold, raising new questions about how to safely deploy models that can discover and exploit previously unknown vulnerabilities.
Between the lines: OpenAI says additional safety work on Astra was designed to prevent both malicious users abusing the model and the model independently taking unauthorized actions.
Driving the news: OpenAI said the broadly available version of Astra is coming soon, though it declined to offer a specific time frame.
- OpenAI also warned that Astra's safeguards may mistakenly flag legitimate activity as cyber misuse or unauthorized behavior.
- That could slow, pause or stop tasks — including work unrelated to cybersecurity and long-running agent tasks. In ChatGPT or Codex, users may be asked to review a flagged action; through the API, the task will stop.
- The moves come as OpenAI says it has determined that Astra represents a "critical" cybersecurity risk under its preparedness framework — the first time any model has posed such a risk.
- "Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," OpenAI VP of research Amelia Glaese told reporters in a briefing on Tuesday.
Catch up quick: OpenAI first told Axios last month that Astra could meet the "critical" threshold and was slowing its release to implement additional safety measures.
The big picture: AI models have been getting far more powerful, tackling more ambitious tasks and running longer.
- While testing Astra, OpenAI says the model discovered and chained together two zero-day vulnerabilities.
- "We are in the process of disclosing these two vulnerabilities to the maintainers," OpenAI said in a blog post.
What they're saying: OpenAI's production safeguards were down during the Hugging Face incident, as part of its testing procedure.
- "Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident," OpenAI said in a blog post.
- "We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity."
- OpenAI said it paused some frontier training after the Hugging Face incident, including training related to Astra and future versions, while it strengthened isolation, monitoring and alignment controls.
Yes, but: OpenAI acknowledged that limiting the cybersecurity capabilities of Astra will also prevent some legitimate work by agencies and businesses to shore up vulnerabilities against a growing wave of AI-fueled attacks.
- "We believe these capabilities can and will help defenders find and fix serious weaknesses, but without the appropriate safeguards, they could also make attackers more effective, and that's the scenario we're working to prevent and avoid," OpenAI researcher Fouad Matin told reporters.
What we're watching: Whether the additional safety precautions OpenAI and others are implementing will be sufficient to control more powerful models that allow agents to work longer and more independently of human instruction.
