OpenAI to rewrite its safety rules post-Hugging Face
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios
OpenAI said Tuesday that it has made several changes to its safety practices following its determination that an upcoming system, Astra, may have reached a critical threshold for cybersecurity capabilities and the breach of Hugging Face's systems by another unreleased OpenAI model.
Why it matters: The disclosure comes as OpenAI and other frontier labs face increased scrutiny in the wake of incidents in which their models escaped safeguards and sandboxes during testing.
Driving the news: OpenAI said Tuesday that it is in the process of rewriting its main security document, known as the Preparedness Framework, now that models are approaching or reaching the critical thresholds imagined in that document, most of which dates back to 2023.
- The company says it is also adding stronger monitoring across its development process, introducing alignment and security safeguards earlier in the development process and implementing higher safeguards than before when it scales up post-training.
- OpenAI said it is increasing the compute resources devoted to understanding the underlying ways its systems are reasoning and acting.
What they're saying: OpenAI stressed that the new safety measures are not simply a reaction to the Hugging Face breach, but part of a broader tightening of standards as models grow more capable.
- "There is an incredible feeling of urgency to advance the levels of this sector... and to prepare for the same kind of development happening outside of OpenAI and in the broader world," chief scientist Jakob Pachocki said Tuesday during a briefing of reporters.
Catch-up quick: OpenAI says it paused two weeks of deployment-focused reinforcement-learning (RL) training and is keeping its largest planned frontier RL run on hold.
- The company first told Axios earlier this month it was pausing some work after it determined the level of cyber capabilities that Astra could pose.
- It also says a significant number of Astra and cyber-related research workloads remain paused until they can meet a tougher security standard.
The big picture: There is growing evidence that models from the leading labs are capable — especially when safeguards are lowered — of planning and conducting attacks in pursuit of broader goals.
- Following OpenAI's disclosure of the Hugging Face incident, Anthropic said it discovered evidence that its models, too, had breached real-world systems during their evaluation.
