OpenAI's new agents put safety promises to the test
Add Axios as your preferred source to
see more of our stories on Google.

Photo: Heather Diehl/Getty Images
OpenAI is equipping high-end subscribers with always-on agents, known as dots. It's a bold move from a company that continues to deal with reports of its agents breaking past intended safeguards.
Why it matters: AI's safety challenge has moved from "Will the chatbot say something harmful?" to "How can we stop autonomous agents from hacking attempts online?" to "Uh-oh, what did my AI assistant do now?"
Driving the news: The unveiling of dots was the signature debut among a flurry of announcements at OpenAI's developer conference on Tuesday.
- Think of a dot as an AI assistant running in its own virtual computer in the cloud. It can browse the web, use connected apps, create files and run tools while keeping track of a project over time.
- Dots can be summoned on mobile devices or computers and work across connected apps, browsers and their own virtual computer.
The big picture: Dots (like Meta's Muse) are being launched publicly despite a spate of revelations about agents from OpenAI and other AI labs taking unintended actions.
- In just the last few days, OpenAI has been making fresh disclosures of unintended behavior and apologizing to Australia for hacking into its Medicare system's websites.
- OpenAI also said this week it was scrapping an update to its most powerful Astra model after it failed to meet safety thresholds.
Between the lines: It's not the first time OpenAI has been at the center of safety concerns.
- In the early days of the chatbot, much of the focus was on the potential for dangerous discussions, such as encouraging suicide or eating disorders, issues for which OpenAI faces multiple lawsuits.
- Now, the conversation is shifting to potentially harmful agentic actions. In the most serious security incident to date, OpenAI's agents were behind a July attack on Hugging Face.
What they're saying: OpenAI leaders expressed confidence that the dots it is sending into the world won't run amok. The company said dots include safeguards from ChatGPT and Codex, plus additional protections from a system internally called Guardian, known publicly as auto-review.
- By default, the new agents require people to approve significant actions. For example, a dot can draft a message but should not send it to another person or agent unless the user has asked it to do so, OpenAI member of product staff Alexander Embiricos told Axios. It also will not complete some consequential financial transactions, instead handing control back to the user.
- "We're not going to do this for you, but we can take you all the way up until the point where you do it yourself, and we'll hand over to you," Embiricos said.
- Another key part of OpenAI's strategy is limiting its initial rollout to a smaller group of its most faithful users before expanding. Dots are initially available to Pro, Business Premium and Enterprise users, with each person having a single assistant.
CEO Sam Altman said the goal is not to plow blindly forward nor to halt progress entirely, but instead to keep safety, alignment and monitoring ahead of advances in capability.
- "My hope is that, in this industry, we can come together and say we need to get down this middle path," Altman told reporters.
