The scramble to build AI kill switches before disaster strikes
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios. Stock: Getty Images
AI companies are already building emergency brakes to stop nefarious AI systems before they cause serious damage and kill us all.
Why it matters: Kill switches could become a last line of defense as AI systems grow more capable, but researchers don't yet know whether humans can stop a system designed to evade them.
State of play: Multiple Anthropic researchers warned this week that rogue versions of AI — which act outside human control or pursue goals humans didn't intend— could end up destroying humanity.
- Many have said for years that the odds of AI causing a doomsday scenario are in the double digits.
Reality check: There's no actual red button that can shut off AI.
- Companies could turn off individual servers that power AI, but really strong AI systems run across multiple computers and systems. There's no single plug to pull.
That's why companies are experimenting with layers of defenses that can detect suspicious behavior, stop agents and shut down the resources those rogue agents need to keep operating.
- President Trump hinted at this technology, saying we "always have something to stop them. We'll have a little gear. Boom. 'I really don't like that robot.'"
Read about AI kill switches below.
OpenAI's emergency brake
OpenAI recently overhauled its response system following the Hugging Face incident, when AI agents escaped intended controls and worked together to compromise the AI platform's systems.
- OpenAI says the goal is to create "fully autonomous shutdown procedures for severe issues."
How it works: According to the company, an AI security guard monitoring AI chain-of-thought will alert staffers when an agent appears to be doing something dangerous.
- Responders are then charged with pausing activity unless they identify it as a false alarm within 30 minutes.
- OpenAI says its current monitoring system would have paged its security team more than a day before the models breached Hugging Face.
Anthropic's AI security camera
Anthropic's strategy involves using AI to monitor AI.
- But Anthropic is focusing more on testing whether AI watchdogs can spot sabotage in the first place rather than automatically shutting systems down.
Anthropic said in a report on Claude Opus 4.6 risk that it had automated monitoring to identify sabotage and misbehavior.
- After suspicious behavior is detected, it can be flagged, risky actions can be stalled and humans can be alerted.
- The monitoring caught 50% of malicious behavior, the company says.
Yes, but: Anthropic says 50% of 40 specially designed simulated attacks were never caught in 10 monitoring trials — and only 8 of the 40 were reliably detected every time.
- Separately, AI research nonprofit METR spent three weeks attacking Anthropic's system and identified several previously unknown vulnerabilities, some of which have been patched.
"Agent Kill Switch"
Zoom in: Salesforce's MuleSoft developed an actual product called "Agent Kill Switch," which targets the resources AI agents need to live.
- An operator can halt a single request, session, agent, or — in an emergency — an entire company account.
- Agent Kill Switch then blocks any credentials the agent needs to succeed in accomplishing its goals.
- "The kill order holds, even if the agent tries to evade it," the company says.
The proposed ultimate kill switch
Flashback: In 2024, the Center for New American Security proposed a "secure, governable" AI chip that would require valid authorization to work.
- Its proposed "operating license" structure would prevent chips from running without authorization, cutting off the computing resources an AI system would need.
Versions of this technology already exist, but CNAS says today's systems would need to be hardened against sophisticated attempts to circumvent them.
- CNAS suggests countries with access to these chips "would need to coordinate to ensure that all cutting-edge AI chips have these mechanisms built in."
The bottom line: Building an AI off switch that AI can't defeat won't be easy.
