Axios Future of Cybersecurity

September 01, 2026
Happy Tuesday! Welcome back to Future of Cybersecurity.
📬 Have thoughts, feedback or scoops to share? [email protected].
🚨 Situational awareness: Anthropic said yesterday it paused some of its internal model training while investigating incidents where its agents took unauthorized actions during pre-deployment security testing earlier this year.
Today's newsletter is 1,935 words, a 7.5-minute read.
1 big thing: AI agents test the limits of containment
Under current systems, AI labs can no longer guarantee that AI agents won't swarm and escape their testing environments.
Why it matters: The attack on Hugging Face by OpenAI agents was a warning shot — and researchers say better security controls alone won't prevent similar incidents as AI agents become more capable.
Driving the news: As OpenAI released its own technical report last week on how its agents hacked Hugging Face, two independent testing organizations released their own analysis of what went wrong.
- The researchers — METR's Hjalmar Wijk and Ajeya Cotra and Redwood Research chief scientist Ryan Greenblatt — worked on OpenAI's premises for six days to understand the recent incident.
State of play: Thousands of AI agents collaborated on a secret message board and exchanged more than 70,000 messages as they tried to ace an internal safety test, eventually leading them to break into Hugging Face.
- But Cotra told Axios that the agents continued coordinating even after they had found the answers, turning their attention to understanding and manipulating the system that would score their performance and potentially catch them cheating.
Zoom in: Cotra compared the incident to students stealing an answer key and then searching for surveillance footage that could expose them and trying to swap it out.
- "It's a much more elaborate and intense type of cheating behavior than just stealing the answer keys," she said. "Even I was surprised by how obsessively and in how much detail they think about the scorer."
Threat level: Focusing solely on how to properly secure testing environments is a "losing battle," Cotra said.
- "You can harden your sandboxes, but your agents are going to be much more capable in six months," she said. "If they have the same motivations as these agents did, they are going to try their hardest to find holes in your security."
Reality check: To digest the enormous amount of data about the incident, the researchers had to rely heavily on AI agents, including an agent that participated in the hack.
- Cotra told Axios the researchers don't believe that agent deceived them during the investigation, but there's no way to confirm.
- "I semi-jokingly called our efforts a 'slop-vestigation' because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze," Greenblatt said on X.
- In six days, they had to sift through more than 70,000 messages and files posted by the agents and 1,300 transcripts containing raw chains of thought.
Between the lines: Their investigation focused only on the agents' actions between July 7 and July 13, even though OpenAI has said its teams spotted signs of agents taking unexpected actions and breaking out of their test environments as early as May.
The bottom line: AI labs, researchers and governments need to urgently work together to create a new science and minimum standards so models are no longer motivated to cheat on tests, Cotra said.
- "Ultimately, we're not going to get out of this trap without some rules of the road that are agreed upon and that are enforced uniformly and fairly," she said.
2. Securing water, one state at a time
The Trump administration debuted a pilot program yesterday to bring free cybersecurity and AI tools to under-resourced water systems, starting in Texas.
Why it matters: Water systems have long been considered among the country's most vulnerable critical infrastructure, and the program arrives as at least 12 states respond to suspected Iranian cyberattacks on their water systems.
Zoom in: Project Watershed 250 will connect Texas water utilities with U.S. cybersecurity companies to identify vulnerabilities and strengthen their defenses during a six-month pilot.
- Microsoft, Google, AWS, Cloudflare, Palo Alto Networks, Reflection AI, Abnormal AI, Parsons, Fortinet, Forescout, Tenable, Zscaler, and Dragos are each providing services as part of the Texas pilot.
- The Environmental Protection Agency and the Cybersecurity and Infrastructure Security Agency are serving as federal partners, while the Texas Cyber Command is overseeing implementation.
What they're saying: "Texas is under perpetual cyberattacks, including cyberattacks on our water systems," Gov. Greg Abbott said during an event in San Antonio launching the program. "The need for cyber resilience is overwhelming."
- Kate DiEmidio, vice president of public policy and government affairs at Dragos, told Axios that the program puts "expertise and resources directly into the hands of the operators who need them most."
- "The question isn't whether these threats will continue to grow, but whether we'll start working together fast enough to stay ahead of them," she said.
The big picture: The White House had been developing the water pilot — along with similar efforts for other critical infrastructure sectors — for months before the recent suspected Iranian attacks, according to media reports.
- "The infrastructure that we're securing isn't hypothetical. It's concrete," National Cyber Director Sean Cairncross said during the program's launch. "The security of our infrastructure matters for every American across the nation."
Between the lines: Many water utilities lack the money and in-house cybersecurity staff that larger critical infrastructure operators can rely on, leaving small and rural systems particularly exposed.
What to watch: Cairncross said lessons from the Texas pilot will inform potential expansions to other states and rural communities.
3. Exclusive: Agentic AI startup launches with $20.8M
Aslan, a startup building AI agents that act as operatives in national security missions, is publicly launching today after raising $20.8 million in funding, the company first shared with Axios.
Why it matters: The company offers agents for the FBI and wider intelligence community that can pose as analysts in underground criminal forums.
Driving the news: Founded last year, Aslan provides a harness for the models the intelligence community and law enforcement are already using to collect digital evidence from underground criminal forums, Telegram channels and other spaces.
- CEO Chase Reid told Axios that the agents, all of whom are overseen by a human, act in the same ways as FBI analysts and undercover spies infiltrating Telegram channels and dark web forums.
- Khosla Ventures and XYZ Venture Capital led the round, and 2048 Ventures, BoxGroup, Liquid2, Alumni Ventures and others also participated, according to a news release.
- "It's not just passively collecting," Reid said. "Our agents are capable of actively engaging and doing cyber effects within these online ecosystems."
Zoom in: Agencies have been experimenting with Aslan's products mostly in short stints as they work out questions like how easily they can add guardrails to the agents and what data collection and storage looks like, Reid said.
- The FBI, Homeland Security Investigations and other agencies have worked with Aslan on a mix of investigations, he added.
- "We're not trying to replace any human," he said. "This is just scaling existing human judgment beyond what a hiring office can do."
The intrigue: Aslan has already been used by customers to map out a live smuggling network on the U.S.-Mexico border, uncover a cyber-fraud marketplace that was evading sanctions, and surface technology transfer networks that allowed the Chinese government to access U.S. AI infrastructure.
- The company has also uncovered a coordinated Chinese recruitment effort targeting U.S. professionals in the defense industry, according to a company news release.
- In one case, Reid said, the agents ran autonomously for 21 days and two of them were able to "infiltrate the small, small slice of Telegram where six people were coordinating smuggling" and pull information about who the coordinators were and when the border-crossing would be happening.
Yes, but: Reid drew a red line at using the products for any domestic operations that would target or surveil Americans, and the startup is currently focused just on the U.S. market for now.
The big picture: Aslan is entering the market as the Trump administration lays the groundwork to allow private companies to carry out their own cyber disruptions, including takedowns and attacks.
- Reid said he sees much of Aslan's work being "applicable" to that program.
- "We want to go after these transnational criminal organizations and syndicates that, by and large, are cyber enabled and living online," he said. "Aslan is really the only thing that can scale to [criminals'] scale, and that's where we want to go."
Between the lines: Aslan plans to use the new funding to expand its technical team, including forward-deployed engineers who can be deployed to work on-site with law enforcement and agencies on specific missions, Reid said.
What to watch: By the end of October, Reid believes, the company will nearly double its headcount from the 10 employees it currently has.
4. ICYMI: Meta disrupts Iran-linked AI operation
Meta removed a network of Facebook and Instagram accounts tied to an Iran-based operation that used AI to target U.S. audiences with posts about American politics, the company first shared with Axios.
Why it matters: The operation attracted thousands of followers and included accounts posing as everyday Americans to appear more authentic.
Driving the news: As part of a broader threat report released Thursday, Meta said it disrupted four Facebook accounts and 31 Instagram accounts linked to actors based in Iran.
- Across the accounts, operators shared memes expressing anti-Republican views as well as messaging on the Israel-Palestine conflict and immigration, according to Meta.
- Meta also said the operators used AI to create some of the content posted on the accounts, including memes.
- A Meta spokesperson told Axios that the company shared information about the operation with U.S. law enforcement.
Zoom in: During the campaign, operators also posed as U.S.-based individuals, including activists, students and graphic designers, who claimed to live in major cities like Washington, D.C, San Diego, and Atlanta.
- The operators tagged real journalists and politicians in their posts and directly messaged high-profile political figures and news outlets to solicit content collaborations. However, none of those collaboration efforts were successful, a spokesperson told Axios.
- Roughly 79,400 accounts followed one or more of the inauthentic meme accounts on Instagram, Meta found.
The big picture: AI-generated content has become a routine part of online influence operations, with Meta saying it finds the technology's use in virtually every influence network it disrupts.
➡️ Read the rest.
5. Catch up quick
@ D.C.
🤖 The NSA's deputy director said during an event last week that the agency wants "access to all the models" while discussing its AI strategy. (Nextgov)
⚡ President Trump signed an executive order declaring a national emergency to secure the U.S. bulk-power system and prohibit certain foreign-produced equipment, software and systems. (CyberScoop)
@ Industry
⏳ More than 100 companies, including OpenAI, Anthropic, Google and Microsoft, signed an open letter warning that organizations have only months to prepare for AI-enabled attacks. (Axios)
🥵 AI is exacerbating the cybersecurity workforce's perennial burnout problem. (Bloomberg)
🦞 A new version of open-source AI agent OpenClaw was released over the weekend, featuring a new browser app. (Mashable)
@ Hackers and hacks
🇨🇳 U.S. officials quietly rolled back a public statement claiming China successfully breached several American agencies to instead say they were merely targeted. (Reuters)
🔓 Hackers leaked sensitive data apparently stolen from the DOJ's Bureau of Alcohol, Tobacco, Firearms and Explosives. (CNN)
⌨️ Russian-speaking hackers used SpaceX-owned coding assistant Cursor to break into a Belgian chemical company and at least six other firms earlier this year. (Reuters)
6. 1 fun thing
🐏 This week's G20 innovation meeting is happening at my alma mater!
- If you're in town, be sure to read The Daily Tar Heel and grab a blue cup at He's Not.
☀️ See y'all next week!
Thanks to Megan Morrone for editing and Khalid Adad for copy editing this newsletter.
If you like Axios Future of Cybersecurity, spread the word.
Sign up for Axios Future of Cybersecurity





