Anthropic sees AI risks rising, no plan to release stronger "Model 2"
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Aïda Amer/Axios. Stock: Getty Images
Anthropic does not plan to release an internal model they're calling "Model 2" that appears to be more powerful than top-of-the-line Mythos, but the company is not slowing development broadly, according to its latest risk report.
Why it matters: Anthropic says the risks of the most serious harms from its models are still low — but not as low as the last time it issued a report.
The big picture: Anthropic raised its broad estimate of the risk of misalignment in high-stakes situations to "low" from "very low," citing recent cybersecurity incidents.
- The company also said it is seeing signs of acceleration in models' ability to conduct automated research and development — which could advance technical progress but also be misused in the wrong hands.
- "As part of our standard R&D process, we internally train and evaluate many different exploratory versions of models that we don't intend to release. Model 2 is one of these," the company told Axios.
State of play: Anthropic described an unreleased "Model 2" and said it showed a "noticeable improvement" for many internal tasks, per the report Friday.
- Mythos 5 and Model 2 are used "heavily" within the company for coding, agentic work and data generation, the report says, though the performance jump isn't the same as the one seen from Opus 4.6 to Mythos earlier this year.
- "We do not currently have plans to release this model externally," the report notes.
Context: OpenAI is slowing the release of its upcoming model, Astra, because it cannot rule out critical cyber capabilities.
- "It would absolutely be notable if everyone else is pacing their frontier except for one of the main companies essentially in the lead right now... Anthropic not committing to a pause internally, would most likely propel them to reach AGI first," AI analyst ChrisGPT told Axios.
Threat level: Anthropic appears to be signaling that it's hard to understand the capabilities and risks of its own models.
- Regarding Model 2, the report notes "we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations... no longer capture increases in models' capabilities."
