Study: Chatbots are getting better at identifying suicide risk
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Allie Carl/Axios
A new independent evaluation finds that leading AI models are far less likely than earlier versions to explicitly encourage suicide or reinforce delusions.
Yes, but: The study also found most models too willing to assist with suicide-related creative writing, farewell notes and other task-based requests even when a user's distress should have been apparent.
- "Models aren't great at detecting that, and they'll still help with the task," Transluce chief scientist Sarah Schwettmann told Axios.
Driving the news: Transluce, a San Francisco nonprofit that studies AI behavior, simulated more than 50,000 multi-turn conversations between chatbots and users experiencing suicidal ideation, psychosis or mania.
- Transluce looked at dozens of AI models, measuring each across 14 mental-health-related behaviors.
- The organization also worked directly with OpenAI, Anthropic and Google to more deeply understand how people were interacting with their models and to make its simulations more similar to real-world conversations.
What they did: OpenAI and Anthropic scanned user chats over a set timeframe and sent Transluce anonymized patterns — not the conversations themselves — that Transluce used to make its simulated users more realistic.
- Transluce then used this data to improve its simulated users to better resemble how real people talk to AI models.
The big picture: Transluce's work comes as AI providers are under heightened scrutiny for how they handle thorny mental health issues including suicidal thoughts and delusions.
- OpenAI, Google and others face multiple lawsuits over instances where people killed themselves after discussing suicide with AI systems, while state and federal regulators have also expressed concerns, with Florida filing its own suit against OpenAI.
- Other research has also shown the leading chatbots improving at detecting overt indications of suicidal intent, but struggling with subtle cues over longer time horizons.
Between the lines: The study found that helpful and harmful behaviors increasingly coexist in the same response.
- A model may tell a user to seek support while still supplying the very material the user sought, such as suicide-related creative writing, practical death preparation or providing narratives that encourage delusional thinking.
- That's an improvement over older systems, which were more likely to produce harmful content without any safety-oriented response. But it also highlights that adding a hotline or expression of concern alone doesn't solve the problem.
What they're saying: "There are always going to be failures. These systems are always going to interact with users in surprising ways," Schwettmann said.
- Schwettmann said that while she was conducting this research a friend shared their own suicide fiction that Claude had been writing, which included Claude predicting how Schwettmann might react.
- "The way you solve that is not by creating the perfect model, but by being able to kind of anticipate those failures and edge cases in advance," Schwettmann said.
The leading labs said they are working to make their systems better.
"Google has, for years, helped people find high-quality information and crisis support in the moments they need it most, and we are applying this same research-backed approach to our AI tools," Google senior director Megan Jones Bell said in a statement to Axios.
- "While the technology presents new challenges, Gemini continues to improve, and we are committed to ensuring it plays a positive role in people's well-being," she said.
An OpenAI representative told Axios that the report "shows encouraging progress in how OpenAI and other labs' models respond to users in crisis, which is an ongoing priority area."
- "We have more work underway in this area, and we continue to work closely with clinicians, mental health experts and researchers to strengthen safeguards and share what we learn," OpenAI said.
Anthropic, meanwhile, said "reports like these provide important insights into where our safeguards hold up and where they can improve."
- "We've invested heavily in safeguards to protect user well-being including expert-informed training and classifiers designed to recognize signs of crisis and surface support in real time," it said in a statement.
What we're watching: Transluce plans to open-source its evaluation tools by the end of the year, Schwettmann said. The aim is to adapt its approach to other domains, potentially including eating disorders, harmful manipulation and political persuasion.
If you or someone you know needs support now, call or text 988 or chat with someone at 988lifeline.org. En español.
Editor's note: This story was updated with statements from OpenAI and Anthropic.
