Researchers have built a tool that highlights possible signs of urgent risk in text conversations. Its promise lies in helping counselors notice clues while keeping people in charge.
The message behind the message
A crisis conversation can turn on a handful of words. A person reaches out. A counselor replies. As the exchange unfolds, the counselor must listen with care and work out whether the person needs immediate help. That is a lot to ask of anyone, especially when every message arrives one fragment at a time.
Researchers at MIT’s McGovern Institute for Brain Research think a new language tool could eventually help. In a study announced on September 24, 2026, they describe a system that looks for signs of suicide risk in text conversations and explains which signs influenced its assessment. The aim is practical: make important clues easier to see when they matter most.
There is a crucial boundary. This research does not establish that a computer can predict who will attempt suicide. It tested how a model classified conversations against existing risk assessments. The team says the system needs further validation before anyone uses it in clinical practice. Still, the work raises a compelling possibility. What if AI helped create a clearer map of language, and a human used that map to ask better questions?
Why a text conversation is hard to read
Text strips away a lot. A counselor cannot see a person’s expression or hear a change in their voice. Messages may arrive slowly. A person might say exactly what they mean, circle around it, or use language that only makes sense after several exchanges. The counselor has to build trust while deciding what to explore next.
That challenge sits at the heart of the new study. The researchers worked with Crisis Text Line, whose trained volunteer counselors provide text-based support. The nonprofit supplied controlled access to de-identified records of roughly 16,000 conversations. Its assessments grouped those conversations into three categories: non-suicidal, suicidal thoughts without imminent risk, and imminent risk.
The last category was the team’s main focus. In this dataset, it included people with a suicide plan or an intention to die within the next 48 hours. That definition describes how the researchers labeled the records; it is not a clock that starts whenever someone uses a particular word. Nor does a phrase, on its own, tell a counselor the whole story. The task is to understand a person, not merely to spot vocabulary.
Start with 49 questions, then build a word list
The researchers began with 49 established risk factors for suicidal thoughts and behaviors. These cover different experiences, from psychological symptoms to the kinds of thoughts that may signal more immediate danger. A familiar problem appeared almost at once: people rarely describe an experience in the tidy language of a clinical questionnaire.
So the team asked a large language model to suggest words and phrases associated with each factor. It was an efficient way to produce a first draft. Then people did the slower, indispensable work. Researchers reviewed and curated the suggestions, and expert clinicians checked whether the terms actually belonged with the concepts they were meant to represent.
The finished Suicide Risk Lexicon contained about 60 words or phrases per factor. Think of it as a carefully edited index. The AI helped assemble candidate entries; experts checked the entries; the final list gave the subsequent model something specific to look for. According to the research team’s paper summary, the initial generation used GPT-4 Turbo. The published workflow does not simply ask a chatbot to read a crisis message and make an opaque judgment.
The smaller model gets its turn
Once the word list was ready, the researchers trained a relatively lightweight machine-learning model. It searched the conversations for language linked to those 49 factors and learned how strongly each factor related to the risk categories in the dataset. The model then assessed conversations it had not seen during training.
Here is the useful twist: the system can show which factors contributed to an estimate. If language tied to active suicidal thoughts or access to lethal means pushes a conversation toward the imminent-risk category, a user can inspect that contribution. The result offers a route to a follow-up question, rather than a mysterious number everyone must accept on faith.
The researchers report that their clinically reviewed lexicon outperformed a more general-purpose language dictionary, LIWC, on the study’s task. It also performed similarly to some more complex deep-learning models. “Some” matters. This is not a claim that the simpler model beats every other approach. The attraction is that it combines useful classification with explanations and modest computing needs. MIT says it can run on a personal computer, which may also ease some cost and privacy concerns.
Which clues carried more weight?

One result stands out. In these crisis conversations, language associated with active suicidal thoughts, direct self-injury, lethal means, and substance use was more strongly associated with the imminent-risk group than language about depressed mood or fatigue. Anxiety, post-traumatic stress disorder, and emotional pain fell into an intermediate range in the team’s analysis.
That does not make depression or emotional pain unimportant. They deserve care and attention. It means the model found different signals more useful for separating the specific risk categories used in this dataset. The distinction is easy to lose when a scientific result gets compressed into a headline.
The broader picture also matters. The US Centers for Disease Control and Prevention notes that suicide risk rarely comes down to one circumstance. Health, relationships, community conditions, and access to support can all shape it. A word list will never contain someone’s entire life. What it can do, if properly validated, is draw attention to a piece of a conversation that might call for closer human attention.
A result is not a prediction of the future
“Predicting risk” sounds dramatic. In this study, it has a narrower meaning. The model learned to distinguish among labels that Crisis Text Line had assigned to past conversations. Researchers then checked whether it could classify previously unseen conversations using those labels. That is a meaningful test of the model’s ability to recognize patterns in text.
It is a different test from following people over time and measuring who later attempts suicide. The MIT announcement does not report evidence of that latter outcome. It also does not say the software has been approved as a clinical diagnostic tool or installed as a replacement for counselors.
This distinction makes the result easier to appreciate. A system can become helpful without claiming supernatural foresight. In a high-stakes conversation, a prompt that says, in effect, “Please look more closely at this factor,” might be useful. But proving that it improves real-world decisions would require more research. Researchers would need to test how the tool behaves in new settings, how often it misses urgent cases, and whether its alerts help the people receiving them.
The snag: words need context
Language loves to complicate neat categories. A person might describe something that happened years ago. They might quote someone else, ask about a friend’s safety, or use slang a dictionary has never seen. A word-matching system can register a phrase while missing its meaning. It can also miss a concern that someone expresses in unfamiliar terms.
The MIT researchers flag these limits directly. A lexicon does not understand context the way a human reader can, and it cannot match phrases absent from its list. That creates room for both false alarms and missed signals. Imagine an alert arriving at the wrong moment, or no alert appearing when a counselor should be especially concerned. Those are reasons to study the system carefully, not reasons to treat the text as a simple scorecard.
There is also the matter of change. People invent new expressions. Different communities use different words. Language in medical records differs from language in an online chat. A model that works on one collection of crisis conversations might need revision and fresh testing elsewhere. Its designers say continued validation and updates would be essential before clinical use.
The people remain the point
The researchers put human oversight at the center of their account. Satrajit Ghosh of MIT’s McGovern Institute says a human will need to remain involved for a long time. That view fits the service where the data originated. Crisis Text Line says a trained counselor answers messages, listens, and works with the person to make a plan for staying safe; supervisors oversee its volunteers.
An alert cannot build rapport. It cannot hear the hesitation behind a sentence or understand every detail of someone’s circumstances. And it should not pretend to know what a person will do next. A counselor can ask, listen to the answer, and adjust. The best potential role for this research, as described by its authors, is to help such a person notice relevant information sooner.
There is a second human layer, too: clinicians helped check the risk-factor language before it went into the model. That work matters because AI-generated suggestions can sound plausible without being useful. Here, human review shaped the tool before any predictions were made. The system’s design is a collaboration between machine-assisted drafting, clinical judgment, and later statistical testing.
Privacy is part of the design question
These messages are intensely personal. That makes the choice of technology more than a contest for the highest score. For the study, Crisis Text Line gave the researchers controlled access to de-identified conversations. MIT says a small model can run locally on a personal computer and that the researchers sometimes use their lexicon alongside larger AI models to guarantee detection of selected terms while maintaining data privacy. (Massachusetts Institute of Technology)
Local operation may reduce the need to send sensitive text to an outside AI service. It does not automatically solve every privacy concern. Any organization considering such a system would still need to decide who can see the messages, how data are protected, and how long they are kept. Those are practical questions raised by the potential use of the tool, not findings that the study has settled.
The team’s approach offers a sensible starting point: do as much useful work as possible with a smaller, inspectable method. Then check whether any additional complexity earns its place. In mental health research, “more powerful AI” is only one part of the equation. Respect for the person behind the message belongs there too.
A tool other researchers can inspect
The project does not end with a paper. The team has shared the lexicon and its construct-tracker software, so other researchers can inspect the approach and build comparable lists for other topics. The project’s documentation describes the suicide-risk lexicon as covering 49 risk factors that clinicians reviewed.
That openness invites useful scrutiny. Other teams can examine the entries, ask whether a category makes sense, and explore where the method falls short. They can also test whether a lexicon built for one research question works for another. Sharing code does not make the restricted crisis conversations public, and it does not establish that a tool is safe to deploy. It does make the method easier to check and improve.
The story has a quieter timeline than a flashy launch announcement suggests. The team’s publicly described lexicon and preprint appeared before this week’s news. The development on September 24, 2026, is the MIT report of the peer-reviewed journal publication and its explanation of the results. The underlying idea has been built and tested over time; the newly reported paper adds a formal research milestone.
The hopeful part is surprisingly human

This story began with AI, but its most promising outcome may be a better question from a person. The new system organizes language associated with known risk factors, shows which factors affect its estimates, and performed well on a retrospective set of crisis conversations. It gives researchers a way to study signals that appear while someone is actually seeking help.
It has clear limits. A matched phrase is not a diagnosis. A risk category is not a forecast of anyone’s future. And a model cannot replace the trust that a counselor builds one message at a time. The authors themselves call for human involvement and much more validation.
That is a hopeful place to land. Good tools help people pay attention. They expose their reasoning, invite correction, and make room for the next compassionate question. If further testing shows this one can do that safely, its biggest achievement may be helping the person on the other side of the screen feel heard a little sooner.
Sources
- MIT News, “Estimating suicide risk from text” (September 24, 2026).
- Low and colleagues, research summary and abstract, Senseable Intelligence Group.
- Suicide Risk Lexicon and construct-tracker software, research team repository.
- Crisis Text Line, frequently asked questions.
- US Centers for Disease Control and Prevention, “Risk and Protective Factors for Suicide” (May 26, 2026).
The Kingy Brief
Get the next Kingy Brief.
Source-checked AI changes, original tests and one practical thing to try.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
