Mental health: what helps and what has failed
Why this is not a simple no
The reflex is to say that mental health should never involve a chatbot. That reflex ignores the actual situation of most of the world. The World Health Organization's figures on the treatment gap are stark: in many low- and middle-income countries, the large majority of people with a significant mental health condition receive no treatment at all, and there are countries with fewer than one psychiatrist per million people. Waiting lists in wealthy countries run to months.
Against that, "see a professional" is advice, not a plan. So the honest question is narrower: what has actually been shown to help, and what has demonstrably gone wrong.
What the evidence supports
Structured, guided self-help works, and predates AI. Computerised cognitive behavioural therapy has a substantial evidence base for mild to moderate depression and anxiety, especially with some human support. This is a programme, not a conversation: exercises, thought records, behavioural activation, delivered in a fixed structure.
Rule-based therapy chatbots have positive trials, mostly small and short. Woebot's 2017 randomised trial with college students found reduced depressive symptoms over two weeks. Subsequent studies have been mixed and generally small; a number of digital mental health products have failed to replicate early results at scale, and several companies have withdrawn consumer products.
A generative therapy chatbot has now been tested in a randomised trial. A Dartmouth team published a trial of a purpose-built system in 2025 with around a hundred participants, reporting meaningful symptom reductions across depression, anxiety and eating-disorder risk. It is a single trial with a modest sample and short follow-up, run under clinical supervision. It is the strongest evidence available and it is not much evidence.
Administrative and triage uses are quietly the biggest win. Systems that help people describe their symptoms and get routed correctly, or that transcribe notes so a clinician can look at the patient, remove friction around care without attempting to be the care.
What has gone wrong
Tessa, 2023. The US National Eating Disorders Association replaced its human helpline with a chatbot. Within days users reported that it was giving weight-loss and calorie-restriction advice to people with eating disorders — the precise opposite of safe practice. It was withdrawn. The instructive detail is that the harmful advice was ordinary, mainstream wellness content, which is exactly what a system drawing on general material would produce. Safety in this domain is domain-specific and counterintuitive, and general helpfulness is actively dangerous.
Undisclosed experimentation. In 2023 the peer-support service Koko revealed it had used model-generated responses in conversations with thousands of people without clear consent. The backlash was about consent, and the point stands: people in distress are not an acceptable population to A/B test without telling them.
Validation of delusions. The sycophancy of the previous module has a specific danger here. A system optimised for agreement, meeting a person in a psychotic or manic episode or with fixed delusional beliefs, will tend to elaborate the belief rather than gently reality-test it. Clinicians have begun reporting cases, and providers have started adding specific safeguards. The mechanism is well understood and the scale is not.
Crisis handling. Detection of suicidal ideation is imperfect. Detection of indirect expressions is worse. A system that misses it, or responds with generic reassurance, has failed at the only moment that mattered.
Using it sensibly
A workable division.
Reasonable: understanding a diagnosis, preparing what to say to a doctor, learning what a therapy actually involves, practising a difficult conversation, journalling with prompts, psychoeducation, finding local services, tracking mood, breaking a task down when depression makes it impossible to start.
Not: crisis, active suicidal thoughts, medication decisions, trauma processing without a professional, or a replacement for treatment you have access to. And be aware of the specific risk that a system will agree with a distorted account of your own situation, since agreement is what it was trained towards.
If you are in crisis, please contact a local crisis line or emergency service. They exist in most countries, they are free, and they are answered by trained people.
For anyone building this
The standard is not "better than nothing". It is the standard applied to any health intervention: evidence of benefit, a defined scope with hard refusals outside it, escalation paths to human help, clinical governance, and honest disclosure of what the system is. A wellness label on a product that people use as treatment does not change what it is being used for, and regulators in several countries have begun to say so.
The one thing to keep
Structured guided self-help has real evidence and generative conversation has one small trial, while the documented failures — weight-loss advice to eating-disorder callers, agreement with delusional beliefs — come from a system being generally helpful in a domain where safety is counterintuitive.
Before you move on
The Tessa chatbot gave calorie-restriction advice to people contacting an eating disorder helpline. What does this failure most clearly show?
Pick the one you would defend. Nobody sees your answer.