Chatbot mental health tools are being built far faster than they are being evaluated, so a randomised trial of one is worth reading carefully. A study published in JMIR mHealth and uHealth tested a CBT-based chatbot with university students, and produced a positive result whose size depends heavily on what it was compared against. [pubmed-loneliness-jul08-2026-source]
What the researchers did
One hundred students at Chinese universities, average age 21 and around 62% female, were randomly assigned to one of two conditions:
- Intervention: interact with a culturally adapted, CBT-based AI chatbot for seven consecutive days
- Control: a waitlist, receiving nothing during the trial
Depression, anxiety and loneliness were measured at baseline, day three and day seven, using standard scales. Financial stress was measured separately.
What they found
Depression and loneliness both improved significantly in the chatbot group and did not change in the waitlist group. The effect sizes were moderate, around 0.7 for depression and 0.6 for loneliness.
Anxiety did not change in either group.
An exploratory analysis found that students reporting high financial stress improved considerably more than those reporting low financial stress, on both depression and loneliness.
The waitlist problem
This is the central thing to understand about the result, and it applies to a very large share of digital mental health research.
The control group got nothing. They knew they got nothing. They had no activity, no daily engagement, nobody paying attention to them, and no expectation of improvement.
So the comparison is not “chatbot versus something else”. It is “chatbot versus being on a list”, and the gap between them includes everything that comes with receiving any structured intervention: the daily prompt to reflect, the sense of doing something about the problem, the expectation that it will help.
Waitlist-controlled trials systematically produce larger effect sizes than trials with active comparisons. The finding is real; the magnitude is inflated by the design, and the question of whether the CBT content specifically mattered is left untouched.
What makes the result more interesting than it might be
The anxiety null.
If the entire effect were attention and expectation, you would expect all three self-reported measures to drift together. They did not. Depression and loneliness moved, anxiety did not, in the same people over the same week.
That pattern is harder to explain by generic study effects than by something specific happening on the two measures that moved. It is not proof, and it is a genuine point in the trial’s favour that the researchers reported the null instead of quietly dropping it.
What seven days cannot tell you
The trial ended when the intervention ended. There is no follow-up.
That leaves the most important practical question unanswered. A week of daily structured reflection improving mood is plausible and not very surprising. Whether anything remains a month later, once the novelty has gone and the chatbot has been deleted, is the thing anyone deciding whether to use one would want to know.
The financial-stress finding needs similar caution. It comes from splitting 100 people into subgroups of roughly 27, which is small enough that the estimate is unstable, and some of the apparent advantage may simply be that people who start worse have more room to improve.
Is your loneliness the kind an app can touch?
Loneliness is not one thing. Tick what fits, to see which kind yours looks like.
0 of 6 ticked
What you are describing is mostly about expectation and withdrawal rather than about having nobody. That is the pattern most likely to respond to structured cognitive work, whether from an app, a self-help course or a therapist, because the belief that contact will go badly is the thing keeping the contact from happening.
Separate the two. The circumstantial part, a move or a change in routine, responds to plans and repeated exposure. The pattern part, the sense of imposing on people, responds to being tested rather than reasoned with. Pick one small instance of the second and test it this week.
If the main issue is that the people are not there yet, the answer is logistical rather than psychological: repeated, low-stakes contact with the same faces. Regularity matters more than depth at the start.
How to read digital mental health trials generally
This study is a reasonable example of a genre, and the questions worth asking of any of them are the same:
- What was the comparison? A waitlist inflates the effect. An active alternative tells you whether the specific approach matters.
- How long was the follow-up? If it ends when the intervention ends, the result is about the week, not about the treatment.
- Who was studied? University students in a single country are not a general population, and a culturally adapted tool is adapted to somebody in particular.
- Were the nulls reported? A paper that reports a measure that did not move is usually more trustworthy than one where everything worked.
The source
These findings are drawn from “Effect of a Cognitive Behavioral Therapy-Based AI Chatbot on Depression and Loneliness in Chinese University Students: Randomized Controlled Trial With Financial Stress Moderation” (Wang Y, Li X, Zhang Q, et al., 2025), published in JMIR mHealth and uHealth. Read the full study on PubMed.