Skip to content
MyFreud

A CBT chatbot cut loneliness in a week, against a waitlist

A trial of a CBT-based AI chatbot in 100 students reduced depression and loneliness in seven days. The comparison group received nothing at all, which matters.

4 min read

Pop-art illustration of a student sitting on a step with a laptop on his knees and a backpack beside him.

Key takeaways

  • The comparison was a waitlist, meaning the control group received nothing. That design reliably overstates how much of the benefit came from the intervention itself.
  • Depression and loneliness improved over seven days. Anxiety did not move, which is a useful sign the results are not simply everyone feeling better about being in a study.
  • Seven days is the entire trial. There is no evidence here about whether anything lasted.
  • The finding that financially stressed students benefited most is exploratory, from subgroups of about 27 people each.
  • One hundred students at Chinese universities is a specific population, and the chatbot was culturally adapted for them.

Chatbot mental health tools are being built far faster than they are being evaluated, so a randomised trial of one is worth reading carefully. A study published in JMIR mHealth and uHealth tested a CBT-based chatbot with university students, and produced a positive result whose size depends heavily on what it was compared against. [pubmed-loneliness-jul08-2026-source]

What the researchers did

One hundred students at Chinese universities, average age 21 and around 62% female, were randomly assigned to one of two conditions:

  • Intervention: interact with a culturally adapted, CBT-based AI chatbot for seven consecutive days
  • Control: a waitlist, receiving nothing during the trial

Depression, anxiety and loneliness were measured at baseline, day three and day seven, using standard scales. Financial stress was measured separately.

What they found

Depression and loneliness both improved significantly in the chatbot group and did not change in the waitlist group. The effect sizes were moderate, around 0.7 for depression and 0.6 for loneliness.

Anxiety did not change in either group.

An exploratory analysis found that students reporting high financial stress improved considerably more than those reporting low financial stress, on both depression and loneliness.

The waitlist problem

This is the central thing to understand about the result, and it applies to a very large share of digital mental health research.

The control group got nothing. They knew they got nothing. They had no activity, no daily engagement, nobody paying attention to them, and no expectation of improvement.

So the comparison is not “chatbot versus something else”. It is “chatbot versus being on a list”, and the gap between them includes everything that comes with receiving any structured intervention: the daily prompt to reflect, the sense of doing something about the problem, the expectation that it will help.

Waitlist-controlled trials systematically produce larger effect sizes than trials with active comparisons. The finding is real; the magnitude is inflated by the design, and the question of whether the CBT content specifically mattered is left untouched.

What makes the result more interesting than it might be

The anxiety null.

If the entire effect were attention and expectation, you would expect all three self-reported measures to drift together. They did not. Depression and loneliness moved, anxiety did not, in the same people over the same week.

That pattern is harder to explain by generic study effects than by something specific happening on the two measures that moved. It is not proof, and it is a genuine point in the trial’s favour that the researchers reported the null instead of quietly dropping it.

What seven days cannot tell you

The trial ended when the intervention ended. There is no follow-up.

That leaves the most important practical question unanswered. A week of daily structured reflection improving mood is plausible and not very surprising. Whether anything remains a month later, once the novelty has gone and the chatbot has been deleted, is the thing anyone deciding whether to use one would want to know.

The financial-stress finding needs similar caution. It comes from splitting 100 people into subgroups of roughly 27, which is small enough that the estimate is unstable, and some of the apparent advantage may simply be that people who start worse have more room to improve.

Is your loneliness the kind an app can touch?

Loneliness is not one thing. Tick what fits, to see which kind yours looks like.

0 of 6 ticked

How to read digital mental health trials generally

This study is a reasonable example of a genre, and the questions worth asking of any of them are the same:

  • What was the comparison? A waitlist inflates the effect. An active alternative tells you whether the specific approach matters.
  • How long was the follow-up? If it ends when the intervention ends, the result is about the week, not about the treatment.
  • Who was studied? University students in a single country are not a general population, and a culturally adapted tool is adapted to somebody in particular.
  • Were the nulls reported? A paper that reports a measure that did not move is usually more trustworthy than one where everything worked.

The source

These findings are drawn from “Effect of a Cognitive Behavioral Therapy-Based AI Chatbot on Depression and Loneliness in Chinese University Students: Randomized Controlled Trial With Financial Stress Moderation” (Wang Y, Li X, Zhang Q, et al., 2025), published in JMIR mHealth and uHealth. Read the full study on PubMed.

Frequently asked questions

What is a waitlist control and why is it a weak comparison?

Participants in a waitlist group are told they will receive the intervention later and meanwhile get nothing. They know they were not selected, they have no activity, no attention and no expectation of improvement. So the comparison measures the intervention plus everything that comes with receiving any intervention at all. A stronger design compares against an active alternative, which is how you find out whether the specific content matters.

Why is it interesting that anxiety did not improve?

Because if the whole result were down to attention and expectation, you would expect every self-reported measure to drift in the same direction. Depression and loneliness moved and anxiety did not, which suggests something more specific was happening. It is the strongest internal evidence the trial provides that the effect was not purely generic.

Do AI chatbots work for mental health?

The evidence base is young, and it is mostly short trials against waitlists. Structured CBT-style chatbots do seem to produce short-term improvements in low mood in non-clinical populations, which is what this trial shows. What remains poorly established is whether effects persist, whether they help people with diagnosed conditions, and how they compare against a real therapist rather than against nothing.

Why would financially stressed students benefit more?

The study cannot say, and this was an exploratory subgroup finding with roughly 27 people per group, so it should be treated as a hypothesis. One plausible reading is access: students under financial strain are least able to pay for support and may have started with more to gain. Another is simply that people who begin with worse scores have more room to improve, which is a statistical artefact rather than an effect.

Should I use a chatbot if I feel lonely?

It is unlikely to do harm, may help in the short term, and is not a substitute for the thing loneliness actually responds to, which is contact with people. The more useful framing is as a stepping stone: something that can lift mood enough to make the harder step of reaching out feel possible. If loneliness is persistent and affecting your health or functioning, it is worth raising with a doctor.

References

  1. 1.Wang Y, Li X, Zhang Q, et al. ( 2025). Effect of a Cognitive Behavioral Therapy-Based AI Chatbot on Depression and Loneliness in Chinese University Students: Randomized Controlled Trial With Financial Stress Moderation. JMIR mHealth and uHealth. pubmed.ncbi.nlm.nih.gov . doi:10.2196/63806