Quick answer: CounselBench, a benchmark built with 100 mental health professionals, found that when you let an AI grade another AI’s mental health answers, the safety problems go missing. CounselBench collected 2,000 expert evaluations of answers from GPT-4, LLaMA 3, Gemini, and human therapists, then stress-tested nine models against 120 adversarial questions written by clinicians. The models scored well on several dimensions and still produced recurring failures (unconstructive feedback, overgeneralization, thin personalization) and were frequently flagged for safety risks, most often unauthorized medical advice. The line that should travel furthest: “LLM judges systematically overrate model responses and overlook safety concerns identified by human experts.” By Matthew Sexton, LCSW, NATC — a Licensed Clinical Social Worker and Certified Narcissistic Abuse Treatment Clinician in private practice.
Almost every claim you have read about an AI system being safe for mental health use rests on a benchmark. And many of those benchmarks were scored by a language model rather than a person, because paying a hundred clinicians to read output is expensive and asking GPT to rate GPT costs almost nothing. CounselBench does not measure how common that practice is. What it measures is what happens to the safety findings when you use it.
A group of researchers did the expensive version. What they found is not that the models are terrible. It is that the cheap method of checking them is blind in exactly the place you would least want it to be.
What CounselBench actually measured
CounselBench has two halves, and the distinction matters.
The evaluation half took real questions people posted to CounselChat, a public forum where members of the public ask mental health questions and therapists answer. Answers from GPT-4, LLaMA 3, Gemini, and the online human therapists were rated by mental health professionals across six clinically grounded dimensions. Crucially, the experts did not just assign scores. They marked specific spans of text and wrote out rationales: this phrase, here, is the problem, and here is why. That produced 2,000 expert evaluations.
The adversarial half went hunting. Clinicians wrote 120 questions specifically designed to trip models into particular failures, the kind of question where a plausible-sounding answer is the dangerous one. Run across nine models, that produced 1,080 responses to examine.
A hundred mental health professionals were involved in building and evaluating this. That number is the reason the study can say anything at all about safety, and it is also the reason almost nobody else does it this way.
Why LLM-as-judge safety scores are unreliable
Here is the sentence, quoted from the paper: LLM judges systematically overrate model responses and overlook safety concerns identified by human experts.
Sit with the word systematically. Not occasionally, not noisily. In a consistent direction — toward the flattering answer.
This matters far beyond one paper, because LLM-as-judge is routine — cheap, fast, and scalable in a way expert review never is. So when a company reports that its model scored well on a safety evaluation, the first thing worth asking is whether the grader was a person. Vendors do not always say. And according to this study, in mental health content the automated method reliably fails to see the category of problem that actually hurts someone.
So a passing score from an automated judge does not mean the output was safe. It means nothing flagged it. Those are very different statements, and the industry routinely reports the second as though it were the first.
| What differed | 100 mental health professionals | LLM judge |
|---|---|---|
| Method | Span-level annotation plus written rationale | Score assignment |
| Safety concerns | Flagged, most often unauthorized medical advice | Systematically overlooked |
| Overall ratings | Lower | Systematically higher |
| Cost per evaluation | High | Near zero |
What the models got wrong, specifically
The paper reports consistent, model-specific failure patterns, meaning each model tends to fail in its own recognizable way, rather than errors scattering randomly. Three recur:
- Unconstructive feedback. Responses that are supportive-sounding and do nothing. Anyone who has sat with a client through a hard week knows the difference between being heard and being handled.
- Overgeneralization. Answers that flatten a specific person’s situation into the general case.
- Limited personalization or relevance. Content that could have been written before the question was asked.
And separately, the safety finding: responses were frequently flagged for safety risks, most notably unauthorized medical advice. That is the model telling someone something about medication, or a condition, or a course of treatment, that it is in no position to tell them.
Worth holding onto: the models scored high on several dimensions. The paper never argues the technology is useless. Its argument is narrower and harder: fluent output, frequently decent, holed in specific repeatable places — a worse problem than being obviously bad, because fluent output does not look like it needs checking.
What CounselBench does not show about AI therapy
I want to draw this line clearly, because the study will get stretched in both directions this week.
CounselBench tested question answering on a public forum. It did not test psychotherapy. There was no course of treatment, no therapeutic relationship, no diagnosis, and no clinical outcome measured. Nobody in this study was anyone’s patient.
It also does not establish that human therapists “won.” The human answers were rated too, and the paper’s contribution is the evaluation infrastructure and the failure taxonomy, not a scoreboard. Anyone telling you this study proves therapists beat GPT-4 is adding something that is not there.
And it says nothing about the uses that dominate actual practice right now: drafting a note you then edit, prepping a claim, summarizing your own writing. Those are different tasks with different risk profiles, and a finding about open-ended advice to strangers does not transfer to them.
One honest note about the date
This is not brand-new research, and you will see it reported as though it were.
The preprint went up June 10, 2025, and has been revised since, most recently in May 2026. What is new is the venue: USC reports it was accepted to ICLR 2026, and the University of Southern California published a write-up this month, which is why it is circulating now.
I mention this because “new study finds” is doing a lot of unearned work in AI coverage generally. A paper’s findings do not get better because a conference accepted it, and they were not worse the year before. If a claim’s credibility depends on the word new, check the date.
The practical version: what to ask a vendor
If you are choosing tools for a practice, this paper hands you one very good question and a couple of follow-ups.
“Who evaluated this, and were they people?”
Then:
- Were the evaluators clinicians? In mental health content, the distinction between supportive and unconstructive is a clinical judgment. A general-purpose crowd worker cannot reliably make it.
- Did the evaluation look for unauthorized medical advice specifically? That was the most-flagged safety category here. If a vendor’s safety evaluation does not name it, ask why.
- Was any part of the scoring done by a model? That alone does not disqualify a tool. For some dimensions model scoring is fine. But if the safety rating came from a model, this paper is direct evidence that the number is optimistic.
- Can you see the failures? The strongest signal a vendor can give you is a list of the things their tool does badly. CounselBench’s whole value is that it catalogs failures. A vendor with no such list has either not looked or is not telling you.
I have written before about who holds authority when an AI scribe drafts the note, and the governance question here is the same one. This paper adds a narrower point underneath that: the measurement layer is part of the governance. If you cannot trust how a tool was checked, its architecture barely matters.
Why this is a systems problem, not a technology problem
The thing I keep returning to is that nobody in this story did anything villainous.
Automated evaluation exists because human expert evaluation is expensive, and it is expensive because clinical judgment is scarce and clinicians’ time is already oversubscribed. The incentive to grade with a model is enormous and it is structural. A safety pipeline built out of the cheapest available judgment will systematically produce reassuring numbers, and everyone in the chain can be acting in good faith while the aggregate result is a set of published scores that overstate how safe things are.
That is the same shape as most of what makes this work harder than it needs to be: a system optimizing for measurable throughput, producing an artifact that looks like accountability. Louder skepticism about AI fixes none of it. Being specific does — what was measured, by whom, and what got left out.
A hundred clinicians reading spans of text found things the machines rated as fine. That is the finding. It cost a great deal more than the alternative, and it is the only reason we know.
If you are weighing an AI tool for your practice and want a second read on what its safety claims actually rest on, book a call.
FAQ
What is CounselBench?
CounselBench is a benchmark built with 100 mental health professionals to evaluate how large language models answer real mental health questions. CounselBench-EVAL contains 2,000 expert evaluations of answers from GPT-4, LLaMA 3, Gemini, and online human therapists, drawn from the public forum CounselChat, rated across six clinically grounded dimensions with span-level annotations and written rationales. CounselBench-Adv is an adversarial set of 120 expert-authored questions producing 1,080 responses across nine LLMs. The paper is arXiv:2506.08584, first posted June 10, 2025, accepted to ICLR 2026.
What did the study find about AI judges evaluating AI?
That LLM judges systematically overrate model responses and overlook safety concerns identified by human experts. This has wide practical reach because evaluating model output with another model is the standard cheap method across the industry. If that method reliably misses safety problems in mental health content, a benchmark score produced that way is not evidence of safety. It is evidence that nothing flagged it.
What kinds of mistakes did the models make?
Consistent, model-specific failure patterns rather than random errors: unconstructive feedback, overgeneralization, and limited personalization or relevance. Responses were also frequently flagged for safety risks, most notably unauthorized medical advice. The models scored high on several dimensions, so this is not a blanket dismissal. The picture is competent-sounding output with specific repeatable holes.
Does this study show AI cannot do therapy?
No, and it does not claim to. CounselBench evaluated question answering on a public forum, not psychotherapy, treatment, or diagnosis. No one in the study was a patient of the model, there was no course of care, and no clinical outcome was measured. The narrower and more useful reading is that when experts examined answers closely enough to annotate specific spans, they found safety problems automated evaluation did not surface.
Disclaimer
This article is for educational and informational purposes only. It discusses published research and does not constitute clinical, medical, or legal advice. Matthew Sexton, LCSW, NATC, is a Licensed Clinical Social Worker in private practice and the founder of Mental Wealth Solutions. Reading this article does not create a therapist-client or advisory relationship. The study described here evaluated question answering on a public forum, not psychotherapy, and its findings should not be extended to clinical care decisions.
Related reading
- AI clinical documentation errors and where the risk actually sits
- Governance that keeps the clinician in control of the scribe
- The 53-minute problem: what the system is actually designed around
- 988 funding stayed flat in 2026
Sources
- Li, Y., Yao, J., Bunyi, J. B. S., Frank, A. C., Hwang, A. H.-C., & Liu, R. “CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering.” arXiv:2506.08584. v1 June 10, 2025; revised May 14, 2026.
- OpenReview submission 8MBYRZHVWT (ICLR 2026). openreview.net
- USC Viterbi School of Engineering. “Can ChatGPT Be Your Therapist? USC Study Tests AI Responses to Mental Health Questions,” July 2026. viterbischool.usc.edu
Frequently asked questions.
- What is CounselBench?
- CounselBench is a benchmark built with 100 mental health professionals to evaluate how large language models answer real mental health questions. It has two parts. CounselBench-EVAL contains 2,000 expert evaluations of answers from GPT-4, LLaMA 3, Gemini, and online human therapists, drawn from questions posted to the public forum CounselChat, with each answer rated across six clinically grounded dimensions plus span-level annotations and written rationales. CounselBench-Adv is an adversarial set of 120 expert-authored questions designed to trigger specific model failures, producing 1,080 responses across nine LLMs. The paper is arXiv:2506.08584, first posted June 10, 2025, and accepted to ICLR 2026.
- What did the study find about AI judges evaluating AI?
- It found that LLM judges systematically overrate model responses and overlook safety concerns identified by human experts. This is the finding with the widest practical reach, because evaluating model output with another model (often called LLM-as-judge) is the standard cheap method across the industry. If that method reliably misses safety problems in mental health content, then a benchmark score produced that way is not evidence of safety. It is evidence that nothing flagged it.
- What kinds of mistakes did the models make?
- The paper reports consistent, model-specific failure patterns rather than random errors, recurring as unconstructive feedback, overgeneralization, and limited personalization or relevance. Responses were also frequently flagged for safety risks, most notably unauthorized medical advice. Notably the models scored high on several dimensions, so this is not a blanket dismissal. The picture is competent-sounding output with a specific and repeatable set of holes.
- Does this study show AI cannot do therapy?
- No, and it does not claim to. CounselBench evaluated question answering on a public forum, not psychotherapy, treatment, or diagnosis. Nobody in the study was a patient of the model, there was no course of care, and no clinical outcome was measured. The honest reading is narrower and more useful: when experts examined answers closely enough to annotate specific spans of text, they found safety problems that automated evaluation did not surface.
If you're the therapist here.
Your clients get 4 sessions a month. The other 26 days they're on their own. VibeCheck is the between-session companion that carries those days back to you — clients check in daily, and you walk in already knowing what kind of week it was. Built by Matthew Sexton, LCSW, NATC.