Move by Move: Measuring and Guiding How Large Language Models Conduct Psychotherapy
This paper addresses an increasingly common yet under-studied phenomenon: users are turning to large language models for emotional support, yet there is almost no measurable analytical framework for how these models actually conduct a psychotherapeutic interaction. The authors propose an ontology of ten categories of 'therapeutic actions'—compact, function-based classifications rooted in the MULTI-60 checklist, validated through annotation by five licensed psychologists and scaled via a judge-based method whose agreement matches expert-level consistency. Applied to real counseling transcripts and model-led conversations, the authors compare action distributions between human therapists and a panel of frontier models: models use inquiry up to three times as much as humans, neglect psychoeducation, and are strongly anchored to context—picking up strategies initiated by human therapists while rarely initiating on their own. Released as a toolkit, this ontology halves the mean deviation from the human action distribution and improves per-turn alignment by 7-9 percentage points without any fine-tuning.
Background and Context
Users are increasingly turning to large language models for emotional support and psychological companionship, yet researchers have lacked a measurable framework for describing how these systems actually conduct a psychotherapeutic interaction. Without a structured account of the therapeutic process, developers cannot reliably assess quality, surface problems, or guide improvement in emotionally supportive applications. This gap motivated the work summarized here, which draws on an arXiv preprint to address a phenomenon that has spread faster than its measurement tools.
The paper's central contribution is an ontology of ten categories of "therapeutic moves." These units are designed to be compact and function-based, and they trace their lineage to the MULTI-60 checklist, a clinical assessment instrument. Rooting the ontology in MULTI-60 is deliberate: it keeps the new categories semantically connected to an existing evaluation framework rather than inventing a parallel vocabulary. The ten categories span common therapeutic behavior units, ranging from inquiry and clarification to psychoeducation and reframing.
Rather than building the ontology purely theoretically, the authors validated it through annotation by five licensed psychologists. This expert labeling establishes clinical validity and reliability. To handle the scale of real-world conversation data without incurring the high cost of expert annotation throughout, the researchers then introduced a judge-based scaling method. They verified that this automated approach achieves agreement matching expert-level consistency, striking a balance between professional rigor and broad usability. The result is a reusable measurement baseline for quantifying model therapeutic behavior.
Deep Analysis
The technical strategy rests on decomposing abstract "therapeutic quality" into observable, statistically tractable action distributions. By encoding the therapeutic process as ten measurable categories, the authors convert a qualitative interaction into a quantitative distribution. This shift makes the question of whether a model "talks like a qualified therapist" answerable with precision instead of subjective impression.
Applying the ontology to both real counseling transcripts and model-led conversations, the authors compared action distributions between human therapists and a panel of frontier models. The comparison revealed a clear set of behavioral deviations worth watching. Models used inquiry up to three times as often as humans, signaling an over-investigative tendency. At the same time, they neglected psychoeducation, the action responsible for transmitting knowledge and supporting cognitive reframing. Perhaps more consequential is the models' strong anchoring to context: they tend to continue strategies that human therapists have already initiated, while rarely launching such strategies on their own.
This passive-following pattern suggests that models function more as responders than as leaders, which could weaken the directionality and initiative of the therapeutic process. To address it, the researchers released the ontology as a toolkit for the models to use. Remarkably, this improvement occurred without any fine-tuning. The intervention halved the mean deviation of the model action distribution from the human distribution and improved per-turn alignment with human therapists by 7 to 9 percentage points. This ablation-style finding indicates that structured process guidance alone, without modifying model weights, can substantially improve behavioral alignment.
Industry Impact
The paper's value extends well beyond constructing a single benchmark. It provides the first measurable, verifiable, and scalable analytical framework for how large language models conduct psychotherapy, opening the black box of quality assessment for emotionally supportive applications. For the open-source community, the ontology and its annotation workflow serve as reusable tools that lower the barrier to entry for subsequent researchers in this field.
For industrial deployment, the deviations the study surfaces—over-inquiry, neglect of psychoeducation, and passive following—offer concrete directions for product design and safety evaluation. The most notable finding is that simply releasing the ontology as a toolkit, with no fine-tuning, meaningfully narrows the behavioral gap between models and human therapists. This delivers a low-cost, low-risk path toward behavioral alignment.
The action ontology can function both as an evaluation benchmark and as a foundation for designing guidance mechanisms. In this way, it pushes emotionally supportive systems from merely being able to converse toward actually conducting therapy, laying a methodological foundation for more reliable and professional AI mental-health assistants.
Outlook
The judge-based scaling method demonstrates that expert-level agreement can be approximated at volume, a result that may encourage similar validation strategies in other clinical-adjacent domains. Keeping the ontology anchored to MULTI-60 suggests future work could tie new assessments directly to established clinical standards rather than building isolated rubrics.
The finding that alignment improves without fine-tuning points toward lightweight, process-level interventions as a viable alternative to weight modification, potentially reducing the cost and risk of aligning emotionally supportive models. Releasing the framework as an open toolkit invites the community to extend the ten categories, refine the judge model, and test the ontology across additional languages and therapeutic modalities.
The behavioral deviations documented here—excessive inquiry, missing psychoeducation, and context anchoring—likely represent starting points for targeted remediation rather than fixed limitations. As the toolkit is adopted, subsequent research can measure whether guided process interventions sustain alignment over longer conversations and across more diverse user populations, ultimately shaping whether AI mental-health assistants can be evaluated and improved with clinical-grade precision.