ResearchPod Summary
MyMentorLLM is a multimodal simulation environment designed to facilitate deliberate practice in psychotherapy. The system automates the mentorship cycle by instantiating three distinct roles: a patient grounded in DSM-5-TR clinical cases (Major Depressive Disorder, Generalized Anxiety Disorder, or Borderline Personality Disorder), a trainee therapist calibrated to empirical competence levels of developing clinicians, and an expert mentor. The researchers generated 2,100 complete training sessions across seven model-modality conditions, including native speech-to-speech and text-only configurations. The evaluation focused on whether these models could maintain psychological fidelity, provide accurate diagnostic feedback, and effectively mirror the emotional dynamics of real-world clinical encounters.
The study demonstrates that LLMs can effectively simulate disorder-congruent emotional profiles, with trainee therapists mirroring the communication styles found in human counseling. A critical finding is that the quality of supervision varies significantly across models. While most LLMs tended to overestimate the competence of the trainee, native speech-to-speech models performed closest to human scoring standards. Furthermore, the inclusion of an expert mentor significantly improved the diagnostic accuracy of the simulated trainees in 5 out of 7 models, and symptom identification accuracy showed a positive correlation with model size.
This work shifts the focus of AI in mental health from merely replacing therapists to enhancing the training pipeline. By providing a scalable, consequence-free environment for deliberate practice, MyMentorLLM addresses the critical shortage of expert supervision and the logistical difficulties of training new clinicians. It offers a robust, interpretable testbed for researchers to evaluate how synthetic agents can support clinical education while highlighting the necessity of rigorous safeguards against model-specific biases and harmful feedback.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.