Safety
From 'Help' to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates el
arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's alpha, Spearman's rho, Pearson's r and Kendall's au. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.
Related
- LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback
- HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
- Sociodemographic Biases in Educational Counselling by Large Language Models
Source: arXiv cs.AI | 2026-07-28