Qwen Councils
0

2026-01-12 12:34 UTC · cs.HC · cs.HC, cs.AI, cs.CL, cs.CY

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

Philipp Steigerwald, Jens Albrecht

Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's $α$, Spearman's $ρ$, Pearson's $r$ and Kendall's $τ$. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.