Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and safety often requires substantial compute, expert input, and labeled data. At the same time, deploying proprietary, cloud-based models for mental health-related interactions raises important privacy and data-governance concerns, given the sensitivities. To address this challenge, we introduce LLUMI setup that can be hosted in-house within protected environments. LLUMI consists of two complementary components: a generation model (GM), which drafts supportive responses to mental health queries, and an improvement model (IM), which revises an initial human-crafted response. We leverage feedback signals from Reddit mental health communities, using community endorsement patterns such as upvotes and downvotes to construct chosen-rejected response pairs for Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO). We further align LLUMI using human evaluation across five dimensions: readability, empathy, connection, actionability, and safety. Our results show that, despite relying on smaller open-source models rather than proprietary cloud-based GPT models, LLUMI achieves comparable performance across linguistic analyses and human evaluations. These findings suggest that open-source models, when trained with community-derived preference signals, can support high-quality mental health support assistance while offering a more privacy-preserving alternative for sensitive support contexts.
翻译:摘要:大语言模型(LLMs)在生成心理健康相关咨询的支持性回复方面展现出潜力,但提升其实用性、共情能力和安全性通常需要大量计算资源、专家输入和标注数据。同时,将专有云模型部署于心理健康交互场景,因其敏感性会引发严重的隐私与数据治理问题。针对这一挑战,我们提出可在受保护环境中内部部署的LLUMI框架。LLUMI由两个互补组件构成:生成模型(GM)负责起草心理健康咨询的支持性回复,改进模型(IM)则对人类初始撰写的回复进行修订。我们利用Reddit心理健康社区中的反馈信号,通过点赞与点踩等社区认可模式构建"入选-淘汰"回复对,用于监督微调(SFT)与直接偏好优化(DPO)。进一步,我们通过可读性、共情力、联结度、可操作性和安全性五个维度的人类评估对LLUMI进行对齐。结果表明,尽管依赖的是小型开源模型而非专有云GPT模型,LLUMI在语言分析与人类评估中仍能达到可比性能。这些发现表明,使用社区衍生偏好信号训练的开源模型,既能支持高质量心理健康辅助,又为敏感支持场景提供了更注重隐私保护的替代方案。