Improving mental health support in developing countries is a pressing need. One potential solution is the development of scalable, automated systems to conduct diagnostic screenings, which could help alleviate the burden on mental health professionals. In this work, we evaluate several state-of-the-art Large Language Models (LLMs), with and without fine-tuning, on our custom dataset for generating concise summaries from mental state examinations. We rigorously evaluate four different models for summary generation using established ROUGE metrics and input from human evaluators. The results highlight that our top-performing fine-tuned model outperforms existing models, achieving ROUGE-1 and ROUGE-L values of 0.810 and 0.764, respectively. Furthermore, we assessed the fine-tuned model's generalizability on a publicly available D4 dataset, and the outcomes were promising, indicating its potential applicability beyond our custom dataset.
翻译:改善发展中国家精神健康支持是一项紧迫需求。潜在解决方案之一是开发可扩展的自动化系统以进行诊断筛查,这有助于减轻精神健康专业人员的负担。在本研究中,我们基于自建数据集评估了多种最先进的大语言模型(含微调与未微调版本),用于从精神状态检查中生成简洁摘要。我们采用成熟的ROUGE指标与人工评估相结合的方式,对四种不同模型的摘要生成能力进行了严格评估。结果表明,我们表现最佳的微调模型优于现有模型,其ROUGE-1与ROUGE-L值分别达到0.810与0.764。此外,我们在公开数据集D4上评估了该微调模型的泛化能力,结果展现了良好潜力,表明该模型在自建数据集之外具有适用性。