Automatically generating short summaries from users' online mental health posts could save counselors' reading time and reduce their fatigue so that they can provide timely responses to those seeking help for improving their mental state. Recent Transformers-based summarization models have presented a promising approach to abstractive summarization. They go beyond sentence selection and extractive strategies to deal with more complicated tasks such as novel word generation and sentence paraphrasing. Nonetheless, these models have a prominent shortcoming; their training strategy is not quite efficient, which restricts the model's performance. In this paper, we include a curriculum learning approach to reweigh the training samples, bringing about an efficient learning procedure. We apply our model on extreme summarization dataset of MentSum posts -- a dataset of mental health related posts from Reddit social media. Compared to the state-of-the-art model, our proposed method makes substantial gains in terms of Rouge and Bertscore evaluation metrics, yielding 3.5% (Rouge-1), 10.4% (Rouge-2), and 4.7% (Rouge-L), 1.5% (Bertscore) relative improvements.
翻译:自动从用户的在线心理健康帖子中生成简短摘要,能够节省咨询师的阅读时间并减轻其疲劳,从而使其能够及时回应那些寻求改善心理状态的人。近年来,基于Transformer的摘要模型为抽象摘要生成提供了一种有前景的方法。这类模型超越了句子选择和抽取式策略,能够处理更复杂的任务,如新词生成和句子释义。然而,这些模型存在一个显著缺陷:其训练策略效率不高,限制了模型的性能。在本文中,我们引入了一种课程学习方法对训练样本进行重新加权,从而实现了高效的学习过程。我们将模型应用于MentSum帖子的极端摘要数据集——该数据集包含来自Reddit社交媒体的心理健康相关帖子。与最先进的模型相比,我们提出的方法在Rouge和Bertscore评估指标上取得了显著提升,相对改进分别为3.5%(Rouge-1)、10.4%(Rouge-2)、4.7%(Rouge-L)和1.5%(Bertscore)。