This paper presents models created for the Social Media Mining for Health 2023 shared task. Our team addressed the first task, classifying tweets that self-report Covid-19 diagnosis. Our approach involves a classification model that incorporates diverse textual augmentations and utilizes R-drop to augment data and mitigate overfitting, boosting model efficacy. Our leading model, enhanced with R-drop and augmentations like synonym substitution, reserved words, and back translations, outperforms the task mean and median scores. Our system achieves an impressive F1 score of 0.877 on the test set.
翻译:本文介绍了为2023年社交媒体健康挖掘共享任务构建的模型。我们的团队针对第一项任务——分类自报新冠诊断的推文——提出了解决方案。该方法采用了一个结合多样化文本增强的分类模型,并利用R-drop技术扩充数据、缓解过拟合,从而提升模型效能。我们的领先模型通过引入基于R-drop的增强策略(如同义词替换、保留词和回译),其表现优于任务平均分和中位数得分。该系统在测试集上取得了0.877的出色F1分数。