In this work, we propose a method for extracting text spans that may indicate one of the BIG5 psychological traits using a question-answering task with examples that have no answer for the asked question. We utilized the RoBERTa model fine-tuned on SQuAD 2.0 dataset. The model was further fine-tuned utilizing comments from Reddit. We examined the effect of the percentage of examples with no answer in the training dataset on the overall performance. The results obtained in this study are in line with the SQuAD 2.0 benchmark and present a good baseline for further research.
翻译:在本文中,我们提出了一种利用问答任务从文本中抽取可能指示大五人格特质(BIG5)的文本片段的方法,该任务包含对提问问题无答案的样本。我们采用了在SQuAD 2.0数据集上微调的RoBERTa模型,并进一步利用Reddit评论进行微调。我们考察了训练数据集中无答案样本比例对整体性能的影响。本研究所获得的结果与SQuAD 2.0基准测试结果一致,为后续研究提供了良好的基线。