There are two competing approaches for modelling annotator disagreement: distributional soft-labelling approaches (which aim to capture the level of disagreement) or modelling perspectives of individual annotators or groups thereof. We adapt a multi-task architecture -- which has previously shown success in modelling perspectives -- to evaluate its performance on the SEMEVAL Task 11. We do so by combining both approaches, i.e. predicting individual annotator perspectives as an interim step towards predicting annotator disagreement. Despite its previous success, we found that a multi-task approach performed poorly on datasets which contained distinct annotator opinions, suggesting that this approach may not always be suitable when modelling perspectives. Furthermore, our results explain that while strongly perspectivist approaches might not achieve state-of-the-art performance according to evaluation metrics used by distributional approaches, our approach allows for a more nuanced understanding of individual perspectives present in the data. We argue that perspectivist approaches are preferable because they enable decision makers to amplify minority views, and that it is important to re-evaluate metrics to reflect this goal.
翻译:对于标注者分歧的建模存在两种竞争性方法:分布式软标注方法(旨在捕捉分歧程度)或建模个体标注者及其群体的视角。我们采用了一种此前在视角建模中表现优异的多任务架构,评估其在SEMEVAL任务11上的性能。具体做法是将两种方法结合,即将预测个体标注者视角作为预测标注者分歧的中间步骤。尽管该方法此前成效显著,我们发现多任务方法在存在明显标注者意见差异的数据集上表现不佳,这表明该架构在视角建模中可能并非始终适用。此外,我们的结果说明:虽然强视角主义方法可能无法在分布法采用的评估指标上达到最先进性能,但该方法能更细致地理解数据中存在的个体视角。我们认为视角主义方法更优,因其能帮助决策者放大少数观点,并有必要重新评估指标以反映这一目标。