Recently, aspect sentiment quad prediction has received widespread attention in the field of aspect-based sentiment analysis. Existing studies extract quadruplets via pre-trained generative language models to paraphrase the original sentence into a templated target sequence. However, previous works only focus on what to generate but ignore what not to generate. We argue that considering the negative samples also leads to potential benefits. In this work, we propose a template-agnostic method to control the token-level generation, which boosts original learning and reduces mistakes simultaneously. Specifically, we introduce Monte Carlo dropout to understand the built-in uncertainty of pre-trained language models, acquiring the noises and errors. We further propose marginalized unlikelihood learning to suppress the uncertainty-aware mistake tokens. Finally, we introduce minimization entropy to balance the effects of marginalized unlikelihood learning. Extensive experiments on four public datasets demonstrate the effectiveness of our approach on various generation templates1.
翻译:摘要:近年来,方面情感四元组预测在基于方面的情感分析领域受到广泛关注。现有研究通过预训练生成式语言模型提取四元组,将原始句子改写为模板化的目标序列。然而,以往工作仅关注生成内容,而忽略了不生成的内容。我们认为考虑负样本同样能带来潜在收益。本文提出一种与模板无关的方法来控制词元级生成,在提升原始学习效果的同时减少错误。具体而言,我们引入蒙特卡洛dropout来理解预训练语言模型的内在不稳定性,从而获取噪声和错误信息。进一步地,我们提出边际化非似然学习来抑制面向不确定性感知的错误词元。最后,引入最小化熵来平衡边际化非似然学习的效果。在四个公开数据集上的大量实验表明,我们的方法在各种生成模板上均具有有效性。