While state-of-the-art NLP models have demonstrated excellent performance for aspect based sentiment analysis (ABSA), substantial evidence has been presented on their lack of robustness. This is especially manifested as significant degradation in performance when faced with out-of-distribution data. Recent solutions that rely on counterfactually augmented datasets show promising results, but they are inherently limited because of the lack of access to explicit causal structure. In this paper, we present an alternative approach that relies on non-counterfactual data augmentation. Our proposal instead relies on using noisy, cost-efficient data augmentations that preserve semantics associated with the target aspect. Our approach then relies on modelling invariances between different versions of the data to improve robustness. A comprehensive suite of experiments shows that our proposal significantly improves upon strong pre-trained baselines on both standard and robustness-specific datasets. Our approach further establishes a new state-of-the-art on the ABSA robustness benchmark and transfers well across domains.
翻译:尽管最先进的自然语言处理模型在方面级情感分析任务上展现出卓越性能,但大量证据表明其缺乏鲁棒性。这种脆弱性尤其体现在面对分布外数据时性能显著下降。近期依赖反事实增强数据集的方法虽具潜力,却因难以获取明确因果结构而存在固有限制。本文提出一种基于非反事实数据增强的替代方案,该方法采用保留目标方面语义的含噪且低成本的数据增强策略,通过建模不同数据版本间的恒等关系来提升鲁棒性。综合实验表明,本方法在标准数据集和鲁棒性专项数据集上均显著优于强预训练基线模型,不仅刷新了方面级情感分析鲁棒性基准的最新纪录,还展现出优秀的跨领域迁移能力。