Intimacy is an essential element of human relationships and language is a crucial means of conveying it. Textual intimacy analysis can reveal social norms in different contexts and serve as a benchmark for testing computational models' ability to understand social information. In this paper, we propose a novel weak-labeling strategy for data augmentation in text regression tasks called WADER. WADER uses data augmentation to address the problems of data imbalance and data scarcity and provides a method for data augmentation in cross-lingual, zero-shot tasks. We benchmark the performance of State-of-the-Art pre-trained multilingual language models using WADER and analyze the use of sampling techniques to mitigate bias in data and optimally select augmentation candidates. Our results show that WADER outperforms the baseline model and provides a direction for mitigating data imbalance and scarcity in text regression tasks.
翻译:亲密性是人际关系的重要组成部分,而语言是传递亲密性的关键手段。文本亲密性分析能够揭示不同语境中的社会规范,并可作为测试计算模型理解社会信息能力的基准。本文提出一种名为WADER的新型弱标注策略,用于文本回归任务中的数据增强。WADER通过数据增强解决数据不平衡与数据稀缺问题,并提供了跨语言零样本任务中的数据增强方法。我们利用WADER对当前最优预训练多语言模型进行性能基准测试,同时分析采样技术在缓解数据偏差及优化增强候选样本选取中的作用。实验结果表明,WADER的性能优于基线模型,为缓解文本回归任务中的数据不平衡与稀缺问题提供了解决方向。