While sentence simplification is an active research topic in NLP, its adjacent tasks of sentence complexification and same-level paraphrasing are not. To train models on all three tasks, we present two new unsupervised datasets. We compare these datasets, one labeled by a weak classifier and the other by a rule-based approach, with a single supervised dataset. Using these three datasets for training, we perform extensive experiments on both multitasking and prompting strategies. Compared to other systems trained on unsupervised parallel data, models trained on our weak classifier labeled dataset achieve state-of-the-art performance on the ASSET simplification benchmark. Our models also outperform previous work on sentence level targeting. Finally, we establish how a handful of Large Language Models perform on these tasks under a zero-shot setting.
翻译:虽然句子简化是自然语言处理中一个活跃的研究课题,但其相邻任务——句子复杂化与同级改写——尚未得到充分探索。为训练模型完成这三项任务,我们提出了两个新的无监督数据集。我们将这些数据集(一个由弱分类器标注,另一个基于规则方法生成)与单个有监督数据集进行了比较。利用这三个数据集进行训练,我们在多任务和提示策略上进行了大量实验。与其他基于无监督平行数据训练的系统相比,在我们弱分类器标注数据集上训练的模型在ASSET简化基准上取得了最先进的性能。我们的模型在句子层级定位方面也超越了先前的工作。最后,我们探究了少量大型语言模型在零样本设定下执行这些任务的表现。