This paper presents various automatic detection methods to extract so called tortured phrases from scientific papers. These tortured phrases, e.g. flag to clamor instead of signal to noise, are the results of paraphrasing tools used to escape plagiarism detection. We built a dataset and evaluated several strategies to flag previously undocumented tortured phrases. The proposed and tested methods are based on language models and either on embeddings similarities or on predictions of masked token. We found that an approach using token prediction and that propagates the scores to the chunk level gives the best results. With a recall value of .87 and a precision value of .61, it could retrieve new tortured phrases to be submitted to domain experts for validation.
翻译:本文提出了多种自动检测方法,用于提取科学论文中所谓的"折磨短语"。这些折磨短语(例如用"flag to clamor"替代"signal to noise")是通过改写工具规避抄袭检测的产物。我们构建了一个数据集,并评估了多种识别此前未被记录的折磨短语的策略。所提出并测试的方法基于语言模型,具体包括嵌入相似度方法和掩码标记预测方法。研究发现,采用标记预测并将得分传播至语块级别的方法效果最佳。该方法在召回率达到0.87、精确率达到0.61的情况下,能够检索出新的折磨短语,供领域专家进行验证。