Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet.
翻译:在事实核查与忠实摘要领域,针对参考文章中信息误述问题的研究已取得显著进展。然而,一个尚未被触及的方面依然存在——识别社交媒体帖子中对关联新闻文章信息的操纵行为。该任务极具挑战性,主要源于此类帖子中普遍存在的个人观点。我们提出了一项新任务——识别社交媒体上的新闻操纵行为,旨在检测社交媒体帖子中的操纵痕迹,并识别被操纵或插入的信息。为研究该任务,我们设计了数据收集方案,并构建了名为ManiTweet的数据集,包含3600对推文及对应新闻文章。分析表明,该任务极具难度,大语言模型的性能表现并不理想。此外,我们开发了一个简洁有效的基准模型,其在ManiTweet数据集上的表现显著优于大语言模型。最后,我们对人工撰写的推文进行了探索性分析,揭示了操纵行为与新闻文章领域及事实性之间的有趣关联,同时发现被操纵的句子更可能概括报道核心情节或结局。