The pervasive influence of misinformation has far-reaching and detrimental effects on both individuals and society. The COVID-19 pandemic has witnessed an alarming surge in the dissemination of medical misinformation. However, existing datasets pertaining to misinformation predominantly focus on textual information, neglecting the inclusion of visual elements, and tend to center solely on COVID-19-related misinformation, overlooking misinformation surrounding other diseases. Furthermore, the potential of Large Language Models (LLMs), such as the ChatGPT developed in late 2022, in generating misinformation has been overlooked in previous works. To overcome these limitations, we present Med-MMHL, a novel multi-modal misinformation detection dataset in a general medical domain encompassing multiple diseases. Med-MMHL not only incorporates human-generated misinformation but also includes misinformation generated by LLMs like ChatGPT. Our dataset aims to facilitate comprehensive research and development of methodologies for detecting misinformation across diverse diseases and various scenarios, including human and LLM-generated misinformation detection at the sentence, document, and multi-modal levels. To access our dataset and code, visit our GitHub repository: \url{https://github.com/styxsys0927/Med-MMHL}.
翻译:虚假信息的广泛传播对个人和社会产生深远且有害的影响。COVID-19疫情期间,医学虚假信息的传播出现了令人担忧的激增。然而,现有关于虚假信息的数据集大多聚焦于文本信息,忽视了视觉元素的纳入,且往往仅关注COVID-19相关虚假信息,忽略了其他疾病的虚假信息。此外,大型语言模型(LLMs,如2022年底开发的ChatGPT)在生成虚假信息方面的潜力在先前研究中未被充分重视。为克服这些局限,我们提出了Med-MMHL,这是一个涵盖多种疾病、面向通用医学领域的新型多模态虚假信息检测数据集。Med-MMHL不仅包含人类生成的虚假信息,还纳入了ChatGPT等LLMs生成的虚假信息。本数据集旨在促进针对不同疾病和多种场景下虚假信息检测方法的综合研究与发展,包括句子级、文档级和多模态层面的人类与LLM生成虚假信息检测。如需访问我们的数据集和代码,请访问GitHub仓库:\url{https://github.com/styxsys0927/Med-MMHL}。