User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages.
翻译:用户生成数据源在揭示不良药物反应(ADRs)方面日益重要,数字世界中相关讨论也愈发频繁。然而,现有临床语料库主要围绕英文科学文献。本研究提出了一个包含德语、法语和日语来源文本的多语种不良药物反应语料库,这些文本采集自患者论坛、社交媒体和临床报告等不同渠道。我们的语料库涵盖12种实体类型、4种属性类型和13种关系类型的标注。该语料库有助于开发面向医疗保健领域的真实世界多语种语言模型。我们提供了相关统计数据以突出语料库面临的特定挑战,并进行了初步实验,为跨语言及单语言场景下的实体及实体关系抽取建立了强基线。