With the advancement of telemedicine, both researchers and medical practitioners are working hand-in-hand to develop various techniques to automate various medical operations, such as diagnosis report generation. In this paper, we first present a multi-modal clinical conversation summary generation task that takes a clinician-patient interaction (both textual and visual information) and generates a succinct synopsis of the conversation. We propose a knowledge-infused, multi-modal, multi-tasking medical domain identification and clinical conversation summary generation (MM-CliConSummation) framework. It leverages an adapter to infuse knowledge and visual features and unify the fused feature vector using a gated mechanism. Furthermore, we developed a multi-modal, multi-intent clinical conversation summarization corpus annotated with intent, symptom, and summary. The extensive set of experiments, both quantitatively and qualitatively, led to the following findings: (a) critical significance of visuals, (b) more precise and medical entity preserving summary with additional knowledge infusion, and (c) a correlation between medical department identification and clinical synopsis generation. Furthermore, the dataset and source code are available at https://github.com/NLP-RL/MM-CliConSummation.
翻译:随着远程医疗的进步,研究人员和医疗从业者正携手开发各种技术,以实现诊断报告生成等医疗操作的自动化。本文首先提出了一项多模态临床对话摘要生成任务,该任务基于医患交互(包括文本和视觉信息)生成对话的简洁摘要。我们提出了一种知识注入、多模态、多任务的医疗领域识别与临床对话摘要生成框架(MM-CliConSummation)。该框架利用适配器注入知识和视觉特征,并通过门控机制统一融合特征向量。此外,我们还构建了一个多模态、多意图的临床对话摘要语料库,其中标注了意图、症状和摘要。通过大量定量和定性实验,我们得出以下发现:(a)视觉信息的关键重要性,(b)通过额外知识注入实现更精确且保留医疗实体的摘要,以及(c)医疗科室识别与临床摘要生成之间的相关性。此外,数据集和源代码已公开于 https://github.com/NLP-RL/MM-CliConSummation。