Factual correctness is often the limiting factor in practical applications of natural language generation in high-stakes domains such as healthcare. An essential requirement for maintaining factuality is the ability to deal with rare tokens. This paper focuses on rare tokens that appear in both the source and the reference sequences, and which, when missed during generation, decrease the factual correctness of the output text. For high-stake domains that are also knowledge-rich, we show how to use knowledge to (a) identify which rare tokens that appear in both source and reference are important and (b) uplift their conditional probability. We introduce the ``utilization rate'' that encodes knowledge and serves as a regularizer by maximizing the marginal probability of selected tokens. We present a study in a knowledge-rich domain of healthcare, where we tackle the problem of generating after-visit care instructions based on patient-doctor dialogues. We verify that, in our dataset, specific medical concepts with high utilization rates are underestimated by conventionally trained sequence-to-sequence models. We observe that correcting this with our approach to knowledge injection reduces the uncertainty of the model as well as improves factuality and coherence without negatively impacting fluency.
翻译:事实正确性通常是医疗等高风控领域中自然语言生成实际应用的关键限制因素。维持事实准确性的核心要求是具备处理罕见词元的能力。本文聚焦于同时出现在源序列与参考序列中的罕见词元——这些词元在生成过程中若被遗漏,将降低输出文本的事实正确性。针对兼具高风控与高知识密度的领域,我们展示了如何利用知识实现两方面目标:(a) 识别同时出现在源序列与参考序列中的关键性罕见词元;(b) 提升其条件概率。我们提出"利用率"概念:该指标通过编码知识并最大化选定词元的边际概率,起到正则化作用。本研究在医疗这一高知识密度领域展开,致力于解决基于医患对话生成诊后护理指导的问题。实验验证,在数据集中,特定具有高利用率的医学概念被传统序列到序列模型所低估。我们观察到,采用知识注入方法修正该偏差后,不仅降低了模型的不确定性,还提升了事实正确性与连贯性,且未对流畅性产生负面影响。