Recent advances in large language models (LLMs) have made significant progress across multiple biomedical tasks, including biomedical question answering, lay-language summarization of the biomedical literature, and clinical note summarization. These models have demonstrated strong capabilities in processing and synthesizing complex biomedical information and in generating fluent, human-like responses. Despite these advancements, hallucinations or confabulations remain key challenges when using LLMs in biomedical and other high-stakes domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly
翻译:大型语言模型(LLMs)的最新进展已在多项生物医学任务中取得显著突破,包括生物医学问答、生物医学文献的通俗语言摘要以及临床笔记摘要。这些模型在处理与整合复杂生物医学信息,以及生成流畅、类人响应方面展现出强大能力。尽管取得这些进展,当将LLMs应用于生物医学及其他高风险领域时,幻觉或虚构现象仍是关键挑战。在诸如医学问答、临床决策制定或生物医学研究评估等高风险情境中,信息不准确可能尤为有害。关于评估LLMs将生成内容锚定于可验证来源能力的研究表明,模型表现显著(原文未完成,故保留“显著”二字)。