In this survey, we review methods that retrieve multimodal knowledge to assist and augment generative models. This group of works focuses on retrieving grounding contexts from external sources, including images, codes, tables, graphs, and audio. As multimodal learning and generative AI have become more and more impactful, such retrieval augmentation offers a promising solution to important concerns such as factuality, reasoning, interpretability, and robustness. We provide an in-depth review of retrieval-augmented generation in different modalities and discuss potential future directions. As this is an emerging field, we continue to add new papers and methods.
翻译:本综述回顾了为辅助和增强生成模型而检索多模态知识的各类方法。这些研究工作聚焦于从外部来源(包括图像、代码、表格、图谱和音频)中检索接地上下文。随着多模态学习与生成式人工智能的影响力日益提升,此类检索增强技术为事实性、推理、可解释性及鲁棒性等关键问题提供了有前景的解决方案。我们深入剖析了不同模态下的检索增强生成方法,并探讨了可能的未来研究方向。鉴于该领域尚处于新兴阶段,我们将持续补充新的论文与方法。