Multimodal Named Entity Recognition (MNER) on social media aims to enhance textual entity prediction by incorporating image-based clues. Existing studies mainly focus on maximizing the utilization of pertinent image information or incorporating external knowledge from explicit knowledge bases. However, these methods either neglect the necessity of providing the model with external knowledge, or encounter issues of high redundancy in the retrieved knowledge. In this paper, we present PGIM -- a two-stage framework that aims to leverage ChatGPT as an implicit knowledge base and enable it to heuristically generate auxiliary knowledge for more efficient entity prediction. Specifically, PGIM contains a Multimodal Similar Example Awareness module that selects suitable examples from a small number of predefined artificial samples. These examples are then integrated into a formatted prompt template tailored to the MNER and guide ChatGPT to generate auxiliary refined knowledge. Finally, the acquired knowledge is integrated with the original text and fed into a downstream model for further processing. Extensive experiments show that PGIM outperforms state-of-the-art methods on two classic MNER datasets and exhibits a stronger robustness and generalization capability.
翻译:社交媒体上的多模态命名实体识别(MNER)旨在通过结合图像线索来增强文本实体预测。现有研究主要聚焦于最大化利用相关图像信息,或从显式知识库中引入外部知识。然而,这些方法要么忽略了向模型提供外部知识的必要性,要么在检索的知识中面临高冗余问题。本文提出PGIM——一个两阶段框架,旨在利用ChatGPT作为隐式知识库,使其启发式生成辅助知识以实现更高效的实体预测。具体而言,PGIM包含一个多模态相似样例感知模块,用于从少量预定义人工样本中选择合适样例。这些样例随后被集成到专为MNER设计的格式化提示模板中,引导ChatGPT生成辅助精炼知识。最终,获取的知识与原始文本融合,输入下游模型进行进一步处理。大量实验表明,PGIM在两个经典MNER数据集上优于现有最先进方法,并展现出更强的鲁棒性和泛化能力。