Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose {\bf{AnyRef}}, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation.
翻译:多模态大语言模型(MLLMs)利用大语言模型作为认知框架,处理各类视觉-语言任务。近期研究致力于赋予MLLMs视觉感知与定位能力,但在提供细粒度像素级感知以及扩展文本特定输入之外的交互方面仍存在差距。本文提出{\bf{AnyRef}}通用MLLM模型,可基于文本、边界框、图像或音频等多模态参考,生成像素级物体感知与自然语言描述。该创新无需特定模态设计,使用户能够以超越文本和区域提示的更高灵活性参与模型交互。通过提出的重聚焦机制,生成的定位输出被引导至更精准地聚焦参考物体,隐含地引入额外像素级监督。这种简单改进利用了LLM推理过程中产生的注意力分数,无需额外计算即可在定位掩码和指代表达方面展现性能提升。仅使用公开训练数据,本模型在多项基准测试中达到领先水平,涵盖多模态指代分割与区域级指代表达式生成任务。