Large vision-language models (LVLMs), while proficient in following instructions and responding to diverse questions, invariably generate detailed responses even when questions are ambiguous or unanswerable, leading to hallucinations and bias issues. Thus, it is essential for LVLMs to proactively engage with humans to ask for clarifications or additional information for better responses. In this study, we aim to shift LVLMs from passive answer providers to proactive engaged partners. We begin by establishing a three-tiered hierarchy for questions of invalid, ambiguous, and personalizable nature to measure the proactive engagement capabilities of LVLMs. Utilizing this hierarchy, we create PIE, (ProactIve Engagement Evaluation) through GPT-4o and human annotators, consisting of 853 questions across six distinct, fine-grained question types that are verified by human annotators and accompanied with well-defined metrics. Our evaluations on \benchmark indicate poor performance of existing LVLMs, with the best-performing open-weights model only achieving an Aggregate Align Rate (AAR) of 0.28. In response, we introduce MACAROON, self-iMaginAtion for ContrAstive pReference OptimizatiON, which instructs LVLMs to autonomously generate contrastive response pairs for unlabeled questions given the task description and human-crafted criteria. Then, the self-imagined data is formatted for conditional reinforcement learning. Experimental results show MACAROON effectively improves LVLMs' capabilities to be proactively engaged (0.84 AAR) while maintaining comparable performance on general tasks.
翻译:大型视觉-语言模型(LVLMs)虽然擅长遵循指令并回答各类问题,但即使在问题模糊或无法回答时,也总会生成详细回应,从而导致幻觉和偏见问题。因此,LVLMs有必要主动与人类互动,以请求澄清或获取额外信息,从而给出更佳回答。本研究旨在将LVLMs从被动的答案提供者转变为主动的积极参与伙伴。我们首先针对无效、模糊和可个性化三类问题建立了一个三层级分类体系,用以衡量LVLMs的主动参与能力。基于此体系,我们利用GPT-4o和人工标注者构建了PIE(主动参与评估)数据集,该数据集包含853个问题,涵盖六种精细划分的问题类型,所有问题均经人工标注者验证并配有明确定义的评估指标。在\benchmark上的评估表明,现有LVLMs表现欠佳,性能最佳的开源权重模型仅获得0.28的总体对齐率(AAR)。为此,我们提出了MACAROON(基于自想象对比偏好优化的方法),该方法指导LVLMs根据任务描述和人工制定的标准,针对未标注问题自主生成对比性回答对。随后,将自想象数据格式化以用于条件强化学习。实验结果表明,MACAROON能有效提升LVLMs的主动参与能力(AAR达0.84),同时在通用任务上保持相当的性能水平。