Human-centered dynamic scene understanding plays a pivotal role in enhancing the capability of robotic and autonomous systems, in which Video-based Human-Object Interaction (V-HOI) detection is a crucial task in semantic scene understanding, aimed at comprehensively understanding HOI relationships within a video to benefit the behavioral decisions of mobile robots and autonomous driving systems. Although previous V-HOI detection models have made significant strides in accurate detection on specific datasets, they still lack the general reasoning ability like human beings to effectively induce HOI relationships. In this study, we propose V-HOI Multi-LLMs Collaborated Reasoning (V-HOI MLCR), a novel framework consisting of a series of plug-and-play modules that could facilitate the performance of current V-HOI detection models by leveraging the strong reasoning ability of different off-the-shelf pre-trained large language models (LLMs). We design a two-stage collaboration system of different LLMs for the V-HOI task. Specifically, in the first stage, we design a Cross-Agents Reasoning scheme to leverage the LLM conduct reasoning from different aspects. In the second stage, we perform Multi-LLMs Debate to get the final reasoning answer based on the different knowledge in different LLMs. Additionally, we devise an auxiliary training strategy that utilizes CLIP, a large vision-language model to enhance the base V-HOI models' discriminative ability to better cooperate with LLMs. We validate the superiority of our design by demonstrating its effectiveness in improving the prediction accuracy of the base V-HOI model via reasoning from multiple perspectives.
翻译:以人为中心的动态场景理解对于提升机器人与自主系统的能力具有关键作用,其中基于视频的人-物交互检测是语义场景理解中的核心任务,旨在全面理解视频中的人-物交互关系,以服务于移动机器人与自动驾驶系统的行为决策。尽管现有的V-HOI检测模型在特定数据集上已取得显著进展,但其仍缺乏类人的通用推理能力以有效推导人-物交互关系。本研究提出V-HOI多LLM协同推理框架,该新型框架由一系列即插即用模块构成,能够通过利用多种现成预训练大语言模型的强大推理能力,提升现有V-HOI检测模型的性能。我们为V-HOI任务设计了一个两阶段的多LLM协同系统:第一阶段设计跨智能体推理机制,驱动LLM从不同维度进行推理;第二阶段执行多LLM辩论,基于不同LLM的异构知识生成最终推理结果。此外,我们引入辅助训练策略,利用大规模视觉语言模型CLIP增强基础V-HOI模型的判别能力,以优化其与LLM的协作效能。通过多视角推理显著提升基础V-HOI模型预测准确率的实验,验证了本设计方案的优越性。