It is important for daily life support robots to detect changes in their environment and perform tasks. In the field of anomaly detection in computer vision, probabilistic and deep learning methods have been used to calculate the image distance. These methods calculate distances by focusing on image pixels. In contrast, this study aims to detect semantic changes in the daily life environment using the current development of large-scale vision-language models. Using its Visual Question Answering (VQA) model, we propose a method to detect semantic changes by applying multiple questions to a reference image and a current image and obtaining answers in the form of sentences. Unlike deep learning-based methods in anomaly detection, this method does not require any training or fine-tuning, is not affected by noise, and is sensitive to semantic state changes in the real world. In our experiments, we demonstrated the effectiveness of this method by applying it to a patrol task in a real-life environment using a mobile robot, Fetch Mobile Manipulator. In the future, it may be possible to add explanatory power to changes in the daily life environment through spoken language.
翻译:对于日常生活支持机器人而言,检测环境变化并执行任务至关重要。在计算机视觉的异常检测领域,概率与深度学习方法常被用于计算图像距离,这些方法通过关注图像像素来衡量差异。与此不同,本研究旨在利用当前发展的大规模视觉语言模型检测日常生活环境中的语义变化。我们提出一种基于视觉问答(VQA)模型的方法,通过向参考图像和当前图像提出多个问题并以句子形式获取答案,从而检测语义变化。与异常检测中基于深度学习的方法相比,该方法无需任何训练或微调,不受噪声影响,且对现实世界中的语义状态变化敏感。在实验中,我们使用移动机器人Fetch Mobile Manipulator在真实生活环境中执行巡逻任务,验证了该方法的有效性。未来,有可能通过口语为日常生活环境的变化增添解释能力。