The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human interaction in real-world settings. This task involves describing complex visual environments, identifying active objects, and interpreting their changes as conveyed through language. Traditional methods, which isolate object captioning and state change detection, offer a limited view of dynamic environments. Moreover, relying on a small set of symbolic words to represent changes has restricted the expressiveness of the language. To address these challenges, in this paper, we introduce the Object State Captioning and State Change Representation (OSCaR) dataset and benchmark. OSCaR consists of 14,084 annotated video segments with nearly 1,000 unique objects from various egocentric video collections. It sets a new testbed for evaluating multimodal large language models (MLLMs). Our experiments demonstrate that while MLLMs show some skill, they lack a full understanding of object state changes. The benchmark includes a fine-tuned model that, despite initial capabilities, requires significant improvements in accuracy and generalization ability for effective understanding of these changes. Our code and dataset are available at https://github.com/nguyennm1024/OSCaR.
翻译:智能模型推断和理解物体状态变化的能力是人工智能研究中一个关键但具有挑战性的方面,尤其是在现实世界中通过人类交互的视角进行考察。这一任务涉及描述复杂的视觉环境、识别主动物体,并解释其通过语言传达的变化。传统方法将物体描述和状态变化检测隔离开来,只能提供对动态环境的有限视角。此外,依赖少量符号化词汇来表示变化限制了语言的表达力。为解决这些挑战,本文提出了物体状态描述与状态变化表示(OSCaR)数据集和基准测试。OSCaR包含来自多种自我中心视频集合的14,084个带注释的视频片段,涉及近1,000个独特物体。它为评估多模态大语言模型(MLLMs)设立了一个新的测试平台。我们的实验表明,尽管MLLMs展现了一定的能力,但它们仍缺乏对物体状态变化的全面理解。该基准测试包括一个经过微调的模型,尽管具备初步能力,但在准确性、泛化能力以及有效理解这些变化方面仍需显著改进。我们的代码和数据集可在 https://github.com/nguyennm1024/OSCaR 获取。