We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMs' ability to understand, remember, and reason about the camera wearer. MyEgo comprises 541 long videos and 5K personalized questions asking about "my things", "my activities", and "my past". Benchmarking reveals that competitive MLLMs across variants, including open-source vs. proprietary, thinking vs. non-thinking, small vs. large scales all struggle on MyEgo. Top closed- and open-source models (e.g., GPT-5 and Qwen3-VL) achieve only~46% and 36% accuracy, trailing human performance by near 40% and 50% respectively. Surprisingly, neither explicit reasoning nor model scaling yield consistent improvements. Models improve when relevant evidence is explicitly provided, but gains drop over time, indicating limitations in tracking and remembering "me" and "my past". These findings collectively highlight the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos. We hope MyEgo and our analyses catalyze further progress in these areas for egocentric personalized assistance. Data and code are available at https://github.com/Ryougetsu3606/MyEgo
翻译:我们首次系统分析了多模态大语言模型在需要“自我定位”——即理解自我中心视频中佩戴摄像头者的能力——的个性化问答中的表现。为此,我们提出了MyEgo,这是首个用于评估多模态大语言模型理解、记忆和推理关于摄像头佩戴者能力的自我中心视频问答数据集。MyEgo包含541个长视频和5,000个关于“我的物品”、“我的活动”和“我的过去”的个性化问题。基准测试显示,不同变体的多模态大语言模型(包括开源与专有模型、具备与不具备推理能力的模型、小规模与大规模模型)在MyEgo上均表现不佳。顶级闭源和开源模型(如GPT-5和Qwen3-VL)仅达到约46%和36%的准确率,分别落后人类表现近40%和50%。令人惊讶的是,显式推理和模型扩展均未带来一致的改进。当显式提供相关证据时,模型表现有所提升,但收益随时间下降,表明模型在追踪和记忆“我”与“我的过去”方面存在局限性。这些发现共同凸显了自我定位和长程记忆在实现自我中心视频个性化问答中的关键作用。我们希望MyEgo及我们的分析能推动这些领域在自我中心个性化辅助方面的进一步进展。数据和代码可在https://github.com/Ryougetsu3606/MyEgo获取。