We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMs' ability to understand, remember, and reason about the camera wearer. MyEgo comprises 541 long videos and 5K personalized questions asking about "my things", "my activities", and "my past". Benchmarking reveals that competitive MLLMs across variants, including open-source vs. proprietary, thinking vs. non-thinking, small vs. large scales all struggle on MyEgo. Top closed- and open-source models (e.g., GPT-5 and Qwen3-VL) achieve only~46% and 36% accuracy, trailing human performance by near 40% and 50% respectively. Surprisingly, neither explicit reasoning nor model scaling yield consistent improvements. Models improve when relevant evidence is explicitly provided, but gains drop over time, indicating limitations in tracking and remembering "me" and "my past". These findings collectively highlight the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos. We hope MyEgo and our analyses catalyze further progress in these areas for egocentric personalized assistance. Data and code are available at https://github.com/Ryougetsu3606/MyEgo


翻译:我们首次系统分析了多模态大语言模型在需要“自我定位”——即理解自我中心视频中佩戴摄像头者的能力——的个性化问答中的表现。为此,我们提出了MyEgo,这是首个用于评估多模态大语言模型理解、记忆和推理关于摄像头佩戴者能力的自我中心视频问答数据集。MyEgo包含541个长视频和5,000个关于“我的物品”、“我的活动”和“我的过去”的个性化问题。基准测试显示,不同变体的多模态大语言模型(包括开源与专有模型、具备与不具备推理能力的模型、小规模与大规模模型)在MyEgo上均表现不佳。顶级闭源和开源模型(如GPT-5和Qwen3-VL)仅达到约46%和36%的准确率,分别落后人类表现近40%和50%。令人惊讶的是,显式推理和模型扩展均未带来一致的改进。当显式提供相关证据时,模型表现有所提升,但收益随时间下降,表明模型在追踪和记忆“我”与“我的过去”方面存在局限性。这些发现共同凸显了自我定位和长程记忆在实现自我中心视频个性化问答中的关键作用。我们希望MyEgo及我们的分析能推动这些领域在自我中心个性化辅助方面的进一步进展。数据和代码可在https://github.com/Ryougetsu3606/MyEgo获取。

0
下载
关闭预览

相关内容

多模态大语言模型下游调优中“保持自我”的重要性
专知会员服务
17+阅读 · 2025年12月15日
视觉自回归模型综述
专知会员服务
45+阅读 · 2024年11月15日
【优青论文】视觉问答技术研究
计算机研究与发展
13+阅读 · 2018年9月21日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
面向国防作战的最佳自主与蜂群无人机技术
专知会员服务
1+阅读 · 12分钟前
《异构人类团队的协作决策过程混合建模研究》
专知会员服务
2+阅读 · 17分钟前
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
2+阅读 · 7月27日
美空军新型反无人机部队初探
专知会员服务
7+阅读 · 7月27日
《防空交战流程的概率建模研究》
专知会员服务
10+阅读 · 7月27日
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
6+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
9+阅读 · 7月26日
相关VIP内容
多模态大语言模型下游调优中“保持自我”的重要性
专知会员服务
17+阅读 · 2025年12月15日
视觉自回归模型综述
专知会员服务
45+阅读 · 2024年11月15日
相关资讯
【优青论文】视觉问答技术研究
计算机研究与发展
13+阅读 · 2018年9月21日
相关基金
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员