TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to missing visual data, neglecting the role of head orientation, and background noise. This study addresses these limitations by introducing EgoAdapt, an adaptive framework designed for robust egocentric "Talking to Me" speaker detection under missing modalities. Specifically, EgoAdapt incorporates three key modules: (1) a Visual Speaker Target Recognition (VSTR) module that captures head orientation as a non-verbal cue and lip movement as a verbal cue, allowing a comprehensive interpretation of both verbal and non-verbal signals to address TTM, setting it apart from tasks focused solely on detecting speaking status; (2) a Parallel Shared-weight Audio (PSA) encoder for enhanced audio feature extraction in noisy environments; and (3) a Visual Modality Missing Awareness (VMMA) module that estimates the presence or absence of each modality at each frame to adjust the system response dynamically.Comprehensive evaluations on the TTM benchmark of the Ego4D dataset demonstrate that EgoAdapt achieves a mean Average Precision (mAP) of 67.39% and an Accuracy (Acc) of 62.01%, significantly outperforming the state-of-the-art method by 4.96% in Accuracy and 1.56% in mAP.
翻译:TTM(是否在与我交谈)任务是理解人类社会交互的核心组成部分,旨在判断谁在与摄像机佩戴者进行对话。传统模型在实际场景中常因视觉数据缺失、忽视头部朝向作用及背景噪声而面临挑战。本研究通过引入EgoAdapt自适应框架解决了上述局限,该框架专为缺失模态下鲁棒的自我中心“是否在与我交谈”说话人检测而设计。具体而言,EgoAdapt包含三个关键模块:(1)视觉说话人目标识别(VSTR)模块,通过捕捉头部朝向(非语言线索)和唇部运动(语言线索),实现对语言与非语言信号的综合解读以解决TTM问题,从而区别于仅聚焦于说话状态检测的任务;(2)并行共享权重音频(PSA)编码器,用于增强噪声环境下的音频特征提取;(3)视觉模态缺失感知(VMMA)模块,通过估计每帧中各模态的存在与否动态调整系统响应。在Ego4D数据集TTM基准上的综合评估表明,EgoAdapt实现了67.39%的平均精度(mAP)和62.01%的准确率(Acc),在准确率和mAP上分别以4.96%和1.56%的幅度显著超越现有最佳方法。