Estimating Emotional Mimicry Intensity (EMI) in naturalistic environments is a critical yet challenging task in affective computing. The primary difficulty lies in effectively modeling the complex, nonlinear temporal dynamics across highly heterogeneous modalities, especially when physical signals are corrupted or missing. To tackle this, we propose TAEMI (Text-Anchored Emotional Mimicry Intensity estimation), a novel multimodal framework designed for the 10th ABAW Competition. Motivated by the observation that continuous visual and acoustic signals are highly susceptible to transient environmental noise, we break the traditional symmetric fusion paradigm. Instead, we leverage textual transcript--which inherently encode a stable, time-independent semantic prior--as central anchors. Specifically, we introduce a Text-Anchored Dual Cross-Attention mechanism that utilizes these robust textual queries to actively filter out frame-level redundancies and align the noisy physical streams. Furthermore, to prevent catastrophic performance degradation caused by inevitably missing data in unconstrained real-world scenarios, we integrate Learnable Missing-Modality Tokens and a Modality Dropout strategy during training. Extensive experiments on the Hume-Vidmimic2 dataset demonstrate that TAEMI effectively captures fine-grained emotional variations and maintains robust predictive resilience under imperfect conditions. Our framework achieves a state-of-the-art mean Pearson correlation coefficient across six continuous emotional dimensions, significantly outperforming existing baseline methods.
翻译:在自然环境中估计情感模仿强度(EMI)是情感计算中一项关键而具挑战性的任务。主要困难在于有效建模高度异质模态间复杂的非线性时序动态,尤其是当物理信号受损或缺失时。为此,我们提出TAEMI(文本锚定的情感模仿强度估计),一种专为第10届ABAW竞赛设计的新型多模态框架。受连续视觉与声学信号极易受瞬时环境噪声影响的观察启发,我们打破了传统的对称融合范式,转而利用文本转录——其天然编码了稳定的、与时间无关的语义先验——作为核心锚点。具体而言,我们引入了文本锚定双重交叉注意力机制,利用这些鲁棒的文本查询主动过滤帧级冗余并对齐含噪的物理流。此外,为防止在无约束真实场景中因数据必然缺失而导致的灾难性性能下降,我们在训练过程中集成了可学习缺失模态张量及模态丢弃策略。在Hume-Vidmimic2数据集上的大量实验表明,TAEMI能有效捕捉细粒情感波动,并在不完美条件下保持稳健的预测弹性。我们的框架在六个连续情感维度上达到了平均皮尔逊相关系数的最优水平,显著超越了现有基线方法。