Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, models must deduce hidden viewer emotions from explicit visual narratives, concrete character-object interactions, and visible textual cues. However, standard Multimodal Large Language Models (MLLMs) typically rely on holistic frame representations, which leave these fine-grained, affect-relevant events implicit and complicate precise emotional reasoning. To address this, we propose a grounded action-centric evidence augmentation framework that enhances video MLLMs' clue extraction and comprehension by introducing explicit event structure and localized visual evidence. Our method extracts temporally ordered subject-verb-object (SVO) triplets and auxiliary visible textual cues from action-centric video descriptions, grounds subject and object entities as visual entity crops, and then enables the MLLM to perform clue-enhanced emotional reasoning based on these extracted structured clues. In this way, action triplets specify "what happens", while grounded visual entity crops anchor "who or what participates in each event" to concrete visual evidence. Experiments on the Pitts dataset show consistent improvements over Qwen2.5-VL and Qwen3-VL baselines. Ablation studies, cross-dataset evaluation on AdsQA, and transfer experiments on an emotion-focused TVQA subset further support the effectiveness and generalization of our approach.
翻译:视频广告中的观众情感预测旨在推断观众在观看时所产生的潜在情感反应。为弥合画面呈现与情感感知之间的差距,模型需从显性的视觉叙事、具体的人物-物体交互以及可见的文字线索中推导出隐藏的观众情感。然而,标准多模态大语言模型(MLLMs)通常依赖全局帧表征,使得这些细粒度的、与情感相关的事件信息隐含化,难以进行精确的情感推理。为解决这一问题,我们提出一种基于具身动作中心的证据增强框架,通过引入显式事件结构与局部化视觉证据,增强视频MLLMs的线索提取与理解能力。该方法从以动作为核心的视频描述中提取时序化的主谓宾三元组及辅助性可见文字线索,将主语与宾语实体锚定为视觉实体裁剪块,进而使MLLM基于这些提取的结构化线索进行增强型情感推理。在此范式下,动作三元组明确界定“发生了什么事”,而具身化视觉实体裁剪块则将“事件中各参与者是谁或什么”锚定到具体视觉证据上。在Pitts数据集上的实验表明,该方法相比Qwen2.5-VL与Qwen3-VL基线取得一致提升。消融研究、AdsQA跨数据集评估以及基于情感驱动TVQA子集的迁移实验进一步验证了本方法的有效性与泛化能力。