Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content. However, their performance in sentiment analysis exhibits acute sensitivity to prompt design, rendering static, uniformly applied prompts inherently suboptimal for capturing the nuanced multimodal cues that vary across inputs. To address this limitation, we propose a Multimodal Adaptive Few-Shot Prompting (MAF) framework, which dynamically retrieves and integrates query-relevant demonstrations to elicit the sentiment reasoning capabilities of MLLMs in a context-sensitive manner. MAF constructs a demonstration retrieval module that holistically encodes facial expressions, scene context, and textual semantics, with a lip movement amplitude detection mechanism introduced for accurate speaker identification in multi-person scenarios. Departing from conventional fixed-weight fusion, a lightweight coefficient generation network is trained to output query-conditioned fusion weights in real time, enabling weighted aggregation of multimodal similarity scores to retrieve the top-K most informative demonstrations. Prediction stability is further enhanced through majority voting over multiple candidate outputs generated by the MLLM. Extensive experiments on public benchmark datasets demonstrate that MAF achieves substantial and consistent performance improvements over the corresponding backbone variants and remains competitive with strong multimodal sentiment-analysis baselines.
翻译:多模态大语言模型(MLLMs)在理解复杂多模态内容方面展现出卓越能力。然而,它们在情感分析中的性能对提示设计高度敏感,导致静态且统一应用的提示 inherently 无法充分捕捉跨输入变化的细微多模态线索。为解决这一局限,我们提出多模态自适应少样本提示(MAF)框架,该框架动态检索并整合与查询相关的示例,以上下文敏感的方式激发MLLMs的情感推理能力。MAF构建了示例检索模块,该模块整体编码面部表情、场景上下文和文本语义,并引入唇动幅度检测机制以实现多人场景中说话者的精准识别。与传统的固定权重融合不同,我们训练了一个轻量级系数生成网络,以实时输出查询条件的融合权重,从而实现多模态相似度分数的加权聚合,以检索最具信息量的前K个示例。通过多轮对MLLM生成的候选输出进行多数投票,进一步增强了预测稳定性。在公开基准数据集上的大量实验表明,MAF相较于对应的骨干变体实现了显著且一致的性能提升,并与强多模态情感分析基线保持竞争力。