Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the \textbf{M}ulti-\textbf{F}aceted \textbf{R}esponsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression.
翻译:面对面交流是包含说话者与倾听者角色的常见场景。现有大多数研究方法集中于生成说话者视频,而倾听者头部的生成仍被广泛忽视。响应式倾听者头部生成是一项重要任务,旨在通过给定说话者视频和倾听者头部图像,模拟面对面交流场景并生成倾听者头部视频。理想的生成式响应倾听视频应能在保持交互模式多样性和倾听者身份信息准确性的同时,对说话者表达态度或观点。为实现此目标,我们提出了**多维响应式倾听者头部生成网络**(MFR-Net)。具体而言,MFR-Net采用概率去噪扩散模型预测多样化的头部姿态和表情特征。为了对说话者视频进行多维响应,同时保持倾听者身份的精确性,我们设计了特征聚合模块,用于增强倾听者身份特征并将其与其他与说话者相关的特征融合。最后,通过微调身份一致性损失的渲染器生成最终的倾听者头部视频。大量实验表明,MFR-Net不仅在多样性和说话者身份信息方面实现了多维响应,在态度和观点表达方面也表现出色。