The recent advancement of Artificial Intelligence Generated Content (AIGC) has led to significant strides in modeling human interaction, particularly in the context of multimodal dialogue. While current methods impressively generate realistic dialogue in isolated modalities like speech or vision, challenges remain in controllable Multimodal Dialogue Generation (MDG). This paper focuses on the natural alignment between speech, vision, and text in human interaction, aiming for expressive dialogue generation through multimodal conditional control. To address the insufficient richness and diversity of dialogue expressiveness in existing datasets, we introduce a novel multimodal dialogue annotation pipeline to curate dialogues from movies and TV series with fine-grained annotations in interactional characteristics. The resulting MM-Dia dataset (360+ hours, 54,700 dialogues) facilitates explicitly controlled MDG, specifically through style-controllable dialogue speech synthesis. In parallel, MM-Dia-Bench (309 highly expressive dialogues with visible single-/dual-speaker scenes) serves as a rigorous testbed for implicit cross-modal MDG control, evaluating audio-visual style consistency across modalities. Extensive experiments demonstrate that training on MM-Dia significantly enhances fine-grained controllability, while evaluations on MM-Dia-Bench reveal limitations in current frameworks to replicate the nuanced expressiveness of human interaction. These findings provides new insights and challenges for multimodal conditional dialogue generation.


翻译:人工智能生成内容(AIGC)的最新进展在模拟人类交互方面取得了显著进步,尤其是在多模态对话的语境中。尽管当前方法在语音或视觉等孤立模态中能够令人印象深刻地生成逼真的对话,但在可控多模态对话生成(MDG)方面仍存在挑战。本文聚焦于人类交互中语音、视觉与文本之间的自然对齐,旨在通过多模态条件控制实现富有表现力的对话生成。为解决现有数据集中对话表现力丰富性和多样性的不足,我们提出了一种新颖的多模态对话标注流水线,从电影和电视剧中筛选出具有细粒度交互特征标注的对话。由此产生的MM-Dia数据集(360+小时,54,700个对话)促进了显式可控的MDG,特别是通过风格可控的对话语音合成。与此同时,MM-Dia-Bench(309个具有高表现力的对话,包含可见的单/双说话人场景)作为隐式跨模态MDG控制的严格测试平台,评估了跨模态的音视频风格一致性。大量实验表明,在MM-Dia上训练显著增强了细粒度可控性,而对MM-Dia-Bench的评估则揭示了当前框架在复现人类交互细腻表现力方面的局限性。这些发现为多模态条件对话生成提供了新的见解和挑战。

0
下载
关闭预览

相关内容

《可控视频生成:综述》
专知会员服务
17+阅读 · 2025年7月24日
多模态基础模型的机制可解释性综述
专知会员服务
43+阅读 · 2025年2月28日
迈向可解释和可理解的多模态大规模语言模型
专知会员服务
41+阅读 · 2024年12月7日
多模态可控扩散模型综述
专知会员服务
39+阅读 · 2024年7月20日
专知会员服务
65+阅读 · 2021年5月29日
数据受限条件下的多模态处理技术综述
专知
22+阅读 · 2022年7月16日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Arxiv
12+阅读 · 2023年5月22日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
9+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
8+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
5+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
10+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
12+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
6+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
12+阅读 · 9月21日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员