We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1) Dialog-based Question Answering (2) Dialog-based Emotion Prediction and (3) Affective emotion explanation generation based on the dialog. Our key contribution is the collection of a large-scale dataset, dubbed AffectVisDial, consisting of 50K 10-turn visually grounded dialogs as well as concluding emotion attributions and dialog-informed textual emotion explanations, resulting in a total of 27,180 working hours. We explain our design decisions in collecting the dataset and introduce the questioner and answerer tasks that are associated with the participants in the conversation. We train and demonstrate solid Affective Visual Dialog baselines adapted from state-of-the-art models. Remarkably, the responses generated by our models show promising emotional reasoning abilities in response to visually grounded conversations. Our project page is available at https://affective-visual-dialog.github.io.
翻译:我们提出情感性视觉对话(Affective Visual Dialog)这一情感解释与推理任务,作为研究视觉情境对话中情感形成机制的平台。该任务涉及三项技能:(1)基于对话的问答;(2)基于对话的情感预测;(3)基于对话的情感解释生成。我们的核心贡献是构建了一个大规模数据集,命名为AffectVisDial,包含5万组10轮视觉情境对话,以及最终的情感归因和基于对话的文本情感解释,总计耗费27,180人时。我们阐释了数据集构建中的设计决策,并介绍了与对话参与者相关的提问者与回答者任务。我们训练并展示了基于最先进模型改编的稳定情感性视觉对话基线方法。值得注意的是,我们的模型生成的回答在响应视觉情境对话时展现出具有前景的情感推理能力。项目页面可访问https://affective-visual-dialog.github.io获取。