In psychoanalysis, generating interpretations to one's psychological state through visual creations is facing significant demands. The two main tasks of existing studies in the field of computer vision, sentiment/emotion classification and affective captioning, can hardly satisfy the requirement of psychological interpreting. To meet the demands for psychoanalysis, we introduce a challenging task, \textbf{V}isual \textbf{E}motion \textbf{I}nterpretation \textbf{T}ask (VEIT). VEIT requires AI to generate reasonable interpretations of creator's psychological state through visual creations. To support the task, we present a multimodal dataset termed SpyIn (\textbf{S}and\textbf{p}la\textbf{y} \textbf{In}terpretation Dataset), which is psychological theory supported and professional annotated. Dataset analysis illustrates that SpyIn is not only able to support VEIT, but also more challenging compared with other captioning datasets. Building on SpyIn, we conduct experiments of several image captioning method, and propose a visual-semantic combined model which obtains a SOTA result on SpyIn. The results indicate that VEIT is a more challenging task requiring scene graph information and psychological knowledge. Our work also show a promise for AI to analyze and explain inner world of humanity through visual creations.
翻译:在精神分析中,通过视觉创作生成对个体心理状态的解读存在显著需求。现有计算机视觉领域的情绪/情感分类与情感描述两项主要任务,难以满足心理解读的要求。为契合精神分析的需求,我们提出一项具有挑战性的任务——视觉情绪解读任务(Visual Emotion Interpretation Task,VEIT)。VEIT要求人工智能通过视觉创作生成对创作者心理状态的合理解读。为支持该任务,我们构建了名为SpyIn(沙盘解读数据集)的多模态数据集,该数据集基于心理学理论支撑并经专业标注。数据分析表明,SpyIn不仅能够支撑VEIT任务,且相较于其他描述数据集更具挑战性。基于SpyIn,我们开展了多项图像描述方法的实验,并提出一种视觉-语义联合模型,在SpyIn上取得了最优结果。实验结果表明,VEIT是一项需要场景图信息与心理学知识的更具挑战性的任务。我们的工作也展示了人工智能通过视觉作品分析与解释人类内心世界的应用前景。