Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descriptions, worldviews, and human social commonsense to put disjointed elements together to form a coherent and engaging human-writeable story. However, most models mainly focus on applying factual information and using taxonomic/lexical external knowledge when attempting to create stories. This paper introduces SCO-VIST, a framework representing the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge. SCO-VIST then takes this graph representing plot points and creates bridges between plot points with semantic and occurrence-based edge weights. This weighted story graph produces the storyline in a sequence of events using Floyd-Warshall's algorithm. Our proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations.
翻译:视觉故事生成旨在根据给定的图像序列自动生成连贯的故事。与图像描述等任务不同,视觉故事应包含事实描述、世界观和人类社交常识,以将零散的元素整合为连贯且引人入胜、可媲美人类创作的故事。然而,现有模型大多仅关注应用事实信息及使用分类/词汇层面的外部知识来构建故事。本文提出SCO-VIST框架,该框架将图像序列表示为包含物体及其关系的图结构,并融入人类行为动机及其社会交互常识知识。随后SCO-VIST将表征情节节点的图,通过基于语义和共现的边权重建立节点间的衔接桥接。该加权故事图经弗洛伊德-沃舍尔算法求解,生成事件序列形式的故事主线。自动评估与人工评估均表明,本框架在视觉基础性、连贯性、多样性和人性化程度等多维度指标上均优于现有方法。