In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that despite not necessarily describing the image may provide useful context that if properly leveraged can be informative. We address this task with a multimodal model that conditions on both textual information from the associated social media post as well as visual signal from the image, and demonstrate that the utility of these two information sources stacks. We put forward a new dataset of 371k images paired with alt-text and tweets scraped from Twitter and evaluate on it across a variety of automated metrics as well as human evaluation. We show that our approach of conditioning on both tweet text and visual information significantly outperforms prior work, by more than 2x on BLEU@4.
翻译:在本工作中,我们提出了一种为社交媒体(特别是Twitter)上分享的图像生成替代文本(alt-text)描述的方法。替代文本不仅仅是图像描述的一个特例,它既要求字面意义上的精确描述,又需结合具体语境。关键的是,Twitter上发布的图像通常伴随用户撰写的文字,这些文字虽未必直接描述图像,但可能提供有用的上下文信息,若恰当利用,将具有信息价值。我们通过一个多模态模型来应对这一任务,该模型条件依赖于社交媒体帖子中的文本信息以及图像的视觉信号,并证明了这两种信息源的效用具有叠加性。我们提出了一个新数据集,包含371k对从Twitter抓取的图像及其替代文本和推文,并基于多种自动化指标及人工评估进行了评测。结果表明,我们基于推文文本和视觉信息双重条件的处理方法显著优于先前工作,在BLEU@4指标上性能提升超过两倍。