Well-formed context aware image captions and tags in enterprise content such as marketing material are critical to ensure their brand presence and content recall. Manual creation and updates to ensure the same is non trivial given the scale and the tedium towards this task. We propose a new unified Vision-Language (VL) model based on the One For All (OFA) model, with a focus on context-assisted image captioning where the caption is generated based on both the image and its context. Our approach aims to overcome the context-independent (image and text are treated independently) nature of the existing approaches. We exploit context by pretraining our model with datasets of three tasks: news image captioning where the news article is the context, contextual visual entailment, and keyword extraction from the context. The second pretraining task is a new VL task, and we construct and release two datasets for the task with 1.1M and 2.2K data instances. Our system achieves state-of-the-art results with an improvement of up to 8.34 CIDEr score on the benchmark news image captioning datasets. To the best of our knowledge, ours is the first effort at incorporating contextual information in pretraining the models for the VL tasks.
翻译:摘要:在企业内容(如营销材料)中,格式良好的上下文感知图像描述和标签对于确保品牌存在感和内容召回率至关重要。鉴于该任务的规模及繁琐性,人工创建和更新这些内容并非易事。我们提出了一种新的统一视觉-语言(VL)模型,基于"万法归一"(OFA)模型,专注于上下文辅助的图像描述任务——即根据图像及其上下文联合生成描述文本。我们的方法旨在克服现有方法中上下文无关(将图像和文本独立处理)的特性。我们通过三类任务的数据集对模型进行预训练以利用上下文:以新闻文章为上下文的新闻图像描述、上下文视觉蕴含,以及从上下文中提取关键词。其中第二个预训练任务是一个新的VL任务,我们为此任务构建并发布了两个数据集,分别包含110万和2200个数据实例。我们的系统在基准新闻图像描述数据集上取得了最先进的结果,CIDEr评分提升高达8.34分。据我们所知,这是首次在VL任务预训练模型中融入上下文信息的尝试。