In the multimedia era, image is an effective medium in search advertising. Dynamic Image Advertising (DIA), a system that matches queries with ad images and generates multimodal ads, is introduced to improve user experience and ad revenue. The core of DIA is a query-image matching module performing ad image retrieval and relevance modeling. Current query-image matching suffers from limited and inconsistent data, and insufficient cross-modal interaction. Also, the separate optimization of retrieval and relevance models affects overall performance. To address this issue, we propose a vision-language framework consisting of two parts. First, we train a base model on large-scale image-text pairs to learn general multimodal representation. Then, we fine-tune the base model on advertising business data, unifying relevance modeling and retrieval through multi-objective learning. Our framework has been implemented in Baidu search advertising system "Phoneix Nest". Online evaluation shows that it improves cost per mille (CPM) and click-through rate (CTR) by 1.04% and 1.865%.
翻译:在多媒体时代,图像已成为搜索广告中的有效媒介。为提升用户体验和广告收入,动态图像广告(DIA)系统应运而生,该系统通过匹配查询与广告图像并生成多模态广告。DIA的核心是查询-图像匹配模块,负责执行广告图像检索与相关性建模。当前查询-图像匹配面临数据有限且不一致、跨模态交互不足等问题,同时检索与相关性模型的独立优化也影响了整体性能。为解决上述问题,我们提出了一种由两部分组成的语言-视觉框架:首先基于大规模图像-文本对训练基础模型以学习通用多模态表征,随后在广告业务数据上微调基础模型,通过多目标学习统一相关性建模与检索任务。该框架已在百度搜索广告系统"凤巢"中落地实施。在线评估表明,该方法使千次展示成本(CPM)和点击率(CTR)分别提升1.04%和1.865%。