This is the first year of the TREC Product search track. The focus this year was the creation of a reusable collection and evaluation of the impact of the use of metadata and multi-modal data on retrieval accuracy. This year we leverage the new product search corpus, which includes contextual metadata. Our analysis shows that in the product search domain, traditional retrieval systems are highly effective and commonly outperform general-purpose pretrained embedding models. Our analysis also evaluates the impact of using simplified and metadata-enhanced collections, finding no clear trend in the impact of the expanded collection. We also see some surprising outcomes; despite their widespread adoption and competitive performance on other tasks, we find single-stage dense retrieval runs can commonly be noncompetitive or generate low-quality results both in the zero-shot and fine-tuned domain.
翻译:这是TREC产品搜索赛道的首年。本年度重点关注可复用集合的构建,以及元数据和多模态数据的使用对检索准确性的影响评估。我们今年利用了包含上下文元数据的新产品搜索语料库。分析表明,在产品搜索领域,传统检索系统表现优异,且通常优于通用预训练嵌入模型。我们的分析还评估了使用简化集合与元数据增强集合的影响,发现扩展集合对检索效果的影响未呈现明确趋势。同时观察到一些令人意外的结果:尽管单阶段稠密检索方法已被广泛采用,并在其他任务中展现出竞争力,但在零样本和微调场景下,其运行结果通常缺乏竞争力或生成低质量检索结果。