Embedding based retrieval has seen its usage in a variety of search applications like e-commerce, social networking search etc. While the approach has demonstrated its efficacy in tasks like semantic matching and contextual search, it is plagued by the problem of uncontrollable relevance. In this paper, we conduct an analysis of embedding-based retrieval launched in early 2021 on our social network search engine, and define two main categories of failures introduced by it, integrity and junkiness. The former refers to issues such as hate speech and offensive content that can severely harm user experience, while the latter includes irrelevant results like fuzzy text matching or language mismatches. Efficient methods during model inference are further proposed to resolve the issue, including indexing treatments and targeted user cohort treatments, etc. Though being simple, we show the methods have good offline NDCG and online A/B tests metrics gain in practice. We analyze the reasons for the improvements, pointing out that our methods are only preliminary attempts to this important but challenging problem. We put forward potential future directions to explore.
翻译:基于嵌入的检索已广泛应用于电子商务、社交网络搜索等多种搜索场景。尽管该方法在语义匹配和上下文搜索等任务中展现出有效性,但仍受困于相关性不可控的问题。本文对2021年初在社交网络搜索引擎中部署的嵌入检索进行实证分析,定义了两类主要失败模式:完整性与杂乱性。前者指严重损害用户体验的仇恨言论和攻击性内容等问题,后者则包含模糊文本匹配或语言错配等不相关结果。我们进一步提出模型推理阶段的处理方法,包括索引优化和定向用户群体干预等。尽管方法简洁,但实验表明其在离线NDCG指标和在线A/B测试中均获得显著增益。通过分析改进原因,我们指出这些方法仅是针对这一重要且具有挑战性问题的初步尝试,并提出了未来值得探索的方向。