Typographical errors are a major source of frustration for visitors of online marketplaces. Because of the domain-specific nature of these marketplaces and the very short queries users tend to search for, traditional spell cheking solutions do not perform well in correcting typos. We present a data augmentation method to address the lack of annotated typo data and train a recurrent neural network to learn context-limited domain-specific embeddings. Those embeddings are deployed in a real-time inferencing API for the Microsoft AppSource marketplace to find the closest match between a misspelled user query and the available product names. Our data efficient solution shows that controlled high quality synthetic data may be a powerful tool especially considering the current climate of large language models which rely on prohibitively huge and often uncontrolled datasets.
翻译:拼写错误是在线市场访客产生挫败感的主要来源。由于这些市场具有领域特异性,且用户搜索的查询通常非常简短,传统的拼写检查解决方案在纠正错别字方面表现不佳。我们提出了一种数据增强方法,以解决注释错字数据不足的问题,并训练循环神经网络学习受上下文约束的领域特异性嵌入。这些嵌入被部署到微软AppSource市场的实时推理API中,用于在拼写错误的用户查询与可用的产品名称之间找到最接近的匹配项。我们的数据高效解决方案表明,受控的高质量合成数据可能是一种强大的工具,尤其是在当前依赖规模庞大且通常不受控制的数据集的大型语言模型背景下。