Sign language video retrieval plays a key role in facilitating information access for the deaf community. Despite significant advances in video-text retrieval, the complexity and inherent uncertainty of sign language preclude the direct application of these techniques. Previous methods achieve the mapping between sign language video and text through fine-grained modal alignment. However, due to the scarcity of fine-grained annotation, the uncertainty inherent in sign language video is underestimated, limiting the further development of sign language retrieval tasks. To address this challenge, we propose a novel Uncertainty-aware Probability Distribution Retrieval (UPRet), that conceptualizes the mapping process of sign language video and text in terms of probability distributions, explores their potential interrelationships, and enables flexible mappings. Experiments on three benchmarks demonstrate the effectiveness of our method, which achieves state-of-the-art results on How2Sign (59.1%), PHOENIX-2014T (72.0%), and CSL-Daily (78.4%).
翻译:手语视频检索在促进听障群体信息获取方面发挥着关键作用。尽管视频-文本检索领域已取得显著进展,但手语本身的复杂性及内在不确定性阻碍了这些技术的直接应用。现有方法通过细粒度模态对齐实现手语视频与文本的映射,然而由于细粒度标注的稀缺性,手语视频固有的不确定性被低估,限制了手语检索任务的进一步发展。为应对这一挑战,我们提出了一种新颖的不确定性感知概率分布检索方法(UPRet),该方法将手语视频与文本的映射过程概念化为概率分布,探索其潜在关联性,并实现灵活映射。在三个基准数据集上的实验证明了我们方法的有效性,在How2Sign(59.1%)、PHOENIX-2014T(72.0%)和CSL-Daily(78.4%)上均取得了最先进的结果。