A Transformer Model for Boundary Detection in Continuous Sign Language

Sign Language Recognition (SLR) has garnered significant attention from researchers in recent years, particularly the intricate domain of Continuous Sign Language Recognition (CSLR), which presents heightened complexity compared to Isolated Sign Language Recognition (ISLR). One of the prominent challenges in CSLR pertains to accurately detecting the boundaries of isolated signs within a continuous video stream. Additionally, the reliance on handcrafted features in existing models poses a challenge to achieving optimal accuracy. To surmount these challenges, we propose a novel approach utilizing a Transformer-based model. Unlike traditional models, our approach focuses on enhancing accuracy while eliminating the need for handcrafted features. The Transformer model is employed for both ISLR and CSLR. The training process involves using isolated sign videos, where hand keypoint features extracted from the input video are enriched using the Transformer model. Subsequently, these enriched features are forwarded to the final classification layer. The trained model, coupled with a post-processing method, is then applied to detect isolated sign boundaries within continuous sign videos. The evaluation of our model is conducted on two distinct datasets, including both continuous signs and their corresponding isolated signs, demonstrates promising results.

翻译：近年来，手语识别（SLR）研究受到广泛关注，其中连续手语识别（CSLR）因其相较于孤立手语识别（ISLR）更高的复杂度成为关键研究领域。CSLR面临的主要挑战之一是如何在连续视频流中准确检测孤立手语的边界。此外，现有模型对人工设计特征的依赖也制约了识别精度的提升。为解决上述问题，我们提出了一种基于Transformer模型的新型方法。与传统模型不同，本方法在提升精度的同时消除了对人工设计特征的依赖。该Transformer模型同时应用于ISLR与CSLR任务。训练阶段采用孤立手语视频，通过Transformer模型对输入视频提取的手部关键点特征进行增强，随后将增强特征送入最终分类层。训练后的模型结合后处理方法，可有效检测连续手语视频中孤立手语的边界。在两个包含连续手语及其对应孤立手语的数据集上的评估结果表明，本方法取得了显著成效。

相关内容

Continuity

关注 4

让 iOS 8 和 OS X Yosemite 无缝切换的一个新特性。 > Apple products have always been designed to work together beautifully. But now they may really surprise you. With iOS 8 and OS X Yosemite, you’ll be able to do more wonderful things than ever before.

Source: Apple - iOS 8

O’Reilly报告：知识图谱崛起——面向现代数据集成和数据结构体系，“The Rise of the Knowledge Graph——Toward Modern Data Integration and the Data Fabric Architecture”

专知会员服务

49+阅读 · 2022年2月18日

【WSDM2020】超越统计关系：将知识关系整合到多标签音乐风格分类的风格关联中（附pdf）

专知会员服务

18+阅读 · 2019年11月23日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日