In this work, we are dedicated to leveraging the denoising diffusion models' success and formulating feature refinement as the autoencoder-formed diffusion process. The state-of-the-art CSLR framework consists of a spatial module, a visual module, a sequence module, and a sequence learning function. However, this framework has faced sequence module overfitting caused by the objective function and small-scale available benchmarks, resulting in insufficient model training. To overcome the overfitting problem, some CSLR studies enforce the sequence module to learn more visual temporal information or be guided by more informative supervision to refine its representations. In this work, we propose a novel autoencoder-formed conditional diffusion feature refinement~(ACDR) to refine the sequence representations to equip desired properties by learning the encoding-decoding optimization process in an end-to-end way. Specifically, for the ACDR, a noising Encoder is proposed to progressively add noise equipped with semantic conditions to the sequence representations. And a denoising Decoder is proposed to progressively denoise the noisy sequence representations with semantic conditions. Therefore, the sequence representations can be imbued with the semantics of provided semantic conditions. Further, a semantic constraint is employed to prevent the denoised sequence representations from semantic corruption. Extensive experiments are conducted to validate the effectiveness of our ACDR, benefiting state-of-the-art methods and achieving a notable gain on three benchmarks.
翻译:在本工作中,我们致力于利用去噪扩散模型的成功,将特征精化构建为自编码器形式的扩散过程。当前最先进的连续手语识别框架包含空间模块、视觉模块、序列模块和序列学习函数。然而,该框架因目标函数和可用小规模基准而导致序列模块过拟合,使得模型训练不充分。为解决过拟合问题,部分连续手语识别研究强制序列模块学习更多视觉时间信息,或通过更具信息量的监督来精化其表示。为此,我们提出一种新型的自编码器形式条件扩散特征精化方法,通过端到端的方式学习编码-解码优化过程,赋予序列表示所需特性。具体而言,该方法中,加噪编码器逐步向序列表示添加携带语义条件的噪声,而去噪解码器则逐步对带有语义条件的含噪序列表示进行去噪。由此,序列表示可被赋予所提供语义条件的语义信息。此外,采用语义约束防止去噪后的序列表示发生语义损坏。大量实验验证了我们方法的有效性,在三个基准上提升了现有最优方法的性能并取得了显著增益。