Text-to-vibration generation converts natural language into haptic feedback, enabling vibration-effect designers to get scenarios-fitted vibrations more efficiently, which shows great potentials in application fields such as metaverse, games, and film to enrich the user experience in interactive scenarios. The core challenge in this field is how to generate accurate, consistent, and complete vibrations according to textual semantics. Very recent autoregressive (AR) approaches (e.g., HapticGen) exhibit limited capacity in fully capturing global dependencies, owing to the inherent sequential nature of their modeling and prevailing data constraints. In this paper, we proposed HapticLDM, the first text-to-vibration generative model built upon Latent Diffusion Models (LDMs). Firstly, with respect to the data, we introduced a text-processing strategy that emphasizes dynamic characteristics to curate high-quality data pairs for fine-grained dynamic modeling. Secondly, HapticLDM incorporates a global denoising mechanism that regulates coherent and stable variations in the temporal envelope. Furthermore, we conduct extensive evaluations, including A/B testing against the state-of-the-art baseline and a user study involving 30 participants. The results demonstrate that our model enhances realism and semantic alignment. Qualitative feedback further indicates that HapticLDM simplifies the haptic design workflow while generating diverse, subtle, and physically precise vibrations.
翻译:文本到振动生成将自然语言转换为触觉反馈,使振动效果设计人员能够更高效地获取适配场景的振动模式,在元宇宙、游戏、影视等应用领域展现出巨大潜力,旨在丰富交互场景中的用户体验。该领域的核心挑战在于如何根据文本语义生成精确、一致且完整的振动信号。近期流行的自回归方法(例如HapticGen)受限于其固有的序列建模特性及现有数据约束,在捕捉全局依赖关系方面能力有限。本文提出了HapticLDM——首个基于潜在扩散模型构建的文本到振动生成模型。首先,在数据层面,我们引入了一种强调动态特性的文本处理策略,以精心构建高质量数据对,实现细粒度动态建模。其次,HapticLDM采用全局去噪机制,调节时间包络中的连贯稳定变化。此外,我们进行了广泛评估,包括与最先进基线的A/B测试及包含30名参与者的用户研究。结果表明,我们的模型在真实感与语义对齐方面均有提升。定性反馈进一步表明,HapticLDM在生成多样、精细且物理精确的振动的同时,简化了触觉设计工作流程。