The spoken language serves as an accessible and efficient interface, enabling non-experts and disabled users to interact with complex assistant robots. However, accurately grounding language utterances gives a significant challenge due to the acoustic variability in speakers' voices and environmental noise. In this work, we propose a novel speech-scene graph grounding network (SGGNet$^2$) that robustly grounds spoken utterances by leveraging the acoustic similarity between correctly recognized and misrecognized words obtained from automatic speech recognition (ASR) systems. To incorporate the acoustic similarity, we extend our previous grounding model, the scene-graph-based grounding network (SGGNet), with the ASR model from NVIDIA NeMo. We accomplish this by feeding the latent vector of speech pronunciations into the BERT-based grounding network within SGGNet. We evaluate the effectiveness of using latent vectors of speech commands in grounding through qualitative and quantitative studies. We also demonstrate the capability of SGGNet$^2$ in a speech-based navigation task using a real quadruped robot, RBQ-3, from Rainbow Robotics.
翻译:口语作为一种便捷高效的交互界面,使非专业用户和残障人士能够与复杂辅助机器人进行交互。然而,由于说话人声音的声学变异性及环境噪声的影响,如何精准对位语言指令仍是一项重大挑战。本文提出一种新颖的语音-场景图对位网络(SGGNet$^2$),该网络通过利用从自动语音识别(ASR)系统中获取的正确识别词与错误识别词之间的声学相似性,实现了对口语指令的鲁棒对位。为融合声学相似性,我们采用NVIDIA NeMo的ASR模型扩展了先前基于场景图的对位网络(SGGNet)。具体方法是将语音发音的潜在向量馈入SGGNet中基于BERT的对位网络。通过定性与定量研究,我们验证了语音指令潜在向量在对位任务中的有效性。此外,我们利用Rainbow Robotics公司的真实四足机器人RBQ-3,在基于语音的导航任务中展示了SGGNet$^2$的实用能力。