Current gesture recognition systems primarily focus on identifying gestures within a predefined set, leaving a gap in connecting these gestures to interactive GUI elements or system functions (e.g., linking a 'thumb-up' gesture to a 'like' button). We introduce GestureGPT, a novel zero-shot gesture understanding and grounding framework leveraging large language models (LLMs). Gesture descriptions are formulated based on hand landmark coordinates from gesture videos and fed into our dual-agent dialogue system. A gesture agent deciphers these descriptions and queries about the interaction context (e.g., interface, history, gaze data), which a context agent organizes and provides. Following iterative exchanges, the gesture agent discerns user intent, grounding it to an interactive function. We validated the gesture description module using public first-view and third-view gesture datasets and tested the whole system in two real-world settings: video streaming and smart home IoT control. The highest zero-shot Top-5 grounding accuracies are 80.11% for video streaming and 90.78% for smart home tasks, showing potential of the new gesture understanding paradigm.
翻译:当前手势识别系统主要专注于识别预定义集合中的手势,未能将这些手势与交互式图形用户界面元素或系统功能建立连接(例如,将“竖拇指”手势与“点赞”按钮相关联)。我们提出GestureGPT,一种新颖的零样本手势理解与接地框架,该框架利用大语言模型(LLM)。基于手势视频中的手部关键点坐标形成手势描述,并将其输入至我们的双智能体对话系统。手势智能体解读这些描述并查询交互上下文(如界面、历史记录、注视数据),由上下文智能体组织并提供相关信息。经过迭代交互后,手势智能体识别用户意图,并将其接地至交互式功能。我们利用公开的第一人称和第三人称手势数据集验证了手势描述模块,并在两种实际场景(视频流媒体和智能家居物联网控制)中测试了整个系统。最高零样本Top-5接地准确率分别为视频流媒体任务80.11%和智能家居任务90.78%,展示了这一新型手势理解范式的潜力。