The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine tune the components of the AR assistant.
翻译:增强现实(AR)助手的概念几十年来一直吸引着人类的想象力,成为现代科幻小说的核心元素。为实现这一目标,需要开发基于人工智能(AI)的方法,这些方法需同时实时感知三维环境、推理物理任务并对执行者进行建模。在此框架下,需要多种传感器生成跨不同模态的数据,例如音频、视频、深度、语音和飞行时间数据。所需的传感器通常是AR头戴设备的一部分,通过视觉、音频和触觉反馈实现执行者感知和交互。AI助手不仅记录执行者的活动,还需要机器学习(ML)模型来理解并协助执行者与物理世界交互。因此,开发此类助手是一项具有挑战性的任务。我们提出ARGUS,一个支持智能AR助手开发的可视化分析系统。该系统是可视化研究人员与ML及AR专家多年合作设计的一部分。这一共同设计过程推动了AR中ML可视化的进展。我们的系统支持对象、动作和步骤检测的在线可视化,以及之前记录的AR会话的离线分析。它不仅可视化多模态传感器数据流,还可视化ML模型的输出。这使得开发者能够深入了解执行者的活动以及ML模型,帮助他们排查问题、改进并微调AR助手的各个组件。