We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.
翻译:我们提出AnyUser,一种统一的机器人指令系统,通过相机图像上的自由手绘草图(可选择辅以语言)实现直观的家用任务指令下达。该系统将多模态输入(草图、视觉、语言)解释为空间语义基元,无需预先地图或模型即可生成可执行的机器人动作。其新颖组件包括用于理解的多模态融合模块和用于鲁棒动作生成的分层策略。通过广泛评估验证了系统有效性:(1)在大规模数据集上的定量基准测试表明,在多种模拟家庭场景中,系统对多样化基于草图的指令具有高精度解释能力。(2)在两个不同机器人平台上的实际环境验证——静态安装的7自由度辅助机械臂(KUKA LBR iiwa)与双臂移动操作平台(Realman RMC-AIDAL)——执行了如定点擦拭与区域清洁等代表性任务,证实了系统在物理环境中锚定指令并可靠执行的能力。(3)一项涵盖不同人口特征(老年人、模拟非语言交流者、低技术素养者)的综合用户研究表明,系统在可用性与任务说明效率上取得显著提升,实现了高任务完成率(85.7%-96.4%)与用户满意度。AnyUser弥合了先进机器人能力与便捷非专业交互需求之间的鸿沟,为适应真实人类环境的实用辅助机器人奠定了基础。