This paper presents the design and evaluation of a novel multi-level LLM interface for supermarket robots to assist customers. The proposed interface allows customers to convey their needs through both generic and specific queries. While state-of-the-art systems like OpenAI's GPTs are highly adaptable and easy to build and deploy, they still face challenges such as increased response times and limitations in strategic control of the underlying model for tailored use-case and cost optimization. Driven by the goal of developing faster and more efficient conversational agents, this paper advocates for using multiple smaller, specialized LLMs fine-tuned to handle different user queries based on their specificity and user intent. We compare this approach to a specialized GPT model powered by GPT-4 Turbo, using the Artificial Social Agent Questionnaire (ASAQ) and qualitative participant feedback in a counterbalanced within-subjects experiment. Our findings show that our multi-LLM chatbot architecture outperformed the benchmarked GPT model across all 13 measured criteria, with statistically significant improvements in four key areas: performance, user satisfaction, user-agent partnership, and self-image enhancement. The paper also presents a method for supermarket robot navigation by mapping the final chatbot response to correct shelf numbers, enabling the robot to sequentially navigate towards the respective products, after which lower-level robot perception, control, and planning can be used for automated object retrieval. We hope this work encourages more efforts into using multiple, specialized smaller models instead of relying on a single powerful, but more expensive and slower model.
翻译:本文提出并评估了一种用于超市机器人辅助顾客的新型多层级LLM界面。该界面允许顾客通过通用查询和特定查询两种方式表达需求。尽管如OpenAI的GPTs等前沿系统具有高度适应性且易于构建和部署,它们仍面临响应时间增加、以及对底层模型进行策略控制以实现定制化用例和成本优化方面的局限。以开发更快速、更高效对话智能体为目标,本文主张采用多个经过微调的、更小规模的专业化LLM,根据用户查询的具体程度和意图类型分别处理不同问题。我们将该方法与基于GPT-4 Turbo构建的专业GPT模型进行对比,通过平衡设计的被试内实验,采用人工社交代理问卷(ASAQ)与定性参与者反馈进行评估。实验结果表明,我们提出的多LLM聊天机器人架构在所有13项评估指标上均优于基准GPT模型,其中在性能表现、用户满意度、用户-代理协作关系及自我形象提升四个关键领域取得了统计学意义上的显著改进。本文还提出一种超市机器人导航方法,通过将最终聊天机器人响应映射至正确货架编号,使机器人能够依次导航至对应商品区域,随后可利用底层机器人感知、控制与规划技术实现自动化物品取放。我们希望这项工作能推动更多研究转向采用多个专业化小型模型,而非依赖单一性能强大但成本更高、响应更慢的模型。