ChatHuman: Language-driven 3D Human Understanding with Retrieval-Augmented Tool Reasoning

Numerous methods have been proposed to detect, estimate, and analyze properties of people in images, including the estimation of 3D pose, shape, contact, human-object interaction, emotion, and more. Each of these methods works in isolation instead of synergistically. Here we address this problem and build a language-driven human understanding system -- ChatHuman, which combines and integrates the skills of many different methods. To do so, we finetune a Large Language Model (LLM) to select and use a wide variety of existing tools in response to user inputs. In doing so, ChatHuman is able to combine information from multiple tools to solve problems more accurately than the individual tools themselves and to leverage tool output to improve its ability to reason about humans. The novel features of ChatHuman include leveraging academic publications to guide the application of 3D human-related tools, employing a retrieval-augmented generation model to generate in-context-learning examples for handling new tools, and discriminating and integrating tool results to enhance 3D human understanding. Our experiments show that ChatHuman outperforms existing models in both tool selection accuracy and performance across multiple 3D human-related tasks. ChatHuman is a step towards consolidating diverse methods for human analysis into a single, powerful, system for 3D human reasoning.

翻译：现有大量方法被提出用于检测、估计和分析图像中人物的属性，包括三维姿态、形状、接触、人-物交互、情绪等。但这些方法各自独立运行，缺乏协同性。针对这一问题，我们构建了语言驱动的人体理解系统——ChatHuman，它能够整合多种不同方法的技术能力。为此，我们微调了一个大型语言模型（LLM），使其能够根据用户输入选择和运用一系列现有工具。通过这种方式，ChatHuman能够融合多个工具的信息，从而比单一工具更准确地解决问题，并利用工具输出提升自身对人体推理的能力。ChatHuman的创新点包括：利用学术文献指导三维人体相关工具的应用，采用检索增强生成模型为处理新工具生成上下文学习样例，以及通过区分和整合工具结果来增强三维人体理解。实验表明，ChatHuman在工具选择准确性和多项三维人体相关任务性能上均优于现有模型。ChatHuman是将多样化人体分析方法整合为一个强大统一的三维人体推理系统的重要探索。

相关内容

TOOLS

关注 1

这个新版本的工具会议系列恢复了从1989年到2012年的50个会议的传统。工具最初是“面向对象语言和系统的技术”，后来发展到包括软件技术的所有创新方面。今天许多最重要的软件概念都是在这里首次引入的。2019年TOOLS 50+1在俄罗斯喀山附近举行，以同样的创新精神、对所有与软件相关的事物的热情、科学稳健性和行业适用性的结合以及欢迎该领域所有趋势和社区的开放态度，延续了该系列。官网链接：http://tools2019.innopolis.ru/

O’Reilly报告：知识图谱崛起——面向现代数据集成和数据结构体系，“The Rise of the Knowledge Graph——Toward Modern Data Integration and the Data Fabric Architecture”

专知会员服务

49+阅读 · 2022年2月18日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

35+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日