Human communication is a complex and diverse process that not only involves multiple factors such as language, commonsense, and cultural backgrounds but also requires the participation of multimodal information, such as speech. Large Language Model (LLM)-based multi-agent systems have demonstrated promising performance in simulating human society. Can we leverage LLM-based multi-agent systems to simulate human communication? However, current LLM-based multi-agent systems mainly rely on text as the primary medium. In this paper, we propose SpeechAgents, a multi-modal LLM based multi-agent system designed for simulating human communication. SpeechAgents utilizes multi-modal LLM as the control center for individual agent and employes multi-modal signals as the medium for exchanged messages among agents. Additionally, we propose Multi-Agent Tuning to enhance the multi-agent capabilities of LLM without compromising general abilities. To strengthen and evaluate the effectiveness of human communication simulation, we build the Human-Communication Simulation Benchmark. Experimental results demonstrate that SpeechAgents can simulate human communication dialogues with consistent content, authentic rhythm, and rich emotions and demonstrate excellent scalability even with up to 25 agents, which can apply to tasks such as drama creation and audio novels generation. Code and models will be open-sourced at https://github. com/0nutation/SpeechAgents
翻译:人类通信是一个复杂且多样化的过程,不仅涉及语言、常识和文化背景等多重因素,还需要多模态信息(如语音)的参与。基于大语言模型(LLM)的多智能体系统在模拟人类社会方面展现出令人瞩目的性能。我们能否利用基于LLM的多智能体系统来模拟人类通信?然而,当前基于LLM的多智能体系统主要依赖文本作为主要媒介。本文提出SpeechAgents,一种基于多模态LLM的多智能体系统,专门用于模拟人类通信。SpeechAgents采用多模态LLM作为单个智能体的控制中心,并利用多模态信号作为智能体间交换信息的媒介。此外,我们提出多智能体调优(Multi-Agent Tuning),在不损害通用能力的前提下增强LLM的多智能体能力。为强化并评估人类通信模拟的有效性,我们构建了人类通信模拟基准(Human-Communication Simulation Benchmark)。实验结果表明,SpeechAgents能够模拟内容连贯、节奏自然、情感丰富的人类通信对话,即使面对多达25个智能体也展现出卓越的可扩展性,可适用于戏剧创作与有声小说生成等任务。代码与模型将开源至https://github.com/0nutation/SpeechAgents。