Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines. However, existing video models struggle to jointly achieve high expressiveness, real-time inference, and long-horizon identity stability, a tension we call the performance trilemma. Conversation is the most comprehensive performance scenario, as characters simultaneously speak, listen, react, and emote while maintaining identity over time. To address this, we present LPM 1.0 (Large Performance Model), focusing on single-person full-duplex audio-visual conversational performance. Concretely, we build a multimodal human-centric dataset through strict filtering, speaking-listening audio-video pairing, performance understanding, and identity-aware multi-reference extraction; train a 17B-parameter Diffusion Transformer (Base LPM) for highly controllable, identity-consistent performance through multimodal conditioning; and distill it into a causal streaming generator (Online LPM) for low-latency, infinite-length interaction. At inference, given a character image with identity-aware references, LPM 1.0 generates listening videos from user audio and speaking videos from synthesized audio, with text prompts for motion control, all at real-time speed with identity-stable, infinite-length generation. LPM 1.0 thus serves as a visual engine for conversational agents, live streaming characters, and game NPCs. To systematically evaluate this setting, we propose LPM-Bench, the first benchmark for interactive character performance. LPM 1.0 achieves state-of-the-art results across all evaluated dimensions while maintaining real-time inference.
翻译:表现——通过视觉、声音和时间行为外化意图、情感与个性——是使角色栩栩如生的关键。从视频中学习此类表现是传统三维制作流程的一个有前景的替代方案。然而,现有视频模型难以同时实现高表现力、实时推理与长期身份稳定性,我们将这一矛盾称为“表现三难困境”。对话是最全面的表现场景,因为角色需要同时进行说话、倾听、反应和情感表达,并在时间推移中维持身份一致性。为解决此问题,我们提出LPM 1.0(大型表现模型),专注于单人全双工视听对话式表现。具体而言,我们通过严格筛选、听说音视频配对、表现理解以及身份感知的多参考提取,构建了一个多模态以人为中心的数据集;训练了一个拥有170亿参数的扩散Transformer(基础版LPM),通过多模态条件实现高度可控且身份一致的表现;并将其蒸馏为因果流式生成器(在线版LPM),用于低延迟、无限长度的交互。在推理时,给定包含身份感知参考的角色图像,LPM 1.0可从用户音频生成倾听视频,从合成音频生成说话视频,并通过文本提示控制动作,所有过程均以实时速度实现身份稳定、无限长度的生成。因此,LPM 1.0可作为对话智能体、直播角色和游戏NPC的视觉引擎。为系统评估此设定,我们提出了LPM-Bench——首个用于交互式角色表现的基准测试。LPM 1.0在所有评估维度上均达到最先进水平,同时保持实时推理。