Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use and often degrade noticeably after aggressive acceleration. We present Hallo-Live, a streaming framework for joint audio-visual avatar generation that combines asynchronous dual-stream diffusion with human-centric preference-guided distillation. To reduce articulation lag in causal generation, we introduce Future-Expanding Attention, which allows each video block to access synchronous audio together with a short horizon of future phonetic cues. To mitigate the quality loss of few-step distillation, we further propose Human-Centric Preference-Guided DMD (HP-DMD), which reweights training samples using rewards from visual fidelity, speech naturalness, and audio-visual synchronization. On two NVIDIA H200 GPUs, Hallo-Live runs at 20.38 FPS with 0.94 seconds latency, yielding 16.0x higher throughput and 99.3x lower latency than the teacher model Ovi. Despite this speedup, it retains strong generation quality, reaching comparable VideoAlign overall score and Sync Confidence score while outperforming other accelerated baselines in the overall quality-efficiency trade-off. Qualitative results further show robust generalization across photorealistic, multi-speaker, and stylized scenarios. To the best of our knowledge, Hallo-Live is the first framework to combine streaming dual-stream diffusion with preference-guided distillation for real-time, text-driven audio-visual generation.
翻译:实时文本驱动的联合音视频头像生成需要高保真且精准同步地合成人像视频与语音,但现有的音视频扩散模型在交互场景中仍过于缓慢,且经过激进加速后往往出现明显质量退化。我们提出Hallo-Live——一种联合音视频头像生成的流式框架,该框架将异步双流扩散与人类中心偏好引导蒸馏相结合。为减少因果生成中的发音延迟,我们引入未来扩展注意力机制,使每个视频块能够同步访问音频信息及短时域的未来语音线索。为缓解少步蒸馏带来的质量损失,我们进一步提出人类中心偏好引导DMD(HP-DMD),该方法利用视觉保真度、语音自然度及音视频同步性的奖励值对训练样本进行重加权。在两块NVIDIA H200 GPU上,Hallo-Live以20.38 FPS的帧率运行,延迟仅0.94秒,吞吐量较教师模型Ovi提升16.0倍,延迟降低99.3倍。尽管实现大幅加速,该框架仍保持优异生成质量,在同步对齐总分与同步置信度得分上达到可比水平,并在整体质量-效率权衡中优于其他加速基线。定性结果进一步表明其在逼真、多说话人及风格化场景中均具有鲁棒泛化能力。据我们所知,Hallo-Live是首个将流式双流扩散与偏好引导蒸馏相结合的实时文本驱动音视频生成框架。