Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusion transformer heads decode the hidden states into frame-level audio and video latents. Experiments on talking portrait benchmarks show Talker-T2AV outperforms dual-branch baselines in lip-sync accuracy, video quality, and audio quality, achieving stronger cross-modal consistency than cascaded pipelines.
翻译:联合音频-视频生成模型已证明,与级联方法相比,统一生成能产生更强的跨模态一致性。然而,现有模型通过遍布整个去噪过程的广泛注意力机制将模态耦合在一起,以完全纠缠的方式处理高层语义与低层细节。这种方法对口型合成而言并非最优:尽管音频与面部运动在语义上相关,但它们的低层实现(声学信号与视觉纹理)遵循不同的渲染过程。在所有层级上强制联合建模会导致不必要的纠缠并降低效率。我们提出Talker-T2AV,一种自回归扩散框架,其中高层跨模态建模在共享主干网络中进行,而低层细化则使用模态专用解码器。一个共享的自回归语言模型在统一的块级标记空间中联合推理音频和视频。两个轻量级扩散Transformer头将隐藏状态解码为帧级音频和视频潜变量。在说话人肖像基准上的实验表明,Talker-T2AV在唇形同步准确性、视频质量和音频质量方面均优于双分支基线,实现了比级联流水线更强的跨模态一致性。