T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Zhe Cao,Tao Wang,Jiaming Wang,Yanghai Wang,Yuanxing Zhang,Jialu Chen,Miao Deng,Jiahao Wang,Yubin Guo,Chenxi Liao,Yize Zhang,Zhaoxiang Zhang,Jiaheng Liu

from arxiv, 41 pages, 13 figures, 12 tables. Accepted at ICML 2026

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.

翻译：文本生成音视频（Text-to-Audio-Video, T2AV）旨在从自然语言中合成了时间连贯的视频与语义同步的音频，但其评估仍呈现碎片化状态，通常依赖于单模态指标或范围狭窄的基准测试，难以捕捉跨模态对齐、指令遵循以及在复杂提示下的感知真实感。为解决此局限，我们提出了T2AV-Compass——一个用于全面评估T2AV系统的统一基准，包含500个多样化且复杂的提示，这些提示通过基于分类法的构建流程生成，以确保语义丰富性与物理合理性。此外，T2AV-Compass引入了一种双层评估框架，整合了针对视频质量、音频质量及跨模态对齐的客观信号级指标，以及用于指令遵循与真实感评估的主观MLLM-as-a-Judge协议。对11个代表性T2AV系统的广泛评估表明，即便最强的模型在人类级真实感与跨模态一致性方面仍存在显著差距，且在音频真实感、精细同步、指令遵循等方面持续出现失败。这些结果揭示了未来模型巨大的改进空间，并凸显了T2AV-Compass作为推动文本生成音视频发展的挑战性诊断测试床的价值。

相关内容

MoDELS

关注 46

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

综述：AI生成视频检测，从视觉取证走向事实保真验证

专知会员服务

11+阅读 · 7月14日

文本、视觉与语音生成的自动化评估方法综述

专知会员服务

20+阅读 · 2025年6月15日

IMAGINE-E：最先进文本到图像模型的图像生成智能评估

专知会员服务

13+阅读 · 2025年2月3日

【NeurIPS2024】通过分解编码和条件控制增强文本到视频生成中的运动效果

专知会员服务

14+阅读 · 2024年11月2日