While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQA) particularly underscores this, presenting unique challenges with its continuous, densely layered audio-visual content, intricate temporal dynamics, and the critical need for domain-specific knowledge. Through a systematic analysis of Music AVQA datasets and methods, this paper identifies that specialized input processing, architectures incorporating dedicated spatial-temporal designs, and music-specific modeling strategies are critical for success in this domain. Our study provides valuable insights for researchers by highlighting effective design patterns empirically linked to strong performance, proposing concrete future directions for incorporating musical priors, and aiming to establish a robust foundation for advancing multimodal musical understanding. We aim to encourage further research in this area and provide a GitHub repository of relevant works: https://github.com/WenhaoYou1/Survey4MusicAVQA.
翻译:尽管近期多模态大语言模型在通用多模态任务上表现出了显著能力,但音乐等专业领域仍需定制化方法。音乐视听问答(Music AVQA)尤其凸显了这一点,其连续且密集交织的视听内容、复杂的时序动态特性以及对领域特定知识的迫切需求构成了独特挑战。通过对Music AVQA数据集及方法的系统分析,本文指出专门化的输入处理、融合时空专用设计的架构以及音乐特定建模策略是此领域成功的关键。本研究通过归纳与高性能表现相关的有效设计模式,提出融合音乐先验知识的具体未来方向,并致力于为推进多模态音乐理解奠定坚实基础。我们旨在鼓励该领域的进一步研究,并提供相关文献的GitHub仓库:https://github.com/WenhaoYou1/Survey4MusicAVQA。