Audio captioning aims to generate text descriptions from environmental sounds. One challenge of audio captioning is the difficulty of the generalization due to the lack of audio-text paired training data. In this work, we propose a simple yet effective method of dealing with small-scaled datasets by leveraging a pre-trained language model. We keep the language model frozen to maintain the expressivity for text generation, and we only learn to extract global and temporal features from the input audio. To bridge a modality gap between the audio features and the language model, we employ mapping networks that translate audio features to the continuous vectors the language model can understand, called prefixes. We evaluate our proposed method on the Clotho and AudioCaps dataset and show our method outperforms prior arts in diverse experimental settings.


翻译:音频字幕生成旨在从环境声音中生成文本描述。其中一个挑战是由于缺乏音频-文本配对训练数据而导致泛化困难。在这项工作中,我们提出了一种简单而有效的方法,通过利用预训练语言模型来处理小规模数据集。我们保持语言模型冻结以维持文本生成的表达能力,仅学习从输入音频中提取全局和时序特征。为了弥合音频特征与语言模型之间的模态差距,我们采用了映射网络,将音频特征转换为语言模型能够理解的连续向量,称为前缀。我们在Clotho和AudioCaps数据集上评估了所提出的方法,并展示了在多种实验设置下,我们的方法优于现有技术。

0
下载
关闭预览

相关内容

【CVPR2022】三元组对比学习的视觉-语言预训练
专知会员服务
33+阅读 · 2022年3月3日
【CVPR2021】基于端到端预训练的视觉-语言表征学习
专知会员服务
38+阅读 · 2021年4月9日
【ICML2020】文本摘要生成模型PEGASUS
专知会员服务
35+阅读 · 2020年8月23日
【论文推荐】文本摘要简述
专知会员服务
69+阅读 · 2020年7月20日
100+篇《自监督学习(Self-Supervised Learning)》论文最新合集
专知会员服务
167+阅读 · 2020年3月18日
文本+视觉,多篇 Visual/Video BERT 论文介绍
AI科技评论
22+阅读 · 2019年8月30日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
ResNet, AlexNet, VGG, Inception:各种卷积网络架构的理解
全球人工智能
20+阅读 · 2017年12月17日
可解释的CNN
CreateAMind
18+阅读 · 2017年10月5日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
4+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2023年5月19日
VIP会员
最新内容
综述 | 多模态大模型的不确定性感知决策
专知会员服务
0+阅读 · 今天14:39
综述 | Deep Academic Survey:有状态闭环综述自动化
专知会员服务
0+阅读 · 今天14:37
综述 | 触觉与力觉感知机器人学习
专知会员服务
6+阅读 · 8月18日
综述 | AI科学家的过去与未来
专知会员服务
6+阅读 · 8月18日
论文 | OmniScientist:全模态全学科AI科学家
专知会员服务
10+阅读 · 8月16日
无人机已改变战场,但并未解决指挥问题
专知会员服务
11+阅读 · 8月14日
相关基金
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
4+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员