Automated audio captioning is multi-modal translation task that aim to generate textual descriptions for a given audio clip. In this paper we propose a full Transformer architecture that utilizes Patchout as proposed in [1], significantly reducing the computational complexity and avoiding overfitting. The caption generation is partly conditioned on textual AudioSet tags extracted by a pre-trained classification model which is fine-tuned to maximize the semantic similarity between AudioSet labels and ground truth captions. To mitigate the data scarcity problem of Automated Audio Captioning we introduce transfer learning from an upstream audio-related task and an enlarged in-domain dataset. Moreover, we propose a method to apply Mixup augmentation for AAC. Ablation studies are carried out to investigate how Patchout and text guidance contribute to the final performance. The results show that the proposed techniques improve the performance of our system and while reducing the computational complexity. Our proposed method received the Judges Award at the Task6A of DCASE Challenge 2022.


翻译:自动音频描述是一项多模态翻译任务,旨在为给定的音频片段生成文本描述。本文提出了一种全Transformer架构,该架构采用[1]中提出的Patchout方法,显著降低了计算复杂度并避免了过拟合。描述生成部分地以文本形式标注的AudioSet标签为条件,这些标签由预训练分类模型提取,该模型经过微调以最大化AudioSet标签与真实描述之间的语义相似性。为缓解自动音频描述中的数据稀缺问题,我们引入了来自上游音频相关任务的迁移学习和扩充的领域内数据集。此外,我们提出了一种将Mixup增强应用于自动音频描述的方法。通过消融实验,我们研究了Patchout和文本引导对最终性能的贡献。实验结果表明,所提出的技术提升了系统性能,同时降低了计算复杂度。我们的方法在DCASE 2022挑战赛的Task6A中获得了评审团奖。

0
下载
关闭预览

相关内容

【AAAI2022】用于视觉常识推理的场景图增强图像-文本学习
专知会员服务
50+阅读 · 2021年12月20日
专知会员服务
30+阅读 · 2021年7月30日
【ICML2020】统一预训练伪掩码语言模型
专知会员服务
27+阅读 · 2020年7月23日
【Google AI】开源NoisyStudent:自监督图像分类
专知会员服务
55+阅读 · 2020年2月18日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
69+阅读 · 2022年6月13日
Arxiv
11+阅读 · 2022年3月16日
Arxiv
19+阅读 · 2021年4月8日
Arxiv
17+阅读 · 2021年3月29日
Arxiv
19+阅读 · 2018年3月28日
VIP会员
最新内容
综述 | 多模态大模型的不确定性感知决策
专知会员服务
0+阅读 · 今天14:39
综述 | Deep Academic Survey:有状态闭环综述自动化
专知会员服务
0+阅读 · 今天14:37
综述 | 触觉与力觉感知机器人学习
专知会员服务
5+阅读 · 8月18日
综述 | AI科学家的过去与未来
专知会员服务
5+阅读 · 8月18日
论文 | OmniScientist:全模态全学科AI科学家
专知会员服务
10+阅读 · 8月16日
无人机已改变战场,但并未解决指挥问题
专知会员服务
11+阅读 · 8月14日
相关VIP内容
相关资讯
相关论文
Arxiv
69+阅读 · 2022年6月13日
Arxiv
11+阅读 · 2022年3月16日
Arxiv
19+阅读 · 2021年4月8日
Arxiv
17+阅读 · 2021年3月29日
Arxiv
19+阅读 · 2018年3月28日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员