Symbolic-control drum generation requires preserving explicit event timing and dynamics while synthesizing acoustically plausible waveforms. We present Sec2Drum-DAC, a conditional latent-diffusion model for symbolic-to-audio drum rendering. The model conditions on event features sampled in physical time at codec-frame locations and predicts standardized principal-component coordinates of frozen DAC summed-codebook embeddings rather than waveform samples. In the evaluated DAC configuration, 72 principal components capture the observed training-frame summed-latent subspace under the stated SVD threshold, yielding a compact continuous denoising target with a deterministic reconstruction path to the 1024-dimensional DAC latent space before waveform decoding. Across 1,733 held-out four-beat windows, PCA diffusion improves paired spectral and transient metrics over deterministic PCA regression and a symbolic rendering baseline, while direct regression remains stronger on phase-sensitive waveform L1. Auxiliary RVQ cross-entropy improves short-step diffusion on mel error, onset-flux cosine, and waveform L1, with the most favorable trade-offs occurring at 6-25 denoising steps depending on the metric.


翻译:符号控制型鼓点生成需要在合成声学上逼真的波形的同时,保留精确的事件时序与力度。我们提出Sec2Drum-DAC——一种面向符号到音频鼓点渲染的条件潜扩散模型。该模型以物理时间在编解码帧位置采样的事件特征为条件,预测冻结DAC累加码本嵌入的标准化主成分坐标,而非波形样本。在所评估的DAC配置中,72个主成分在指定SVD阈值下捕获了训练帧累加潜变量子空间,从而在波形解码前,为1024维DAC潜空间提供了一条具有确定性重构路径的紧凑连续去噪目标。在1,733个保留的四拍窗口上,PCA扩散在成对的频谱和瞬态指标上优于确定性PCA回归及符号渲染基线,而直接回归在相位敏感的波形L1上仍表现更强。辅助RVQ交叉熵在短步扩散中改进了梅尔误差、起始流余弦相似度及波形L1,根据指标不同,最优折衷出现在6至25个去噪步长区间。

0
下载
关闭预览

相关内容

DAC:Design Automation Conference。 Explanation:设计自动化会议。 Publisher:ACM。 SIT: https://dblp.uni-trier.de/db/conf/dac/
基于扩散模型和流模型的推理时引导生成技术
专知会员服务
17+阅读 · 2025年4月30日
《扩散模型》最新教程,141页ppt
专知会员服务
79+阅读 · 2024年12月2日
去噪扩散概率模型,46页ppt
专知会员服务
64+阅读 · 2023年1月4日
【NeurIPS2022】GENIE:高阶去噪扩散求解器
专知会员服务
18+阅读 · 2022年11月13日
Deformable Kernels,用于图像/视频去噪,即将开源
极市平台
13+阅读 · 2019年8月29日
【泡泡点云时空-PCL源码解读】PCL中的点云配准方法
泡泡机器人SLAM
71+阅读 · 2019年6月16日
【泡泡点云时空-PCL源码解读】ICP点云精配准算法
泡泡机器人SLAM
187+阅读 · 2019年5月22日
使用 FastAI 和即时频率变换进行音频分类
AI研习社
11+阅读 · 2019年5月9日
使用RNN-Transducer进行语音识别建模【附PPT与视频资料】
人工智能前沿讲习班
74+阅读 · 2019年1月29日
变分自编码器VAE:一步到位的聚类方案
PaperWeekly
25+阅读 · 2018年9月18日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
Arxiv
0+阅读 · 5月6日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
8+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
14+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
10+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员