Symbolic-control drum generation requires preserving explicit event timing and dynamics while synthesizing acoustically plausible waveforms. We present Sec2Drum-DAC, a conditional latent-diffusion model for symbolic-to-audio drum rendering. The model conditions on event features sampled in physical time at codec-frame locations and predicts standardized principal-component coordinates of frozen DAC summed-codebook embeddings rather than waveform samples. In the evaluated DAC configuration, 72 principal components capture the observed training-frame summed-latent subspace under the stated SVD threshold, yielding a compact continuous denoising target with a deterministic reconstruction path to the 1024-dimensional DAC latent space before waveform decoding. Across 1,733 held-out four-beat windows, PCA diffusion improves paired spectral and transient metrics over deterministic PCA regression and a symbolic rendering baseline, while direct regression remains stronger on phase-sensitive waveform L1. Auxiliary RVQ cross-entropy improves short-step diffusion on mel error, onset-flux cosine, and waveform L1, with the most favorable trade-offs occurring at 6-25 denoising steps depending on the metric.
翻译:符号控制型鼓点生成需要在合成声学上逼真的波形的同时,保留精确的事件时序与力度。我们提出Sec2Drum-DAC——一种面向符号到音频鼓点渲染的条件潜扩散模型。该模型以物理时间在编解码帧位置采样的事件特征为条件,预测冻结DAC累加码本嵌入的标准化主成分坐标,而非波形样本。在所评估的DAC配置中,72个主成分在指定SVD阈值下捕获了训练帧累加潜变量子空间,从而在波形解码前,为1024维DAC潜空间提供了一条具有确定性重构路径的紧凑连续去噪目标。在1,733个保留的四拍窗口上,PCA扩散在成对的频谱和瞬态指标上优于确定性PCA回归及符号渲染基线,而直接回归在相位敏感的波形L1上仍表现更强。辅助RVQ交叉熵在短步扩散中改进了梅尔误差、起始流余弦相似度及波形L1,根据指标不同,最优折衷出现在6至25个去噪步长区间。