Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity reconstruction, strong predictability under a language model, and cross-track correspondence. We introduce DuoTok, a source-aware dual-track tokenizer that addresses this trade-off through staged disentanglement. DuoTok first pretrains a semantic encoder, then regularizes it with multi-task supervision, freezes the encoder, and applies hard dual-codebook routing while keeping auxiliary objectives on quantized codes. A diffusion decoder reconstructs high-frequency details, allowing tokens to focus on structured information for sequence modeling. On standard benchmarks, DuoTok achieves a favorable predictability-fidelity trade-off, reaching the lowest cnBPT while maintaining competitive reconstruction at 0.75 kbps. Under a held-constant dual-track language modeling protocol, enBPT also improves, indicating gains beyond codebook size effects. Controlled diagnostics show larger predictability costs under cross-track corruption and larger gains from longer context, suggesting that models trained on DuoTok tokens use cross-track structure and non-local history.
翻译:音频分词连接了连续波形与多轨音乐语言模型。在双轨建模中,标记需同时保持三种特性:高保真重建、语言模型下的强可预测性,以及跨轨对应关系。我们提出DuoTok,一种源感知的双轨分词器,通过分阶段解耦应对这一权衡。DuoTok首先预训练语义编码器,随后通过多任务监督进行正则化;冻结编码器后,在量化码上保持辅助目标的同时应用硬双码本路由。扩散解码器重建高频细节,使标记能够专注于结构化信息以进行序列建模。在标准基准测试中,DuoTok实现了有利的可预测性—保真度权衡,在0.75 kbps码率下达到最低的cnBPT,同时保持具有竞争力的重建性能。在恒定双轨语言建模协议下,enBPT同样得到改善,表明提升超越了码本尺寸效应。对照诊断显示,跨轨破坏导致更高的可预测性代价,而长上下文带来更大的收益,这表明基于DuoTok标记训练的模型充分利用了跨轨结构和非局部历史信息。