Discrete diffusion models offer a simple and stable likelihood-based framework for sequence generation, recently extended to any-length settings via token insertion. Principled reward-guided fine-tuning for any-length discrete diffusion, however, remains largely unexplored. We introduce Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding (A2D2), a unified framework for reward-guided fine-tuning of any-length discrete diffusion models via joint optimization of the insertion and unmasking policies together with a quality-based inference schedule. We derive the Radon-Nikodym derivative for the joint insertion-unmasking path measures, enabling theoretically guaranteed convergence to the intractable reward-tilted sequence distribution without requiring target samples. Building on this, we establish unmasking and insertion quality as tractable approaches for minimizing decoding error and introduce the Adaptive Joint Decoding (AJD) loss, which provably yields the optimal path measure that generates the reward-tilted distribution. Empirically, A2D2 improves reward optimization while enhancing generation flexibility and accuracy over prior fixed-length fine-tuning and inference-time guidance methods.
翻译:离散扩散模型为序列生成提供了一种简单且稳定的基于似然的框架,并最近通过令牌插入扩展至任意长度设置。然而,针对任意长度离散扩散的基于奖励的准则化微调在很大程度上仍未得到探索。我们提出面向任意长度离散扩散的自适应解码微调(A2D2),这是一个通过联合优化插入与掩码策略以及基于质量的推理调度来实现任意长度离散扩散模型奖励引导微调的统一框架。我们推导了联合插入-掩码路径测度的拉东-尼科迪姆导数,使得在无需目标样本的情况下,理论上能保证收敛到难以处理的奖励倾斜序列分布。基于此,我们建立了掩码与插入质量作为最小化解码误差的可处理方法,并引入自适应联合解码(AJD)损失,该损失可证明地生成产生奖励倾斜分布的最优路径测度。实验结果表明,相较于先前的固定长度微调和推理时引导方法,A2D2在提升奖励优化的同时,增强了生成灵活性与准确性。