We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere $\mathbb S^{d-1}$. There the von Mises-Fisher (vMF) distribution induces a natural noise process and admits a closed-form conditional score. The conditional velocity is in general intractable. Exploiting the radial symmetry of the vMF density we reduce the continuity equation on $\mathbb S^{d-1}$ to a scalar ODE in the cosine similarity, whose unique bounded solution determines the velocity. The marginal velocity and marginal score on $(\mathbb S^{d-1})^L$ both decompose into posterior-weighted tangent sums that differ only by per-token scalar weights. This gives access to both ODE and predictor-corrector (PC) sampling. The posterior is the only learned object, trained by a cross-entropy loss. Experiments compare the vMF path against geodesic and Euclidean alternatives. The combination of vMF and PC sampling significantly improves results on Sudoku and language modeling.
翻译:我们研究在连续嵌入空间中学习离散序列生成模型的问题。现有方法通常操作于欧几里得空间或概率单纯形上,而本研究转而采用球面 $\mathbb S^{d-1}$ 作为工作空间。在此空间中,von Mises-Fisher(vMF)分布引出一个自然的噪声过程,并具有闭式条件得分。条件速度通常难以处理。利用vMF密度的径向对称性,我们将 $\mathbb S^{d-1}$ 上的连续性方程简化为关于余弦相似度的标量常微分方程,其唯一有界解决定了速度。$(\mathbb S^{d-1})^L$ 上的边缘速度和边缘得分均分解为后验加权正切和,两者仅因每词元的标量权重而异。这同时提供了常微分方程和预测-校正(PC)采样的实现路径。后验是唯一需学习的对象,通过交叉熵损失进行训练。实验将vMF路径与测地线和欧几里得替代方案进行了比较。vMF与PC采样的结合显著提升了数独和语言建模任务的结果。