When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) $\textit{with emergent creative capabilities}$. The core idea of an AM is to reliably recover stored data points as $\textit{memories}$ by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of $\textit{training}$ and $\textit{test}$ examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.
翻译:语言扩散模型何时会记忆训练数据,以及如何定量评估其真实生成模式?我们通过证明基于均匀分布的离散扩散模型(UDDMs)本质上表现为具有涌现创造能力的联想记忆(AMs)来解答这些问题。联想记忆的核心思想是通过建立围绕数据点的独特吸引域,可靠地恢复存储的数据点作为记忆。历史上,像Hopfield网络这样的模型使用显式能量函数来保证这些稳定吸引子。我们扩展了这种观点,利用能量并非严格必要这一观察,因为吸引域也可以通过条件似然最大化形成。通过评估训练和测试示例的令牌恢复能力,我们在UDDMs中识别出一个由训练数据集大小驱动的尖锐记忆-泛化转变:随着数据集增大,训练示例周围的吸引域缩小,未见测试示例周围的吸引域扩大,直到两者最终收敛到相同水平。关键的是,我们仅通过预测令牌序列的条件熵即可检测这种转变:记忆阶段的特征为条件熵趋近于零,而泛化阶段中大多数令牌的条件熵保持有限。因此,条件熵为部署模型中的记忆-泛化转变提供了实用探测工具。