We study the problem of multi-bit watermarking for non-autoregressive large language models (LLMs). We introduce an information-theoretic model inspired by diffusion language models (DLM), in which the encoder has limited non-causal access to token distributions within each token block. This formulation enables an information-theoretic characterization of the non-causal watermarking capacity, in which knowledge of LLM cover statistics is leveraged to enable a multi-bit covert embedding. We study the information-theoretic limits of the model by combining Gelfand--Pinsker and channel synthesis coding techniques and obtain an exact characterization of the capacity. The embedding strategy is further optimized across blocks using a constrained Markov decision process (CMDP) and we develop an explicit algorithm based on polar codes following the information-theoretic principles. We simulate the error performance on LLaDA, and provide empirical total variation (TV) analysis as a function of key randomness.
翻译:暂无翻译