Low-precision pretraining (FP8, MXFP4, NVFP4) is now standard for frontier language models, yet the literature is almost entirely achievability -- algorithms and empirical scaling laws -- with no matching characterization of what is information-theoretically possible. We study a B-bit quantized stochastic first-order oracle: an optimizer interacts for T rounds and receives, each round, a B-bit adaptive public-coin description of its stochastic gradient. Our main contribution is an exact reduction from optimizing a strongly convex quadratic family to interactively compressed Gaussian mean estimation -- under the B-bit oracle the query carries no information, so optimization collapses exactly onto a sequential distributed-estimation problem. This yields two unconditional lower bounds, a communication bound TB = Omega(d) and a statistical bound T = Omega(sigma^2 d / eps^2), and the sharp product-form bound T = Omega((sigma^2 d / eps^2) max{1, d/B}). The product form is also unconditional: a B-bit transcript carries at most O(TB / sigma^2) of Fisher trace about the mean, so bits rather than dimension limit the recoverable information, and combined with the multivariate van Trees inequality this gives the bound directly, without bounded-likelihood-ratio truncation. We give a near-matching achievability result with exact per-round bit accounting under a bounded-dynamic-range oracle, tight up to a logarithmic factor; the lower bound is for truly Gaussian (unbounded) gradients, and closing this oracle gap is left open. A sequential rate-distortion perspective extends the reduction to correlated and drifting oracles and corrects an earlier conjecture: positive noise correlation raises the bound by (1+rho)/(1-rho) rather than relaxing it. The bounds give an information-theoretic baseline for any low-bit gradient path, not an optimality claim about deployed FP4 systems.
翻译:低精度预训练(FP8、MXFP4、NVFP4)已成为前沿语言模型的标准做法,然而现有文献几乎完全集中在可行性研究——算法与经验标度律——缺乏对信息论可能性的匹配刻画。我们研究了一种B比特量化随机一阶预言机:优化器与系统交互T轮,每轮接收一个B比特的自适应公共随机数描述其随机梯度。我们的主要贡献在于,将强凸二次族优化问题精确约化为交互式压缩高斯均值估计问题——在B比特预言机下,查询本身不携带信息,因此优化问题精确塌缩为序列分布式估计问题。由此得到两个无条件下界:通信下界TB = Ω(d)和统计下界T = Ω(σ²d/ε²),以及尖锐的乘积形式下界T = Ω((σ²d/ε²)max{1, d/B})。该乘积形式同样是无条件的:B比特转录本在均值上至多携带O(TB/σ²)的Fisher迹,因此限制可恢复信息的是比特数而非维度,结合多元van Trees不等式可直接得到该下界,无需有界似然比截断。我们给出了近乎匹配的可行性结果,在有限动态范围预言机下实现了精确的每轮比特核算,紧致性仅差对数因子;下界针对真正高斯(无界)梯度,而弥合该预言机差距的问题留待后续研究。序贯率失真视角将约化扩展到相关漂移预言机,并修正了先前的猜想:正噪声相关性将下界提升(1+ρ)/(1-ρ)倍而非放松。该下界为任何低比特梯度路径提供了信息论基准,而非对已部署FP4系统的最优性断言。