Low-precision pretraining (FP8, MXFP4, NVFP4) is now standard for frontier language models, yet the literature is almost entirely achievability -- algorithms and empirical scaling laws -- with no matching characterization of what is information-theoretically possible. We study a B-bit quantized stochastic first-order oracle: an optimizer interacts for T rounds and receives, each round, a B-bit adaptive public-coin description of its stochastic gradient. Our main contribution is an exact reduction from optimizing a strongly convex quadratic family to interactively compressed Gaussian mean estimation -- under the B-bit oracle the query carries no information, so optimization collapses exactly onto a sequential distributed-estimation problem. This yields two unconditional lower bounds, a communication bound TB = Omega(d) and a statistical bound T = Omega(sigma^2 d / eps^2), and the sharp product-form bound T = Omega((sigma^2 d / eps^2) max{1, d/B}). The product form is also unconditional: a B-bit transcript carries at most O(TB / sigma^2) of Fisher trace about the mean, so bits rather than dimension limit the recoverable information, and combined with the multivariate van Trees inequality this gives the bound directly, without bounded-likelihood-ratio truncation. We give a near-matching achievability result with exact per-round bit accounting under a bounded-dynamic-range oracle, tight up to a logarithmic factor; the lower bound is for truly Gaussian (unbounded) gradients, and closing this oracle gap is left open. A sequential rate-distortion perspective extends the reduction to correlated and drifting oracles and corrects an earlier conjecture: positive noise correlation raises the bound by (1+rho)/(1-rho) rather than relaxing it. The bounds give an information-theoretic baseline for any low-bit gradient path, not an optimality claim about deployed FP4 systems.


翻译:低精度预训练(FP8、MXFP4、NVFP4)已成为前沿语言模型的标准做法,然而现有文献几乎完全集中在可行性研究——算法与经验标度律——缺乏对信息论可能性的匹配刻画。我们研究了一种B比特量化随机一阶预言机:优化器与系统交互T轮,每轮接收一个B比特的自适应公共随机数描述其随机梯度。我们的主要贡献在于,将强凸二次族优化问题精确约化为交互式压缩高斯均值估计问题——在B比特预言机下,查询本身不携带信息,因此优化问题精确塌缩为序列分布式估计问题。由此得到两个无条件下界:通信下界TB = Ω(d)和统计下界T = Ω(σ²d/ε²),以及尖锐的乘积形式下界T = Ω((σ²d/ε²)max{1, d/B})。该乘积形式同样是无条件的:B比特转录本在均值上至多携带O(TB/σ²)的Fisher迹,因此限制可恢复信息的是比特数而非维度,结合多元van Trees不等式可直接得到该下界,无需有界似然比截断。我们给出了近乎匹配的可行性结果,在有限动态范围预言机下实现了精确的每轮比特核算,紧致性仅差对数因子;下界针对真正高斯(无界)梯度,而弥合该预言机差距的问题留待后续研究。序贯率失真视角将约化扩展到相关漂移预言机,并修正了先前的猜想:正噪声相关性将下界提升(1+ρ)/(1-ρ)倍而非放松。该下界为任何低比特梯度路径提供了信息论基准,而非对已部署FP4系统的最优性断言。

0
下载
关闭预览

相关内容

面向多模态智能的下一个Token预测:综述
专知会员服务
26+阅读 · 2024年12月30日
使用有限数据微调语言模型的实用指南
专知会员服务
27+阅读 · 2024年11月18日
【博士论文】基于信息论的泛化理论方法,274页pdf
专知会员服务
54+阅读 · 2024年6月3日
【博士论文】信息论视角下的泛化理论方法,274页pdf
专知会员服务
51+阅读 · 2024年4月28日
绝对干货 | 随机梯度下降算法综述
菜鸟的机器学习
15+阅读 · 2017年10月30日
精品公开课 | 随机梯度下降算法综述
七月在线实验室
13+阅读 · 2017年7月11日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
0+阅读 · 今天15:49
综述 | Self-Evolving Coding Agents:自进化编程智能体
专知会员服务
0+阅读 · 今天13:16
美海军陆战队将三型无人机整合入统一战场网络
专知会员服务
2+阅读 · 今天9:39
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 今天9:12
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 今天9:08
相关VIP内容
面向多模态智能的下一个Token预测:综述
专知会员服务
26+阅读 · 2024年12月30日
使用有限数据微调语言模型的实用指南
专知会员服务
27+阅读 · 2024年11月18日
【博士论文】基于信息论的泛化理论方法,274页pdf
专知会员服务
54+阅读 · 2024年6月3日
【博士论文】信息论视角下的泛化理论方法,274页pdf
专知会员服务
51+阅读 · 2024年4月28日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员